The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Stop sending the model the whole page at every step. Start with a shallow accessibility snapshot, search it for the control you need, then inspect only that control’s subtree. Keep a compact record of the task and current page instead of appending every observation to the conversation. Add screenshots only when the task genuinely depends on visual information.
Why browser agents run out of useful context
Browser automation produces observations as well as actions. A page snapshot, screenshot, error message, and prior conversation all compete for room in the model’s context. If an agent repeatedly appends full-page observations—including navigation, repeated menus, and unrelated content—the prompt grows while the useful signal gets harder to find. Old observations can also become stale after a page changes.
The scale can be substantial: Prune4Web (2025) reports that web-agent DOM structures can range from 10,000 to 100,000 tokens. That is not a universal size for every page or agent; it illustrates why observations should be treated as budgeted inputs rather than free logging.
The practical goal is not to minimize every observation regardless of cost. It is to send enough current evidence to make the next decision, with a visual check when the task requires one, and no more.
#1 Best Overall
Use a progressively narrower observation loop
- Begin with a shallow semantic snapshot. Ask for the page’s accessibility information at a small depth. Playwright Agent CLI documents
snapshot --depth=4as a way to limit output on complex pages. Treat four as a starting example, not a universal ideal. - Search the snapshot before taking another one. If the first result contains the relevant section, use text or a regular-expression search to locate the control. Playwright’s CLI documentation recommends
findwhen the task is locating one element; the result can return matching nodes with nearby context instead of another full tree. - Expand only the relevant subtree. Once you identify a likely section, request its subtree rather than the entire page. This keeps repeated navigation, unrelated lists, and other page regions out of the next model input.
- Act, then observe the changed state. Give the browser a narrow action—such as click, fill, select, or navigate—and return a compact result. Use deterministic waits and URL checks in code where possible rather than asking the model to reconsider unchanged output.
- Re-snapshot after navigation or a meaningful state change. Reuse a ref only while the page state that produced it remains current. Playwright’s guidance says refs are invalidated after navigation; take a fresh snapshot and target the current page state instead of replaying an old ref.
This is progressive disclosure: page-level orientation first, a search result next, and a subtree only if needed to understand or operate the control. If a shallow view does not contain the target, increase depth or scope deliberately instead of defaulting to a full, deep snapshot on every turn.
Choose the representation that fits the task
| Representation | Best use | Cost or limitation |
|---|---|---|
| Accessibility snapshot | Finding and operating controls described by accessible names, roles, and text. | It may omit visual relationships that matter to the task, and a full tree can still be large. |
| Filtered semantic tree or search result | Locating one target or extracting a small, task-relevant region. | A filter can hide context the task actually needs; broaden the scope if the target is ambiguous. |
| Screenshot | Checking layout, canvas, charts, image-heavy content, or an ambiguous icon-only control. | Visual input can be high-token compared with semantic text. A screenshot does not replace a useful control tree for ordinary interaction. |
| Raw HTML | Tasks that explicitly require markup or attributes not represented in the semantic view. | It can include large amounts of markup irrelevant to the next action. Scope or extract narrowly when possible. |
Playwright MCP describes accessibility snapshots as low-token text and screenshots as high image-token inputs. Its approach uses accessibility snapshots instead of screenshots by default. That makes semantic output a sensible default for ordinary controls, not a rule against visual inspection. For a canvas, chart, image-heavy page, or unclear icon, capture a visual view as an exception; use it to resolve the question, then stop carrying it forward once it is no longer needed.
Keep working memory compact and current
Do not use the conversation transcript as the agent’s only state store. Maintain a short working record that replaces older observations as the page changes. A useful record contains:
Rank #2
- Goal: the specific outcome the agent is trying to achieve.
- Page identity: the current URL or a short description of the current page and state.
- Completed actions: only actions that affect what the agent should do next.
- Extracted values: the few values needed to finish the task, with brief evidence if necessary.
- Blockers: what prevented progress, if anything.
- Next decision: the next action or the exact uncertainty to resolve.
After an action changes the page, replace the previous snapshot with a concise summary and the latest relevant evidence. Keep a short excerpt when it justifies the next decision; discard unrelated page content and old screenshots. This compact-state approach is an engineering pattern built on Playwright’s scoped-snapshot and re-snapshot mechanics, not a guarantee that a particular memory format will suit every task.
Filter very large pages by the task
On a long page, a task-guided relevance filter can select accessibility-tree lines that match the current goal before they reach the model. FocusAgent presents this as a way to trim large web-agent context. A filter is useful when the page is too large for a simple depth limit or one search, but it introduces a recall trade-off: if relevant evidence is filtered out, the agent can make a confident decision from an incomplete view.
Keep the filter narrow enough to reduce noise but allow a recovery path. If no likely target appears, widen the search terms, inspect a broader subtree, or increase snapshot depth. Do not silently treat “no match in the filtered result” as proof that an element does not exist.
Rank #3
Measure savings without sacrificing task success
Compare observation strategies on the same tasks and target sites. At minimum, track:
- Input tokens per observation and cumulative context tokens.
- Browser round trips, latency, and retries.
- Stale-ref failures and other recovery events.
- Task success and whether the agent recovered correctly after a miss.
Compare full snapshots with depth-limited snapshots, subtree snapshots, and find-based retrieval. Token count alone is not enough: the smallest prompt is not useful if it causes more missed targets or retries. A 2025 paper, Building Browser Agents, reports approximately 85% success on WebGames across 53 challenges for a hybrid design using accessibility snapshots, selective vision, browser tooling, and prompt engineering. That is a reported benchmark result for those challenges, not a universal success rate or proof that one pruning technique alone caused the result.
Recommended Free Tools
There is no established universal token-reduction percentage, context limit that applies to every model, or guaranteed success improvement from a specific pruning strategy. Measure on the model, sites, and task distribution you actually use.
Rank #4
Troubleshoot common context and targeting failures
- The shallow snapshot does not show the control: Increase depth or inspect the relevant parent subtree. Search using likely visible text or role-related wording before recapturing the whole page.
- Search returns no match: The text may differ from the visible label, or the target may be outside the captured scope. Broaden the search or snapshot scope; do not infer absence from one narrow result.
- A ref or target fails after navigation: It may belong to the old page state. Take a fresh snapshot and identify the control again rather than replaying the stale ref.
- The agent cannot distinguish an icon or visual region: Use a targeted screenshot or visual probe. This is a suitable exception for ambiguous icon-only controls, charts, canvas, or layout-dependent decisions.
- The model keeps reasoning over unchanged output: Move deterministic waits, URL checks, and failure handling into the browser-driving code. Return the result of the action and only the current evidence needed for the next decision.
- Filtering removes needed evidence: Relax the task-guided filter and inspect a broader region. Check task success and recovery along with token counts so pruning does not hide a quality regression.
Or skip the browser setup
For a visual check, ScreenshotNeo can return a webpage screenshot through one GET request; it is a screenshot API, not a replacement for accessibility snapshots or a way to prune an agent’s DOM. Its MCP server also provides screenshot tools for AI agents, including Claude, Cursor, and other MCP clients. Cookie banners are accepted and removed along with 60+ known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies page verdict and billing status in headers.
Example with cURL (see the ScreenshotNeo documentation):
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
ScreenshotNeo offers 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000 shots. Every feature is on every plan. If you want a browser screenshot without setting up the capture locally, sign up for the free plan.
Best Value
Frequently Asked Questions
Does reducing the snapshot depth change what the browser can interact with?
No. Depth limits the observation returned to the agent; it does not by itself remove page content or prevent the browser from interacting with the page.
Should an agent retain a screenshot for later steps?
Only if later decisions still depend on that visual evidence. Otherwise, retain a brief extracted fact or decision and discard the image from working context.
Can one context-pruning strategy be assumed to work across sites and models?
No. Compare token use and task success on the sites, model, and tasks your agent must handle.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

