Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallTo reduce context usage in a multi-step AI automation, send each model call only the instructions, conversation history, tool definitions, and data it needs for that step. Keep large or reusable material outside the prompt, retrieve relevant portions on demand, trim tool inputs and outputs, and compact stale conversation state when appropriate. Prompt caching can lower the cost of repeated input, but it does not make that input take up less context.
What counts as context in an automation?
Context is the assembled input a model can see for a request—not just the latest prompt. It may include system and developer instructions, the current user turn, earlier messages, application or editor state, referenced files, tool definitions, and tool results. An agent can therefore use more context on every step even when the new instruction is short. Microsoft’s overview of agent context describes these ingredients: Understand context in AI agents.
As an Amazon Associate I earn from qualifying purchases.
Start by inspecting the actual request your application sends at representative points in a run. Attribute token use by category where your provider exposes usage data. Look for repeated stable instructions, irrelevant history, oversized schemas, stale tool results, and large source material that later steps do not need. Optimizing before identifying which of these is growing risks cutting useful information while leaving the main source of bloat untouched.
Recommended Free Tools
Reduce what each step receives
Make instructions specific to the task
Use a focused instruction set for the current job rather than placing every possible rule into one universal prompt. Keep shared requirements that truly apply, but avoid carrying unrelated workflow guidance into every request. Keep dynamic, task-specific information distinct from stable instructions so it can be added only when needed.
#1 Best Overall
- Built for Local AI Development: AMD Ryzen AI Halo is designed for local AI development and inference, featuring 128GB unified memory and support for up to 200B parameter models to build and run intensive AI workloads locally.
- 128GB Unified Memory: Features 128GB LPDDR5x unified memory at 8000 MT/s with 256 GB/s memory bandwidth, providing a shared memory pool across the CPU, GPU, and NPU to support larger AI models.
- AMD Ryzen AI Max+ 395 Processor: Features 16 cores, 32 threads, and Zen 5 architecture, paired with AMD Radeon 8060S integrated graphics featuring 40 RDNA 3.5 compute units and an AMD XDNA 2 NPU with up to 50 TOPS.
- Linux AI Developer Platform: Purpose-built for Linux-based AI development with full AMD ROCm software support and preloaded tools, models, and workflows optimized for local AI development.
- Compact, Connected Design: Includes a 2TB M.2 SSD, 10GbE LAN, Wi-Fi 7, Bluetooth 5.4, USB-C connectivity, and HDMI 2.1b.
Reference only relevant source material
Include only the files, records, or documents needed for the current decision. For a large corpus, store the material in a filesystem or retrieval layer and ask the model to open or parse the relevant portions on demand. OpenAI describes this approach in its account of giving the Responses API a computer environment: From model to agent: Equipping the Responses API with a computer environment.
This changes the design question from “How can I compress everything?” to “What must this step actually see?” If later work needs only a small excerpt, pass that excerpt or a retrieval pointer instead of repeating the full source. Keep exact values that matter—such as identifiers, code, or constraints—in durable storage or explicitly in the next step’s input.
Keep tools and their results lean
Trim tool definitions and expose tools on demand
Tool descriptions and schemas are part of the model-visible request. Remove redundant wording and fields the model does not need, while preserving required parameters, clear calling instructions, and safety constraints. If a platform supports loading tool definitions on demand, expose only the tools relevant to the current task. Anthropic’s tool-context guide recommends tool search when a toolset grows past roughly 20 tools or when baseline context use becomes noticeable; that is a vendor heuristic, not a universal cutoff: Manage tool context.
Prevent intermediate results from accumulating
Tool outputs become part of the conversation history when they are returned to the model. For small deterministic operations, consider batching work in application code or using a platform feature that performs programmatic tool calls without placing every intermediate result into the conversational transcript. Return concise, structured summaries with identifiers or retrieval pointers when a later step can fetch full details as needed.
Some platforms also provide context editing to remove old tool results after their purpose has passed. These capabilities and their continuation semantics vary by provider, model, and API; check the relevant documentation before designing around them. Anthropic documents tool search, programmatic tool calling, and context editing in its tool context guide.
Compact long-running state carefully
When history has grown too large or become stale, compaction can replace it with a smaller continuation state. OpenAI documents server-side threshold compaction and a standalone compact endpoint. For the standalone endpoint, its output is the canonical next context and should be passed through as returned; for server-side compaction, follow the documented input-array or response-ID chaining pattern rather than manually pruning the transcript. See OpenAI’s compaction guide for the supported patterns.
Rank #3
If you control the summarization instructions, specify what the next step must retain:
- The objective and constraints.
- Decisions made and completed actions with their outcomes.
- Exact identifiers, code snippets, and library choices needed to continue.
- Open questions, blockers, and the next action.
Summaries are not a substitute for durable state when exact data matters. Validate critical values against their source rather than relying on a summary to preserve them perfectly. AWS Bedrock’s compaction guidance, for example, calls out preserving code snippets, library choices, and decisions about retries and rate limiting: Compaction – Amazon Bedrock.
Compaction also has a cost. AWS documents an additional sampling step that affects billing and rate limits, and notes that compaction may be followed by a cache miss. Measure whether the reduction in later context outweighs the added work for your automation.
Start a fresh session for unrelated work
When an automation moves to an unrelated task, do not carry forward a full transcript merely because it is available. Start a new session where the platform permits it, and send a focused handoff only when the work needs to continue. Include the task, applicable constraints, decisions, current result, blockers, and next action. Session boundaries and whether state persists vary by platform, so verify the behavior for the application you use.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Understand what prompt caching does—and does not do
Prompt caching reuses processing for a matching prefix, potentially reducing the cost of repeated input. Cached tokens still occupy the model’s context window. Anthropic states: “Prompt caching doesn’t reduce the number of tokens in context, but it reduces what you pay for them on subsequent requests.” See Anthropic’s tool-context documentation.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →To improve the chance of a cache match, keep stable instructions and shared reference material at the beginning of the request, with timestamps and user-specific values later. Append new turns rather than rewriting old ones where the API’s format allows it. A cache hit is not guaranteed, and summarizing, compacting, or truncating history can change the prefix and interrupt reuse. OpenAI’s caching guide says cached input may receive a discount of up to 95%; that figure is model- and pricing-dependent, not a general savings promise: Prompt caching | OpenAI API.
Best Value
Measure context reduction separately from cost savings
Track the outcomes independently so a lower bill is not mistaken for a smaller request:
- Context occupancy: input token counts per step, including tool definitions and returned data where reported.
- Compaction overhead: extra calls, tokens, charges, or rate-limit use associated with compaction where exposed.
- Cache use: cached-input token counts and cache-read usage reported by the provider.
- Workflow performance: latency, call count, and whether the next step still has the information it needs.
Compare representative runs before and after a change. A cache hit can reduce repeated processing cost without lowering context occupancy; compaction or selective retrieval targets occupancy but may add latency, lose detail, or affect cache continuity. Provider features, model support, regions, and API continuation requirements differ, so confirm current documentation for the specific deployment.
Quick Recap
Choose an approach by its trade-offs
| Approach | Effect on context | Key trade-off |
|---|---|---|
| Selective retrieval or references | Only the portions needed for the step enter the request. | Requires retrieval or parsing logic; relevant information must be found accurately. |
| Lean tool schemas and concise results | Reduces baseline tool-definition size and accumulated result text. | Descriptions must remain clear enough for correct, safe tool use. |
| Batching or programmatic tool calls | Can keep intermediate results out of conversational history. | Availability and API semantics are platform-specific; batching may reduce visibility into intermediate steps. |
| Context editing | Removes stale material from model-visible context where supported. | Deleted details may no longer be available unless stored elsewhere. |
| Compaction | Replaces accumulated history with a smaller continuation state. | Adds work and can lose detail or interrupt cache reuse; preserve exact facts deliberately. |
| Prompt caching | Does not reduce context occupancy. | Can lower repeated-input processing cost when prefixes match; discounts depend on model and pricing. |
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




