Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix Now×
Skip to content
World desk6 min

How to Reduce Context Usage in Multi-Step AI Automations

A practical guide to reducing model-visible input across multi-step AI automations without losing the state each step needs.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To reduce context usage in a multi-step AI automation, send each model call only the instructions, conversation history, tool definitions, and data it needs for that step. Keep large or reusable material outside the prompt, retrieve relevant portions on demand, trim tool inputs and outputs, and compact stale conversation state when appropriate. Prompt caching can lower the cost of repeated input, but it does not make that input take up less context.

What counts as context in an automation?

Context is the assembled input a model can see for a request—not just the latest prompt. It may include system and developer instructions, the current user turn, earlier messages, application or editor state, referenced files, tool definitions, and tool results. An agent can therefore use more context on every step even when the new instruction is short. Microsoft’s overview of agent context describes these ingredients: Understand context in AI agents.

As an Amazon Associate I earn from qualifying purchases.

Start by inspecting the actual request your application sends at representative points in a run. Attribute token use by category where your provider exposes usage data. Look for repeated stable instructions, irrelevant history, oversized schemas, stale tool results, and large source material that later steps do not need. Optimizing before identifying which of these is growing risks cutting useful information while leaving the main source of bloat untouched.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Reduce what each step receives

Make instructions specific to the task

Use a focused instruction set for the current job rather than placing every possible rule into one universal prompt. Keep shared requirements that truly apply, but avoid carrying unrelated workflow guidance into every request. Keep dynamic, task-specific information distinct from stable instructions so it can be added only when needed.

#1 Best Overall
AMD Ryzen™ AI Halo - Personal AI Desktop Computer - Developer Platform - Linux OS
  • Built for Local AI Development: AMD Ryzen AI Halo is designed for local AI development and inference, featuring 128GB unified memory and support for up to 200B parameter models to build and run intensive AI workloads locally.
  • 128GB Unified Memory: Features 128GB LPDDR5x unified memory at 8000 MT/s with 256 GB/s memory bandwidth, providing a shared memory pool across the CPU, GPU, and NPU to support larger AI models.
  • AMD Ryzen AI Max+ 395 Processor: Features 16 cores, 32 threads, and Zen 5 architecture, paired with AMD Radeon 8060S integrated graphics featuring 40 RDNA 3.5 compute units and an AMD XDNA 2 NPU with up to 50 TOPS.
  • Linux AI Developer Platform: Purpose-built for Linux-based AI development with full AMD ROCm software support and preloaded tools, models, and workflows optimized for local AI development.
  • Compact, Connected Design: Includes a 2TB M.2 SSD, 10GbE LAN, Wi-Fi 7, Bluetooth 5.4, USB-C connectivity, and HDMI 2.1b.

Reference only relevant source material

Include only the files, records, or documents needed for the current decision. For a large corpus, store the material in a filesystem or retrieval layer and ask the model to open or parse the relevant portions on demand. OpenAI describes this approach in its account of giving the Responses API a computer environment: From model to agent: Equipping the Responses API with a computer environment.

This changes the design question from “How can I compress everything?” to “What must this step actually see?” If later work needs only a small excerpt, pass that excerpt or a retrieval pointer instead of repeating the full source. Keep exact values that matter—such as identifiers, code, or constraints—in durable storage or explicitly in the next step’s input.

Keep tools and their results lean

Trim tool definitions and expose tools on demand

Tool descriptions and schemas are part of the model-visible request. Remove redundant wording and fields the model does not need, while preserving required parameters, clear calling instructions, and safety constraints. If a platform supports loading tool definitions on demand, expose only the tools relevant to the current task. Anthropic’s tool-context guide recommends tool search when a toolset grows past roughly 20 tools or when baseline context use becomes noticeable; that is a vendor heuristic, not a universal cutoff: Manage tool context.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Prevent intermediate results from accumulating

Tool outputs become part of the conversation history when they are returned to the model. For small deterministic operations, consider batching work in application code or using a platform feature that performs programmatic tool calls without placing every intermediate result into the conversational transcript. Return concise, structured summaries with identifiers or retrieval pointers when a later step can fetch full details as needed.

Some platforms also provide context editing to remove old tool results after their purpose has passed. These capabilities and their continuation semantics vary by provider, model, and API; check the relevant documentation before designing around them. Anthropic documents tool search, programmatic tool calling, and context editing in its tool context guide.

Compact long-running state carefully

When history has grown too large or become stale, compaction can replace it with a smaller continuation state. OpenAI documents server-side threshold compaction and a standalone compact endpoint. For the standalone endpoint, its output is the canonical next context and should be passed through as returned; for server-side compaction, follow the documented input-array or response-ID chaining pattern rather than manually pruning the transcript. See OpenAI’s compaction guide for the supported patterns.

If you control the summarization instructions, specify what the next step must retain:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • The objective and constraints.
  • Decisions made and completed actions with their outcomes.
  • Exact identifiers, code snippets, and library choices needed to continue.
  • Open questions, blockers, and the next action.

Summaries are not a substitute for durable state when exact data matters. Validate critical values against their source rather than relying on a summary to preserve them perfectly. AWS Bedrock’s compaction guidance, for example, calls out preserving code snippets, library choices, and decisions about retries and rate limiting: Compaction – Amazon Bedrock.

Compaction also has a cost. AWS documents an additional sampling step that affects billing and rate limits, and notes that compaction may be followed by a cache miss. Measure whether the reduction in later context outweighs the added work for your automation.

Start a fresh session for unrelated work

When an automation moves to an unrelated task, do not carry forward a full transcript merely because it is available. Start a new session where the platform permits it, and send a focused handoff only when the work needs to continue. Include the task, applicable constraints, decisions, current result, blockers, and next action. Session boundaries and whether state persists vary by platform, so verify the behavior for the application you use.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Understand what prompt caching does—and does not do

Prompt caching reuses processing for a matching prefix, potentially reducing the cost of repeated input. Cached tokens still occupy the model’s context window. Anthropic states: “Prompt caching doesn’t reduce the number of tokens in context, but it reduces what you pay for them on subsequent requests.” See Anthropic’s tool-context documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To improve the chance of a cache match, keep stable instructions and shared reference material at the beginning of the request, with timestamps and user-specific values later. Append new turns rather than rewriting old ones where the API’s format allows it. A cache hit is not guaranteed, and summarizing, compacting, or truncating history can change the prefix and interrupt reuse. OpenAI’s caching guide says cached input may receive a discount of up to 95%; that figure is model- and pricing-dependent, not a general savings promise: Prompt caching | OpenAI API.

Measure context reduction separately from cost savings

Track the outcomes independently so a lower bill is not mistaken for a smaller request:

  • Context occupancy: input token counts per step, including tool definitions and returned data where reported.
  • Compaction overhead: extra calls, tokens, charges, or rate-limit use associated with compaction where exposed.
  • Cache use: cached-input token counts and cache-read usage reported by the provider.
  • Workflow performance: latency, call count, and whether the next step still has the information it needs.

Compare representative runs before and after a change. A cache hit can reduce repeated processing cost without lowering context occupancy; compaction or selective retrieval targets occupancy but may add latency, lose detail, or affect cache continuity. Provider features, model support, regions, and API continuation requirements differ, so confirm current documentation for the specific deployment.

Choose an approach by its trade-offs

Approach Effect on context Key trade-off
Selective retrieval or references Only the portions needed for the step enter the request. Requires retrieval or parsing logic; relevant information must be found accurately.
Lean tool schemas and concise results Reduces baseline tool-definition size and accumulated result text. Descriptions must remain clear enough for correct, safe tool use.
Batching or programmatic tool calls Can keep intermediate results out of conversational history. Availability and API semantics are platform-specific; batching may reduce visibility into intermediate steps.
Context editing Removes stale material from model-visible context where supported. Deleted details may no longer be available unless stored elsewhere.
Compaction Replaces accumulated history with a smaller continuation state. Adds work and can lose detail or interrupt cache reuse; preserve exact facts deliberately.
Prompt caching Does not reduce context occupancy. Can lower repeated-input processing cost when prefixes match; discounts depend on model and pricing.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.