Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →A long-running agent is a workflow with explicit continuation points. It is not one request that stays open for a long time. To make an agent survive approvals, external events, retries, and process restarts, do four things. Give each run a durable ID. Persist its state at every step boundary. Pick one owner for conversation state. Resume from stored state instead of from a live process. Add a durable orchestration engine only when waits, retries, or restarts exceed what the SDK’s own continuation can cover.
This guide turns “asynchronous” into concrete design decisions. It draws on OpenAI’s current Agents SDK and API documentation, checked on 5 October 2026. It does not name a universal winner among runtimes. The official pages don’t compare cost, latency, or reliability across options, so this article makes no benchmark claims.
What counts as long-running agent work
Agent work becomes “long-running” when it crosses a boundary that one process and one request can’t safely hold. Four boundaries matter in practice:
- Human waits: a person must approve a refund, a deployment, or an email, and may take minutes or days.
- External events: the next step depends on a webhook, a build finishing, or a reply arriving.
- Retries: a model call, tool call, or downstream API fails and must be repeated without repeating completed side effects.
- Process boundaries: a deploy, crash, or autoscale event kills the worker mid-task.
A short task with none of these needs only a normal agent loop. OpenAI’s Agents SDK documentation describes a single run as executing that loop. Anything longer needs a deliberate strategy for carrying state into the next turn (OpenAI Agents SDK, Running agents).
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
The workflow spine: four things every long-running agent needs
1. A durable run ID
Create an identifier you control when the task starts, and attach every approval, event, log line, and retry to it. Without one, you can’t route a late approval or webhook back to the right work.
2. Persisted state at step boundaries
Decide what “state” means for your agent, then write it down before you wait. A practical record usually holds:
- the run ID, status (running, waiting for approval, waiting for event, failed, done), and timestamps;
- the conversation or session reference, or the serialized run state;
- pending actions awaiting a decision, with the arguments the agent proposed;
- a list of completed side effects, each with an idempotency key;
- the sandbox or workspace reference, if the agent uses one.
This list is a design sketch, not a schema from any vendor.
3. Explicit step boundaries
Choose where the agent may stop: before a consequential tool call, after a tool returns, or on receiving a human decision. Persist at each stop. Between stops, the agent can run as an ordinary loop.
Recommended Free Tools
4. A resume path that doesn’t depend on the original process
Any worker should be able to load the stored state and continue. If resuming requires the original process to still be alive, you have a long request, not a long-running workflow.
Choosing a state strategy: who owns the conversation?
The SDK documentation describes two families of continuation. Pick one per run (Running agents).
| Approach | How continuation works | Who stores state | Fits when |
|---|---|---|---|
| Client-managed | Your application carries history forward, or uses SDK sessions | Your application or its session store | You need your own retention rules, audit trail, or database control; you want to inspect or edit history |
| Server-managed | Continuation through conversation IDs or response chaining | The service | You prefer not to store and replay history yourself |
One documented constraint: the SDK says session persistence cannot be combined with server-managed conversation settings in the same run. Mixing them is a configuration error, not a hybrid mode. Decide up front, because switching later means migrating the state of in-flight runs.
Which owner is right depends on your deployment. If your workflow engine or database already holds the authoritative record of each run, client-managed state keeps one source of truth. If the service already holds the conversation and your own record stores only the reference and status, server-managed continuation means less to build. OpenAI’s Agents overview also separates a managed Agents API, an application-run SDK, and direct API use. Where your runtime lives determines who ends up responsible for state (OpenAI API, Agents).
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Rank #3
Approval as a persisted pause, not an open request
Human review can outlast any request timeout and any process. The Agents SDK’s human-in-the-loop guide (in the JavaScript SDK docs) handles this with interruptible runs: the run stops at a step needing approval, its state can be serialized, and you resume it once a decision exists (OpenAI Agents SDK, Human-in-the-loop). The Python and JavaScript documentation are separate pages, so confirm the equivalent behavior in the SDK language you use.
A robust pause/resume path looks like this:
- Run until interruption. The agent proposes a consequential action and the run stops instead of executing it.
- Serialize and store. Save the run state under your run ID with status “waiting for approval”, plus the proposed action and its arguments.
- Release the worker. End the request or free the process. Nothing should block while a person decides.
- Notify the reviewer. Send the proposed action and enough context to decide: a queue item, a message, or a dashboard entry.
- Receive the decision. Record approve or reject, who decided, and when. Make this write idempotent so a double click can’t resume twice.
- Rehydrate and resume. Any worker loads the stored state, applies the decision, and continues the run.
- Handle expiry. Decide what happens if no one answers: escalate, cancel, or keep waiting. Store the deadline with the pending action.
Approvals aren’t the only thing that fits this shape. A webhook or a finished job is the same pattern with a different trigger: stop, persist, release, resume on the event.
When the SDK’s continuation is enough, and when to add a durable engine
You may not need a separate orchestrator. If waits are short, a lost run can be safely restarted, and your application already stores run state, SDK-level continuation plus your own database may suffice.
The official guidance changes when execution itself must survive failure. OpenAI’s wording: “The integrations below are for durable orchestration when runs may span long waits, retries, or process restarts.” (OpenAI API, Running agents). The SDK documentation names Dapr, Temporal, Restate, and DBOS as integrations. The current API guide describes Temporal as supporting durable, long-running workflows, including human-in-the-loop tasks. The pages don’t rank these four or say one is best for every workload.
Rank #4
Signs you’ve outgrown hand-rolled persistence
- You’re writing your own timers, retry schedulers, or “stuck run” sweepers.
- A worker crash mid-run leaves you unsure which side effects already happened.
- Waits run for hours or days and must survive deploys.
- Many concurrent runs need uniform visibility into their state.
Axes for comparing runtime options
Use these axes to compare Dapr, Temporal, Restate, DBOS, or a plain queue-plus-database design. They are decision criteria, not published benchmark results. Fill in the answers from each product’s own documentation and a trial with your workload.
| Axis | Question to answer |
|---|---|
| State ownership | Who stores workflow state, and where does the conversation history live? |
| Crash recovery | Does execution continue after a worker or process restart, and from which point? |
| Retries and duplicates | How are retries configured, and how do you prevent a repeated side effect? |
| Waiting | How do approvals and external events resume a paused run? Can it wait days? |
| Operations | What must your team deploy and run (services, databases, workers)? |
| Isolated execution | Does the agent need a sandbox for files and commands, and how does it fit? |
| Observability | How do you inspect, audit, and evaluate runs? |
Put controls at consequential boundaries
OpenAI’s guardrails guide describes two controls relevant here: input checks before expensive or side-effecting work, and human review for approval decisions (OpenAI API, Guardrails and human review). Long-running work raises the stakes of both. A bad input that passes early can burn hours of runtime or trigger an irreversible action late in the run.
- Validate before you spend. Check inputs before launching a sandbox, calling costly tools, or starting a multi-step plan.
- Require approval where actions are hard to undo. Payments, deletions, outbound messages, and production changes belong behind a persisted pause, not inside an uninterrupted loop.
- Record the decision. Store who approved what, with the exact arguments approved. If arguments change after approval, require a new approval.
Retries and duplicate side effects
Durable engines make retries possible, but they don’t make your tools safe to repeat. The sources don’t say how each integration deduplicates side effects, so check this for your chosen runtime. As a general design rule, give every side-effecting tool call an idempotency key derived from the run ID and step, and check your completed-effects record before acting. Then a resumed run can replay a step without charging a card twice.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Use a sandbox when the agent needs a workspace
If the agent has to read and write files, run commands, install packages, or reach the network under control, run it in an isolated environment instead of your application host. OpenAI’s sandbox guide covers this case. It also describes snapshots and resumable state, which suit work that pauses for review or a later event (OpenAI API, Sandbox agents).
Best Value
Treat the sandbox as part of your workflow state. Store its reference next to the run record. Decide how long a paused workspace lives. Make sure approvals refer to what is actually in the workspace at resume time.
Observability and evaluation
A paused run is invisible unless you make it visible. At minimum, log every state transition under the run ID: started, interrupted, approved or rejected, resumed, retried, failed, completed. Track the age of runs sitting in each waiting state so you can spot stalled approvals or missed events. Keep tool inputs and outputs for audit, subject to your own privacy and retention rules.
For evaluation, test the resume path itself, not just the agent’s answers. Kill a worker mid-run and confirm the run completes. Deliver an approval twice and confirm only one resume happens. Replay a step and confirm no duplicate side effect. Measure cost and latency on your own workload. The official pages reviewed publish no comparative figures, so any number you rely on should come from your own tests.
Workload-based selection checklist
| If your workload looks like this | Start with |
|---|---|
| Seconds to minutes, safe to restart, no approvals | A plain SDK run with your own timeout and retry handling |
| Occasional approvals, waits of minutes to hours, your app already has a database | Interruptible runs, serialized state in your store, and a resume endpoint; pick client-managed or server-managed state, not both |
| Waits of days, many concurrent runs, crashes that must not lose progress | A durable workflow engine such as Temporal, Dapr, Restate, or DBOS, compared on the axes above |
| Agent edits files or executes commands | A sandbox, with its reference stored in the run record |
| Irreversible side effects | Input validation up front, approval before the action, idempotency keys on every call |
Before you ship, confirm each of these:
- Every run has an ID you generate, and every event and approval carries it.
- You chose one state owner and verified the configuration doesn’t mix sessions with server-managed conversation settings.
- The agent can stop at defined boundaries, and state is stored before each wait.
- Any worker can resume a run, and resume is idempotent.
- Waits have deadlines and an expiry policy.
- Side-effecting tools are protected against duplicates.
- Guardrails run before expensive or consequential work.
- You’ve tested crash, duplicate approval, and replay cases.
These docs change often. Check the linked pages for current API names and behavior before you build.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




