Recommended Free Tools
A playbook and a runbook do different jobs, and responders who blur them start fixing things before they know what broke. A playbook guides discovery and narrows the search toward a root cause. A runbook gives the steps to mitigate a cause you already understand. For failures in CLI agents and sandboxes built on the OpenAI Agents API, the working order is: identify which layer failed, check whether the session is still usable, and only then decide whether to retry, repair, or recreate.
Playbook or runbook: which document a scenario needs
AWS’s operational guidance defines playbooks by their investigative purpose: “Playbooks are step-by-step guides used to investigate an incident.” (AWS Well-Architected Framework, OPS07-BP04). The runbook is the companion document that describes mitigation once the cause is understood. Use the comparison below to decide which one a scenario needs and what each must contain.
As an Amazon Associate I earn from qualifying purchases.
| Question | Playbook (investigation) | Runbook (mitigation) |
|---|---|---|
| Purpose | Discover what happened and move toward root cause | Resolve a known cause |
| Starting point | A symptom, an alert, or a failure with no explanation yet | A confirmed cause or a recognised failure class |
| Evidence | Logs, events, and request or session identifiers, gathered as discovery proceeds | The cause is established; evidence confirms the scenario matches |
| Tools and permissions | Name any special tools and elevated permissions up front | Name the changes to be made and the permissions needed to make them |
| Expected output | A scoped incident and a stated root cause, or a clean escalation | A restored service or session and a verified result |
| Escalation trigger | The cause is still unknown after the defined discovery steps | Mitigation fails, or the next step falls outside the runbook’s authorisation |
AWS’s GuardDuty material frames the moment after a finding as the question “Now what?” A playbook exists so that the team answers it with defined steps, not improvisation.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →What a scenario runbook must contain
AWS’s guidance for security response playbooks (SEC10-BP04) calls for scenario-specific documents that state a goal, prerequisites, owners, an escalation path, technical response steps, and expected outcomes. The same structure works for operational scenarios, including agent and sandbox failures. Keep each section concrete enough that an on-call engineer can act without reconstructing context.
#1 Best Overall
Overview and goal
Name the alert or symptom, the scope it covers, and what “resolved” means. A goal such as “a new session can complete a turn with its input files” is testable; “fix the agent” is not.
Prerequisites
- The logs you will read and where they live
- The detection mechanism that fired, and the alert text you expect to see
- The tools, dashboards, and access you need, with the permissions confirmed in advance
Contacts, responsibilities, and escalation
- The owner of the scenario and the person who sends stakeholder updates
- The condition that triggers escalation, such as a diagnosis that has not narrowed after the defined discovery steps, and the next contact
Response steps
Every step should say what to inspect, the query or command to run, the result that means “continue,” and the decision that follows. “Check the logs” is not a step. “Retrieve the session and read its status field; if the status is not usable, go to the recreate branch” is.
Expected outcomes
List the observable end states that close the scenario, so that responders know when to stop.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallRank #2
Investigate outside-in before you mitigate
For operational troubleshooting, AWS’s guidance (OPS07-BP04) implies an outside-in order: discover the symptoms, scope the impact, gather evidence, identify the root cause, and then hand off to the mitigation runbook. Two additions make this sequence usable in practice. First, state the stakeholder update plan when the incident is declared, so updates do not depend on whoever is debugging. Second, write the escalation route into the playbook, so that a stalled diagnosis has a defined exit rather than an open-ended hunt.
The five response phases
AWS’s security framework groups response actions into five phases. Treat them as a coverage checklist for each scenario, not as a replacement for scenario-specific commands or authorisation boundaries.
- Detect. Confirm the alert came from the documented mechanism and matches the alert text in the prerequisites.
- Analyze. Scope the affected sessions, environments, and users, and establish what changed.
- Contain. Limit further impact within the authority the runbook grants, before the cause is removed.
- Eradicate. Remove the cause.
- Recover. Restore the affected resource and verify the expected outcome.
What failed: the request, the turn, the session, or the environment?
The OpenAI Agents API exposes failures at four layers, and each has a specific place to look. Classify the layer before changing anything, because the same symptom (an agent that stops producing output) can originate at any of them. The layer table is OpenAI-specific; the sequence is the reusable part.
| Failure layer | Where to inspect | What it indicates |
|---|---|---|
| API request | The HTTP status and the response error object |
The request itself returned an error |
| Turn | Retrieve the turn, then inspect its status and error | A single turn failed at runtime |
| Session | Retrieve the session, then inspect its status and error | The session itself may have failed |
| Environment | The environment error event, then the sandbox troubleshooting guidance | Setup or sandbox execution failed |
Should I retry, repair, or recreate the session?
OpenAI’s “Errors and recovery” documentation states: “A failed turn doesn’t always mean the session has failed.” That sentence sets the order of operations. Check session status before you decide anything else.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problems- The turn failed and the session is still usable. Determine whether the session can continue. If it can, correct the cause before continuing; do not discard a working session.
- The session itself failed. Correct the underlying cause, then create a new session with the inputs it needs.
- The error matches a known class. Apply the specific recovery in the table below.
- The error matches no known class, or the same error returns after correction. Stop retrying, preserve the evidence described later in this article, and escalate.
| Error or symptom | What the guidance points toward | Recovery |
|---|---|---|
| Connection failure or timeout | Executor startup and network access | Diagnose connectivity before any retry; see the sandbox checks below |
sandbox_error |
Setup commands, packages, input files, or environment details | Correct the setup, package, input, or environment issue that the reported error identifies |
| Incompatible executor version | The executor version in use | Upgrade first, then create a new session |
idle_timeout |
The session went idle | Create a new session and supply the inputs again |
| Expired environment | The environment is no longer available to the session | Create a new session and resubmit the inputs |
These error names and recoveries come from OpenAI’s documentation for the Agents API. They are not a vendor-neutral taxonomy and should not be applied as-is to another vendor’s CLI agent.
Sandbox setup, network, and file checks
The environment layer has its own checks. Work through them in the order below, and record what each one returns.
Rank #4
Setup, packages, and input files
When a setup or sandbox execution error occurs, check the setup commands, the packages they install, the input files the run depends on, and the environment error the platform reported. Compare the reported error text with the setup commands line by line before rerunning anything.
Blocked requests and redirects
If a sandbox request is blocked, inspect the network settings first. Then list every host the request reaches, including hosts reached through redirects. A check that examines only the first host can miss the block.
Free tools Windows power users keep installed
One-click scans. No signup required.
Live file operations
Before live file operations, confirm the sandbox is connected. If the environment has expired, the recovery is a new session with the input files resubmitted; reattaching to the expired environment is not a recovery path in the guidance.
Hosted or self-hosted sandbox
OpenAI’s hosted sandbox guide says OpenAI provisions and connects the environment. A self-hosted sandbox is for cases that need a custom image, custom compute, or a private network. The choice changes which controls you have and which failure surfaces you own.
| Dimension | OpenAI-hosted sandbox | Self-hosted sandbox |
|---|---|---|
| Who provisions and connects the environment | OpenAI | Your team |
| When the guidance indicates it | Not stated as a specific use case in the cited guide | Cases requiring a custom image, custom compute, or a private network |
| Image control | Not stated in the cited guide | Custom image supported |
| Network control | Not stated in the cited guide | Private network supported |
| Setup and connectivity checks in this article | Apply; the cited guide does not say which connectivity settings you can change directly | Apply, with the network and image settings under your control |
Preserve evidence before you retry or escalate
Do not retry blindly. Each retry against a failing sandbox or session can change the state you need to diagnose, and it replaces one error with another. Before any retry or escalation, record the following for the incident:
- The observable symptom, in the words the user or system reported
- The event or error identifier, and the HTTP status where one exists
- The affected session or environment identifier
- The change made, with its time
- The outcome you expected from that change
The vendor documentation specifies where to inspect and how to recover, but it does not prescribe this record format. Treat the format as an editorially recommended practice. OpenAI’s guidance does explicitly recommend keeping the request ID when a status or file-list request keeps returning server errors, and that identifier belongs in any escalation.
Validate the runbook before a real incident
AWS states that response arrangements should be validated before an actual incident. Its Incident Detection and Response guidance describes a scheduled GameDay as an end-to-end simulation in which participants can observe how the runbook unfolds and refine its instructions. GameDay scheduling requires advance coordination; confirm the current lead-time requirements on the AWS service page before planning, because they are not reproduced here. The AWS material gives operational scheduling details rather than performance statistics, so it does not establish how much a GameDay improves response times.
Review the runbook whenever one of these changes, and treat this as an operational recommendation rather than a quoted AWS requirement:
Quick Recap
- The workload the runbook covers
- The alerts that trigger it
- The permissions the responders hold
- The tools the steps depend on
- The escalation contacts
What these sources do and do not cover
- AWS guidance covers playbook and runbook structure and AWS service procedures.
- OpenAI guidance covers the Agents API error layers, the named error classes, and sandbox troubleshooting for the hosted and self-hosted options described above.
- Neither source provides a vendor-neutral CLI agent error taxonomy or a universal diagnostic command. The layered triage in this article is a method; the error names and recoveries are OpenAI-specific.
- The official AWS and OpenAI pages were checked on 7 October 2026. Re-check the OpenAI error and sandbox pages before relying on exact error names, since vendor documentation changes.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




