A trace can show that an AI agent called a model, ran a tool, and stopped. It cannot, by itself, prove that the agent completed the requested work or delivered the result where the user expects it. There are status vocabularies for particular systems and stages, but the sources available here do not establish one finalized, universal vocabulary for long-running agent tasks.
Why a finished trace is not proof of a finished task
“Done” can describe several different events, and they are not interchangeable:
As an Amazon Associate I earn from qualifying purchases.
- A span ended: one recorded operation, such as a model request or tool call, has ended.
- A tool call succeeded: a tool returned successfully. That says something about the operation, not necessarily about the user’s request.
- The agent run reached a terminal state: the run stopped, whether because it completed, failed, was cancelled, or timed out.
- The deliverable was verified: the requested result exists and is visible on the surface the consumer expects.
These distinctions matter especially when a task involves multiple steps or asynchronous handoffs. A green indicator for one successful operation can coexist with an incomplete task. A useful completion signal must describe the task outcome, not merely the last event observed.
What existing status vocabularies do—and do not—cover
OpenTelemetry’s CI/CD task conventions
OpenTelemetry’s CI/CD semantic conventions include task results such as success, failure, error, skip, cancellation, and timeout, as well as pipeline states pending, executing, and finalizing. The page labels these conventions Release Candidate. They provide useful vocabulary in the CI/CD context; they do not establish a universal lifecycle for every long-running AI-agent task.
#1 Best Overall
Agent Arc’s proposed long-running work phases
The Agent Arc Status Protocol v0.2, a draft last updated June 14, 2026, proposes started, milestone, heartbeat, done, and blocked phases for a unit of authorized agent work. Its motivation is that teams can otherwise create siloed progress surfaces that do not interoperate; that is the draft’s rationale, not a measured industry statistic.
The draft puts a particularly important condition on its terminal signal: “An emitter MUST verify completion from the consumer’s vantage point before emitting done (i.e. the deliverable is visible on the surface the consumer expects).” Under that draft, reporting an incomplete task as done is a conformance violation. This is a proposal, not a finalized standard.
Rank #2
Telemetry-management status is a different layer
OpAMP is OpenTelemetry’s agent-management protocol. It concerns management and status reporting for telemetry collection agents, including package-installation outcomes; it does not define whether an AI assistant fulfilled a user’s request. OpAMP is marked Beta.
Recommended Free Tools
A broader IETF draft is still a working document
The Agent Runtime Telemetry System is an IETF Internet-Draft dated July 2026. It describes a broad telemetry framework, including task-completion and output-validation signals, but it is not a finalized IETF standard. The draft indicates an expiration date of January 7, 2027.
What a useful agent-completion signal should record
The following is a practical design synthesis, not a standardized field schema. Keep detailed execution traces separate from a task-lifecycle record when traces alone cannot answer whether the user’s work is complete.
- Stable task identifier: preserve the same identifier across agent steps, tools, and asynchronous boundaries so updates can be correlated.
- Lifecycle timestamps: record when the task started and when it was updated, and distinguish updates from terminal outcomes.
- Progress events: emit meaningful milestones or heartbeats for work that takes time. A heartbeat indicates activity, not success.
- Explicit non-success states: represent blocked work and terminal failure distinctly from completion; include cancellation or timeout where they apply.
- Terminal outcome: state whether the task succeeded, failed, or ended for another reason instead of inferring the result from the final span.
- Consumer-side completion check: verify the deliverable on the surface the expected consumer uses before emitting
done.
The Agent Arc draft gives a default cadence floor of five minutes and a default silence window of twenty minutes. Those are defaults in that draft, not universal operating requirements; teams should set monitoring behavior to fit their own task duration and alerting needs.
How to combine tracing with task-level status
OpenTelemetry spans are useful for understanding execution. AWS guidance recommends tracing spans across reasoning, model, tool, memory, retrieval, and handoff operations, and recommends custom metrics such as task success and failure rates when they are not captured implicitly. A task-level event or metric can answer a different question: did the requested work reach a verifiably complete outcome?
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Keep those signals correlated, but do not collapse them. A trace explains what the agent and its dependencies did; a task outcome records whether the requested result was achieved and checked. AWS’s CloudWatch generative AI observability guidance is one vendor-specific implementation reference, not an open standard.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Choosing an implementation path
The available material supports two broad approaches, without establishing feature parity or a performance comparison:
| Approach | What it can represent | Key design question | Trade-off |
|---|---|---|---|
| Built-in vendor instrumentation for a supported platform | Platform-supported execution traces and, where configured, task-level metrics | Does it capture the user-task outcome and consumer-visible delivery, or mainly model and infrastructure activity? | Can reduce setup effort within that platform, but portability and framework coupling need consideration. |
| Framework-specific spans plus custom task events or metrics | Execution details and a task lifecycle shaped to the application’s needs | Can the task identifier and outcome be carried across agent, tool, and asynchronous boundaries, and can delivery be verified? | Offers control over the lifecycle vocabulary, while requiring application-level instrumentation and maintenance. |
Before choosing, check whether the approach can correlate agent and tool activity, verify delivery from the consumer’s perspective, and preserve a task-level result independent of infrastructure status. A transport-agnostic draft vocabulary may help shape events, but adopting a draft does not make it a finalized standard.
A practical definition of “done”
For a long-running agent, treat done as a claim about the requested deliverable, not as a synonym for “the process stopped.” The strongest signal pairs a terminal task outcome with evidence that the expected consumer can see the result. Keep execution traces for diagnosis, and use an explicit lifecycle signal for the question those traces cannot settle on their own.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




