Build a reproducible AI agent evaluation lab by treating each run as a controlled experiment: define a reviewable task, fix its inputs and environment, isolate execution, record both the answer and tool activity, and compare repeated runs with a saved baseline. Docker Compose can make the lab’s services, mounts, networks, and configuration explicit, but Compose alone does not make results reproducible. You must also record the agent, model, prompt, task, dependencies, image versions, and relevant resource settings.
What the lab needs to control
An evaluation is only comparable when the new run and the baseline face equivalent conditions. Write down what the agent receives, what it may access, how the task is prepared, what counts as success, and which outputs are retained. Keep those decisions separate from the score itself: a passing result without enough run detail is difficult to investigate or reproduce.
As an Amazon Associate I earn from qualifying purchases.
- Task: the user input, expected behavior, and any task-specific success criteria.
- Environment: setup steps, working directory, mounted files, available tools, resource limits, and network access.
- Agent configuration: model and provider, agent version, prompt or system instructions, and tool configuration.
- Evidence: final response, tool calls, logs, score components, and task outputs.
- Comparison: repeat count, baseline identity, and any tolerance used to handle noisy scores.
Keep this information with each run. Docker Agent’s evaluation documentation describes session definitions with a user question, expected tool calls, optional response criteria, setup, and working-directory fixtures. That is a useful example of making a test inspectable; it is not a complete, pinned Docker Compose recipe.
Organize the project around tasks and runs
Store task definitions and their fixtures in version control so a reviewer can see what changed between evaluations. Keep run outputs outside the task definitions: otherwise generated files can accidentally become inputs to the next run.
#1 Best Overall
agent-eval-lab/
compose.yaml
tasks/
task-001/
case.yaml
workspace/
setup.sh
runner/
results/
This is a suggested project layout, not a prescribed format. Each task definition should identify its input, expected tool behavior or response properties, setup requirements, working directory, and scoring criteria. Record a task revision or commit identifier in the run metadata so that a later comparison can establish which fixture was evaluated.
Map the lab to Docker Compose
Use Compose to make the runner and its boundaries visible and reviewable. The sources for this topic do not establish a verified, pinned Compose manifest or a universal service layout, so treat the following as design decisions to validate against the agent and workload you choose—not as a ready-to-run reference file.
Rank #2
Runner
Define the runner as the service that launches one task, collects results, and exits. Build or select its image deliberately; pin the image to an immutable version or digest when the runner supports it, and record that identifier. Lock language packages and other dependencies using the package manager’s lockfile. These are implementation practices for reducing drift, not conventions specified by Docker Agent’s evaluation guide.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsTask workspace and results
Mount the selected task fixture at a known path and make the working directory explicit. Write reports and logs to a separate results location that survives the runner container’s exit. If task execution should not modify its source fixture, mount that fixture read-only and give the task a distinct writable workspace. Workspace-Bench describes this kind of task-local setup, including separate HOME, temporary, and cache paths for its protocol.
Rank #3
Network, credentials, and isolation
Decide which services and external endpoints the agent needs, then limit access to those requirements. Keep provider credentials out of task files and version control; pass them through the runtime’s secret or environment mechanism, and avoid writing their values to logs or result artifacts. A container boundary helps define the experiment, but it does not by itself make untrusted agent code safe: permissions, mounts, network access, and host-side processes remain part of the security design.
Credential handling varies by runner. Docker Agent’s own evaluation documentation says provider API keys are forwarded automatically, while GITHUB_TOKEN and GH_TOKEN are not forwarded automatically; its GitHub Copilot setup requires explicit handling. Its LLM judge runs on the host, rather than inside the evaluation container. Do not assume these behaviors apply to a Compose implementation using another agent or judge.
Define how a task is scored
Score the behavior the task is intended to test, rather than relying on a single impression of the final answer. For a tool-using task, preserve action-level evidence as well as response quality. Keep deterministic checks distinct from judgments made by a language model.
Recommended Free Tools
| Score component | What it can reveal | How to interpret it |
|---|---|---|
| Tool-call accuracy | Whether the agent selected and used expected tools or actions | Docker Agent documents a tool-call F1 metric. The appropriate expected calls depend on the task. |
| Response relevance | Whether the final response satisfies the task’s relevance criteria | Docker Agent documents an LLM judge for relevance statements; keep this judgment distinguishable from deterministic checks. |
| Output size | Whether the response is within a useful size category | Docker Agent includes output size as an evaluation dimension. Define what size is suitable for your task rather than treating one threshold as universal. |
| Cost | How much a run costs, when the runner reports it | Docker Agent documents cost reporting, but cost is not used by its regression gate. |
Use only metrics that answer a real evaluation question, and retain the underlying response and tool trace so a surprising score can be reviewed. A composite score can hide a failure mode—for example, a relevant answer produced after the wrong tool action—so preserve component scores even when you also calculate an aggregate.
Best Value
- Docker, Docker Swarm, Docker Compose, Programmer, Developer, Coding, Programming, Software Engineer, Code, DevOps, Deploy, Deployment, Kubernetes, Salt, Puppet, Chef, Terraform, Container, AWS, Azure, Cloud, Geek, Funny, Computer, Software, Tech, IT
- Integration, Scrum, Compile, Compilation, Science, Bug, Debug, Python, Linux, Java, Javascript, Scala, Dotnet, Kotlin
- Lightweight, Classic fit, Double-needle sleeve and bottom hem
Run, repeat, and compare evaluations
- Validate the Compose configuration. Run
docker compose configbefore execution to catch configuration problems and inspect the resolved settings. - Start from a clean task state. Apply the task’s setup steps and fixture, and make sure prior run output is not visible as input unless the test explicitly requires it.
- Execute the selected task suite. Record the Compose project configuration and the agent, model, prompt, task, dependency, and image identifiers with the run.
- Repeat selected tasks. Repeated runs help reveal variation, especially when model behavior or an LLM judge is nondeterministic. Record the repeat count and retain individual results instead of only an average.
- Compare with a saved baseline. Compare like with like: use the same task revision, fixtures, resource profile, tools, and credential permissions. Investigate changes in tool behavior and individual score components, not only the aggregate.
- Keep the artifacts. Save the report, logs, session data, and task outputs needed to explain the result. Docker Agent’s documented result directory includes JSON, logs, and a database.
Docker Agent supports repeat counts and comparison against a saved prior run. Its documentation warns that an LLM judge can vary, so a regression tolerance may help avoid noisy aggregate gates. A transition from pass to fail still gates according to its documented behavior. Set any tolerance intentionally and retain the per-run results; otherwise a threshold can obscure a genuine change.
Choose isolation and resources for the workload
Fresh containers and consistent resource limits make benchmark conditions easier to compare. Workspace-Bench describes a protocol that creates and removes a fresh container for each task, uses task-local paths, and fixes the resource profile across tasks. Its documented defaults are 2 CPUs, 8 GiB of memory, 512 PIDs, and 20 GiB of writable task storage. Those figures describe Workspace-Bench’s protocol, not universal requirements or recommended settings for every lab.
Choose limits based on the work being evaluated, then keep them the same across compared configurations. Also record whether each run has equivalent access to tools, credentials, and network services. A model comparison is misleading if one configuration has more memory, a different workspace, or broader tool access than another.
Read scores within their benchmark scope
A benchmark result applies to its named tasks, agent setup, and scoring method; it is not a general ranking of agent quality. OpenAI reports a 21.0% average replication score for Claude 3.5 Sonnet (New) with open-source scaffolding as the best-performing tested agent in its PaperBench evaluation. That figure belongs to that model-and-scaffolding configuration and PaperBench’s replication tasks; it should not be used to predict performance on an unrelated task suite.
When publishing or sharing your own lab’s results, name the task suite and revision, agent and model configuration, runtime conditions, scoring method, and repeat count. Separate deterministic checks from judge-based scores, and report variation when repeated runs differ. That makes the score interpretable without implying that it measures every aspect of agent quality.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




