AgentClash

Web · Self-hosted · API · paid plans from $49/mo

Freedom report

Two barsScore 6.4

  • Free tierA free tier is on its own pricing page
  • Open codeNo open-source code on record
  • Runs widely1 of 6 device platforms
  • DocumentedPlans, terms and facts published

AgentClash is an open-source platform for evaluating AI agents on multi-turn tasks in a sandbox. It scores tool choices, cost, latency, recovery, and final results, then can preserve a failed run as a regression test for later evaluations. Each agent runs in a fresh Firecracker microVM with an isolated filesystem and network; the sandbox is removed after the run. YAML challenge packs can define tools, policy, scoring, and starting state. Available tools include file operations, data queries, HTTP, shell, and test runners. Scoring combines deterministic, mathematical, behavioral, and language-model judges with configurable weights and consensus. Provider adapters support OpenAI, Anthropic, Gemini, xAI, Mistral, and OpenRouter. CI/CD checks can run from GitHub Actions, a webhook, or the CLI, and fail builds when correctness, cost, latency, or required evidence regresses. AgentClash is MIT licensed and can be self-hosted as a full stack or used through its hosted backend. The Free plan includes 25 evaluation runs per month. Pro costs 49.00 USD per month billed monthly, or $39 / month ($468 / yr) with annual billing.

Who it is for

AgentClash suits teams evaluating multi-step AI agents for coding, research, SRE, operations, codebase questions, or support. It is also relevant to teams that want evaluation failures turned into regression checks in CI/CD.

What is good

  • Scores tool choices, cost, latency, recovery, and results.
  • Failed runs can become regression tests.
  • Runs agents in isolated, disposable microVMs.
  • Supports several model providers and OpenRouter.
  • Can be self-hosted or used through a hosted backend.

What to know first

  • Free plan allows 25 evaluation runs per month.
  • Pro costs 49.00 USD per month billed monthly.
  • Free plan requires a BYO LLM API key and sandbox token.

Verdict

AgentClash combines sandboxed agent evaluation with replayable regression tests and CI/CD checks. The Free plan has a monthly run limit; Pro costs 49.00 USD per month billed monthly or $39 / month ($468 / yr) with annual billing.

AgentClash plans and pricing

All plans
Free Free 1 workspace · 25 eval runs / month · up to 4 models per run · 7-day replay retention · BYO LLM API key · BYO E2B sandbox token · community support agentclash.dev · 1 Oct 2026
Pro $49/mo Billed monthly; annual billing $39 / month ($468 / yr) 500 eval runs / workspace / month · up to 8 models per run · 30-day replay retention · hosted sandbox with included credit · private challenge packs · CI integration · 3 concurrent eval runs · email support < 1 business day agentclash.dev · 1 Oct 2026
Team $100/mo Billed monthly; annual billing $80 / month ($960 / yr) 2,000 eval runs / workspace / month · up to 12 models per run · 90-day replay retention · 10 concurrent eval runs · multiple workspaces · workspace-level audit log · Slack notifications · priority email support < 4 business hours agentclash.dev · 1 Oct 2026
Enterprise Not published Custom SSO / SAML · org-wide audit logs · unlimited replay retention · 99.9% uptime SLA · dedicated support channel · custom MSA / billing terms agentclash.dev · 1 Oct 2026

Compared on AI agent evaluation tools

Free plan
Yes
Paid from
$39/mo
Evaluation methods
hybrid
Tool-call checks
Yes
Trace ingestion
Yes
Safety evaluations
Yes
Regression runs
Yes

Best AgentClash alternatives

See all 20