Best AI LLM Evaluation Tools in 2026

In short: Vellum is ranked #1 of 30 as of 8 October 2026, ahead of Weights & Biases and Evidently AI. The best-ranked option with a free plan is Weights & Biases. The lowest first paid tier on this page is Opik at $19/mo.

Choosing an LLM evaluation tool starts with the questions you need to answer about a model or prompt. Compared on evaluation methods, model support, and safety evaluations, options such as Vellum, Weights & Biases, and Evidently AI can be weighed against your assessment needs. Prompt versioning is another criterion when prompt changes matter, while API access and deployment help you compare how a tool may fit into a workflow. The entries also show free-plan availability and paid-from pricing, so you can consider access and cost alongside capabilities. Review the listed criteria in light of what you want to evaluate and how you expect to use the tool.

30 AI LLM evaluation tools ranked on what their makers publish — plans and prices, free tiers, platforms and the facts on their own pages.

30ranked
22free plans on this page
$19/molowest paid tier
8 Oct 2026last checked

22 of the 25 here have a free tier you can use, 0 have open-source code on record, and 0 have full bars: free, open, on most of your devices and documented.

#AppSignalFree tierOpen codeDevicesFromScore
1VellumWeb · Windows · Mac · Android · iPhone · Self-hosted · Extension Three bars—5 of 6$30/mo6.7
2Weights & BiasesWeb · Windows · Mac · Linux · iPhone · Self-hosted · API Three bars—5 of 6$60/mo6.7
3Evidently AIWeb · Windows · Mac · Linux · Self-hosted · API Three bars—4 of 6$80/mo6.6
4PromptfooWeb · Windows · Mac · Linux · Self-hosted · API Three bars—4 of 6Free6.6
5DeepEvalWindows · Mac · Linux · Self-hosted Three bars—3 of 6Free6.5
6GiskardWeb · Linux · Self-hosted · API Two bars—2 of 6Free6.5
7OpikWeb · Linux · Self-hosted · API Two bars—2 of 6$19/mo6.5
8Rhesis AIWeb · Linux · Self-hosted · API Two bars—2 of 6Free6.5
9BraintrustWeb · Self-hosted · API Two bars—1 of 6$249/mo6.4
10Confident AIWeb · Self-hosted · API Two bars—1 of 6$200/mo6.4
11GalileoWeb · Self-hosted · API Two bars—1 of 6$100/mo6.4
12HoneyHiveWeb · Self-hosted · API Two bars—1 of 6Free6.4
13LangfuseWeb · Self-hosted · API Two bars—1 of 6$29/mo6.4
14LangSmithWeb · Self-hosted · API Two bars—1 of 6$39/mo6.4
15LangWatchWeb · Linux · Self-hosted · API Two bars—2 of 6€29/mo6.4
16Maxim AIWeb · Self-hosted · API Two bars—1 of 6$29/mo6.4
17NVIDIA NeMo EvaluatorLinux · Self-hosted · API Two bars—1 of 6Free6.4
18Arize PhoenixWeb · Self-hosted · API Two bars—1 of 6Free6.3
19LM Evaluation HarnessLinux · Self-hosted · API Two bars—1 of 6Free6.3
20RAGCheckerSelf-hosted Two bars—0 of 6Free6.3
21TruLensSelf-hosted · API Two bars—0 of 6Free6.3
22Inspect AIAPI · Extension Two bars—0 of 6Free6.2
23OpenAI EvalsWeb · Self-hosted · API One bar——1 of 6—5.5
24OpenCompassWindows · Linux · Self-hosted · API One bar——2 of 6—5.5
25Pydantic EvalsLinux One bar——1 of 6—5.5
Compare all 25 in a table
#AppScoreFree planFromFree planPaid fromEvaluation methodsModel support
1Vellum6.7Free plan$30/moYes30 /mo——
2Weights & Biases6.7Free plan$60/moYes60 /mo——
3Evidently AI6.6Free plan$80/moYes———
4Promptfoo6.6Free planFreeYes———
5DeepEval6.5Free planFreeYes———
6Giskard6.5Free planFreeYes———
7Opik6.5Free plan$19/moYes19 /mo——
8Rhesis AI6.5Free planFreeYes—offline evaluation, online trace metrics, LLM-as-a-judge, custom metrics, single-turn testing, multi-turn testing, adversarial red-teamingOpenAI, Anthropic, Google Gemini, Azure OpenAI, Mistral, Cohere, Groq, Together AI, Perplexity, Replicate, Ollama, vLLM, LiteLLM Proxy
9Braintrust6.4Free plan$249/moYes249 /mo——
10Confident AI6.4Free plan$200/moYes200 /mo——
11Galileo6.4Free plan$100/moYes100 /mo——
12HoneyHive6.4Free planFreeYes———
13Langfuse6.4Free plan$29/moYes29 /mo——
14LangSmith6.4Free plan$39/moYes———
15LangWatch6.4Free plan€29/moYes———
16Maxim AI6.4Free plan$29/moYes———
17NVIDIA NeMo Evaluator6.4Free planFree——Built-in benchmarks, exact match, fuzzy match, multiple-choice regex, answer-line, numeric match, code sandbox, LLM-as-judge, JSON-schema validation, regression comparison, quality gatesOpenAI-compatible model endpoints, NVIDIA API Catalog, vLLM, NIM, local vLLM, SGLang, TensorRT-LLM, vision-language models
18Arize Phoenix6.3Free planFreeYes———
19LM Evaluation Harness6.3Free planFree————
20RAGChecker6.3Free planFree————
21TruLens6.3Free planFree————
22Inspect AI6.2Free planFree————
23OpenAI Evals5.5No———basic exact/match evaluations, model-graded evaluations, custom evaluation logic, academic benchmarks, meta-evaluationsOpenAI API models and custom CompletionFunction implementations
24OpenCompass5.5No———objective; subjective; discriminative; generative; LLM-as-a-judgeHugging Face models; API-based models; custom models; OpenAI; Anthropic; Gemini; Qwen; GLM; DeepSeek
25Pydantic Evals5.5No—Yes—Deterministic checks; custom evaluators; LLM judges; G-Eval; performance checks; report evaluators; span-based evaluation; agentic trajectory evaluationOpenAI; Anthropic; Gemini; xAI; Bedrock; Cerebras; Cohere; Groq; Hugging Face; Mistral; OpenRouter; and other listed Pydantic AI providers

Is your app on this list?

Numbered spots on this list can be sponsored, and a sponsored row is labelled as paid.

Questions about this list

Which AI LLM evaluation tool is ranked first on Freedom251?

Vellum is ranked #1 of 30 with a score of 6.7. Weights & Biases is second and Evidently AI third.

How many of these have a free plan?

22 of the 25 on this page publish a free plan on their own pricing pages.

Which is the cheapest paid option?

On this page, Opik has the lowest first paid tier we found: $19/mo.

How is this list ranked?

Ranked free-and-open first: a usable free tier and open-source code, then the platforms it runs on and its documentation.

More in AI Tools

All AI tools lists