Best AI LLM Evaluation Tools in 2026
Updated
In short: Vellum is ranked #1 of 30 as of 8 October 2026, ahead of Weights & Biases and Evidently AI. The best-ranked option with a free plan is Weights & Biases. The lowest first paid tier on this page is Opik at $19/mo.
Choosing an LLM evaluation tool starts with the questions you need to answer about a model or prompt. Compared on evaluation methods, model support, and safety evaluations, options such as Vellum, Weights & Biases, and Evidently AI can be weighed against your assessment needs. Prompt versioning is another criterion when prompt changes matter, while API access and deployment help you compare how a tool may fit into a workflow. The entries also show free-plan availability and paid-from pricing, so you can consider access and cost alongside capabilities. Review the listed criteria in light of what you want to evaluate and how you expect to use the tool.
30 AI LLM evaluation tools ranked on what their makers publish — plans and prices, free tiers, platforms and the facts on their own pages.
22 of the 25 here have a free tier you can use, 0 have open-source code on record, and 0 have full bars: free, open, on most of your devices and documented.
| # | App | Signal | Free tier | Open code | Devices | From | Score |
|---|---|---|---|---|---|---|---|
| 1 | VellumWeb · Windows · Mac · Android · iPhone · Self-hosted · Extension | Three bars | — | 5 of 6 | $30/mo | 6.7 | |
| 2 | Weights & BiasesWeb · Windows · Mac · Linux · iPhone · Self-hosted · API | Three bars | — | 5 of 6 | $60/mo | 6.7 | |
| 3 | Evidently AIWeb · Windows · Mac · Linux · Self-hosted · API | Three bars | — | 4 of 6 | $80/mo | 6.6 | |
| 4 | PromptfooWeb · Windows · Mac · Linux · Self-hosted · API | Three bars | — | 4 of 6 | Free | 6.6 | |
| 5 | DeepEvalWindows · Mac · Linux · Self-hosted | Three bars | — | 3 of 6 | Free | 6.5 | |
| 6 | GiskardWeb · Linux · Self-hosted · API | Two bars | — | 2 of 6 | Free | 6.5 | |
| 7 | OpikWeb · Linux · Self-hosted · API | Two bars | — | 2 of 6 | $19/mo | 6.5 | |
| 8 | Rhesis AIWeb · Linux · Self-hosted · API | Two bars | — | 2 of 6 | Free | 6.5 | |
| 9 | BraintrustWeb · Self-hosted · API | Two bars | — | 1 of 6 | $249/mo | 6.4 | |
| 10 | Confident AIWeb · Self-hosted · API | Two bars | — | 1 of 6 | $200/mo | 6.4 | |
| 11 | GalileoWeb · Self-hosted · API | Two bars | — | 1 of 6 | $100/mo | 6.4 | |
| 12 | HoneyHiveWeb · Self-hosted · API | Two bars | — | 1 of 6 | Free | 6.4 | |
| 13 | LangfuseWeb · Self-hosted · API | Two bars | — | 1 of 6 | $29/mo | 6.4 | |
| 14 | LangSmithWeb · Self-hosted · API | Two bars | — | 1 of 6 | $39/mo | 6.4 | |
| 15 | LangWatchWeb · Linux · Self-hosted · API | Two bars | — | 2 of 6 | €29/mo | 6.4 | |
| 16 | Maxim AIWeb · Self-hosted · API | Two bars | — | 1 of 6 | $29/mo | 6.4 | |
| 17 | NVIDIA NeMo EvaluatorLinux · Self-hosted · API | Two bars | — | 1 of 6 | Free | 6.4 | |
| 18 | Arize PhoenixWeb · Self-hosted · API | Two bars | — | 1 of 6 | Free | 6.3 | |
| 19 | LM Evaluation HarnessLinux · Self-hosted · API | Two bars | — | 1 of 6 | Free | 6.3 | |
| 20 | RAGCheckerSelf-hosted | Two bars | — | 0 of 6 | Free | 6.3 | |
| 21 | TruLensSelf-hosted · API | Two bars | — | 0 of 6 | Free | 6.3 | |
| 22 | Inspect AIAPI · Extension | Two bars | — | 0 of 6 | Free | 6.2 | |
| 23 | OpenAI EvalsWeb · Self-hosted · API | One bar | — | — | 1 of 6 | — | 5.5 |
| 24 | OpenCompassWindows · Linux · Self-hosted · API | One bar | — | — | 2 of 6 | — | 5.5 |
| 25 | Pydantic EvalsLinux | One bar | — | — | 1 of 6 | — | 5.5 |
Compare all 25 in a table
| # | App | Score | Free plan | From | Free plan | Paid from | Evaluation methods | Model support |
|---|---|---|---|---|---|---|---|---|
| 1 | Vellum | 6.7 | Free plan | $30/mo | Yes | 30 /mo | — | — |
| 2 | Weights & Biases | 6.7 | Free plan | $60/mo | Yes | 60 /mo | — | — |
| 3 | Evidently AI | 6.6 | Free plan | $80/mo | Yes | — | — | — |
| 4 | Promptfoo | 6.6 | Free plan | Free | Yes | — | — | — |
| 5 | DeepEval | 6.5 | Free plan | Free | Yes | — | — | — |
| 6 | Giskard | 6.5 | Free plan | Free | Yes | — | — | — |
| 7 | Opik | 6.5 | Free plan | $19/mo | Yes | 19 /mo | — | — |
| 8 | Rhesis AI | 6.5 | Free plan | Free | Yes | — | offline evaluation, online trace metrics, LLM-as-a-judge, custom metrics, single-turn testing, multi-turn testing, adversarial red-teaming | OpenAI, Anthropic, Google Gemini, Azure OpenAI, Mistral, Cohere, Groq, Together AI, Perplexity, Replicate, Ollama, vLLM, LiteLLM Proxy |
| 9 | Braintrust | 6.4 | Free plan | $249/mo | Yes | 249 /mo | — | — |
| 10 | Confident AI | 6.4 | Free plan | $200/mo | Yes | 200 /mo | — | — |
| 11 | Galileo | 6.4 | Free plan | $100/mo | Yes | 100 /mo | — | — |
| 12 | HoneyHive | 6.4 | Free plan | Free | Yes | — | — | — |
| 13 | Langfuse | 6.4 | Free plan | $29/mo | Yes | 29 /mo | — | — |
| 14 | LangSmith | 6.4 | Free plan | $39/mo | Yes | — | — | — |
| 15 | LangWatch | 6.4 | Free plan | €29/mo | Yes | — | — | — |
| 16 | Maxim AI | 6.4 | Free plan | $29/mo | Yes | — | — | — |
| 17 | NVIDIA NeMo Evaluator | 6.4 | Free plan | Free | — | — | Built-in benchmarks, exact match, fuzzy match, multiple-choice regex, answer-line, numeric match, code sandbox, LLM-as-judge, JSON-schema validation, regression comparison, quality gates | OpenAI-compatible model endpoints, NVIDIA API Catalog, vLLM, NIM, local vLLM, SGLang, TensorRT-LLM, vision-language models |
| 18 | Arize Phoenix | 6.3 | Free plan | Free | Yes | — | — | — |
| 19 | LM Evaluation Harness | 6.3 | Free plan | Free | — | — | — | — |
| 20 | RAGChecker | 6.3 | Free plan | Free | — | — | — | — |
| 21 | TruLens | 6.3 | Free plan | Free | — | — | — | — |
| 22 | Inspect AI | 6.2 | Free plan | Free | — | — | — | — |
| 23 | OpenAI Evals | 5.5 | No | — | — | — | basic exact/match evaluations, model-graded evaluations, custom evaluation logic, academic benchmarks, meta-evaluations | OpenAI API models and custom CompletionFunction implementations |
| 24 | OpenCompass | 5.5 | No | — | — | — | objective; subjective; discriminative; generative; LLM-as-a-judge | Hugging Face models; API-based models; custom models; OpenAI; Anthropic; Gemini; Qwen; GLM; DeepSeek |
| 25 | Pydantic Evals | 5.5 | No | — | Yes | — | Deterministic checks; custom evaluators; LLM judges; G-Eval; performance checks; report evaluators; span-based evaluation; agentic trajectory evaluation | OpenAI; Anthropic; Gemini; xAI; Bedrock; Cerebras; Cohere; Groq; Hugging Face; Mistral; OpenRouter; and other listed Pydantic AI providers |
Is your app on this list?
Numbered spots on this list can be sponsored, and a sponsored row is labelled as paid.
Questions about this list
Which AI LLM evaluation tool is ranked first on Freedom251?
Vellum is ranked #1 of 30 with a score of 6.7. Weights & Biases is second and Evidently AI third.
How many of these have a free plan?
22 of the 25 on this page publish a free plan on their own pricing pages.
Which is the cheapest paid option?
On this page, Opik has the lowest first paid tier we found: $19/mo.
How is this list ranked?
Ranked free-and-open first: a usable free tier and open-source code, then the platforms it runs on and its documentation.














