What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Test an LLM application by defining observable success criteria, running representative inputs through the complete system, grading the results with checks suited to the task, and investigating failures. Keep a versioned evaluation set and rerun it when prompts, models, tools, retrieval, or application code change. A single score is not proof that an application is reliable: it describes only the system, data, graders, and conditions used to produce it.

What an LLM application test should measure

An evaluation is a set of inputs plus grading logic that measures whether a system did what you intended. Start with the behavior that matters to the user, not a generic quality score. For a support assistant, success might mean answering the question accurately from approved material, citing the relevant passage, and refusing to invent a policy. For a classifier, success may be returning one of a fixed set of labels. For an agent, it may mean completing a task and leaving an external system in the correct state.

Write criteria so that a reviewer or program can tell what passed. “Be helpful” is too broad by itself; “identify the requested refund window, cite the policy source, and do not promise an exception” is more testable. OpenAI’s guidance breaks evaluation into describing the task, running test inputs, and analyzing results for iteration: Working with evals.

  • Input: the user request and any relevant conversation history or context.
  • System under test: the model together with prompts, retrieval, tools, application logic, and safeguards.
  • Expected evidence: reference answers, labels, required behavior, forbidden behavior, or the desired final state.
  • Grader: exact code, a human rubric, a model-based judgment, or a combination.

Build a test set that resembles actual use

A useful evaluation set represents the real task distribution, not just easy prompts that make the system look good. Begin with typical requests and add edge cases, adversarial inputs, and examples drawn from user feedback or production failures when appropriate. Have people with relevant subject-matter expertise create or review reference answers and labels. OpenAI recommends diverse examples and expert input; see its evaluation best practices.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Keep each case explicit and maintainable

Store cases in a version-controlled format such as JSONL. Include the input, any required context, the expected result or grading criteria, and a stable case ID. For RAG, preserve the expected source or relevant passages as well as the user question. For tool-using systems, record what actions are permitted and what outcome should result.

{"id":"refund-window-01","question":"How long do I have to request a refund?","expected_contains":"30 days","required_source":"refund-policy","tags":["typical","policy"]}
{"id":"refund-unknown-02","question":"Can I get a refund after the stated window?","expected_contains":"contact support","required_source":"refund-policy","tags":["edge"]}

The values above are illustrative; replace them with your own policy, expected response, and application-specific source identifiers. Keep dataset versions alongside the model, prompt, and code version used for each run. When a real failure occurs, turn it into a case where doing so is safe and useful. This makes a previously observed regression visible next time.

Include difficult and misuse cases

Add malformed inputs, ambiguous requests, missing context, conflicting instructions, attempts to override system rules, requests for private data, and requests that should be refused or redirected. Choose scenarios relevant to your product rather than treating a generic adversarial list as complete. Google’s safety evaluation guidance and OpenAI’s red-teaming guide describe ways to examine system-specific safety risks.

Choose graders that match the requirement

There is no single grader that is best for every quality. Prefer the least subjective method that captures the requirement, then add human review where judgment is necessary. A combined evaluation is often more informative than asking one model to assign a vague overall score.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Requirement Suitable evidence or grader Watch for
Exact format, required fields, allowed labels, or valid JSON Parser, schema validation, or exact match A structurally valid response can still be factually wrong.
Required content or prohibited phrase Deterministic checks, carefully scoped to the task Substring checks can miss paraphrases or accept misleading context.
Factual correctness, tone, or answer quality Human rubric; model grader calibrated against human-labeled examples Model graders can be inconsistent or favor longer answers or a particular answer position.
Retrieval quality and grounding Check retrieved passages against relevant references; separately judge answer support Good retrieval does not guarantee a correct answer, and a plausible answer may lack support.
Agent task completion Tool-call and trace checks plus final environment state A convincing final message does not prove the intended action happened.

For subjective qualities, write a rubric with observable criteria and examples of passing and failing behavior. If using a model grader to handle volume, compare its judgments with human labels on a representative sample and inspect disagreements. Pairwise comparisons can be useful, but randomize which answer appears first and account for position and verbosity bias. These cautions are covered in OpenAI’s guidance on human evaluation and model graders.

Test the whole application and the parts that can fail

Run cases through the same relevant application path users rely on, but measure components separately when doing so helps explain failures. A single end-to-end pass tells you that something went wrong; component evidence can help locate why.

For RAG: separate retrieval from generation

Record the retrieved documents or chunks for each case. Check whether the relevant material appeared in the context, then grade whether the generated answer is correct and supported by that context. If retrieval missed the needed source, changing the answer prompt may not solve the root problem. If the needed source was present but ignored or misrepresented, investigate generation, context formatting, and instructions. OpenAI’s evaluation recommendations include assessing retrieval and answer quality as distinct parts of a RAG workflow.

For agents: grade actions, traces, and outcomes

An agent is not just a final text response. Evaluate the task, the tools available, the harness that runs the model, and the environment where actions occur. Capture tool calls and intermediate traces when possible, then verify the final state: for example, whether the right record was updated or the requested file was created. Run repeated trials for tasks where model or environment variability matters. Anthropic explains task definitions, trials, graders, transcripts, outcomes, and evaluation harnesses in Demystifying evals for AI agents.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For browser-facing features: test rendered output when it matters

If users interact with an LLM feature in a web interface, text-level tests will not catch every UI regression. Add browser checks for the states that matter, such as an answer appearing in the correct panel or a citation link being usable. Screenshot comparison can help inspect visual changes, but it should complement—not replace—checks of the underlying content and behavior.

Automate a small regression evaluation

Start with a compact suite that can run after meaningful changes. Expand it as you learn which cases distinguish good behavior from failure. OpenAI recommends continuous evaluation on changes and monitoring for nondeterminism in its evaluation best practices.

  1. Freeze a baseline. Run the current dataset and save case-level outputs, grader results, and system configuration.
  2. Run on relevant changes. Include prompt, model, retrieval, tool, safeguard, and application-code changes that could affect behavior.
  3. Compare, do not just average. Review pass rates by task or risk category and inspect individual regressions, new failures, and changed outputs.
  4. Investigate before accepting. Determine whether a changed result is a genuine regression, an acceptable trade-off, or noise from a nondeterministic output.
  5. Promote new cases. Add reproducible failures to the versioned set, then rerun after a fix.

Example: a minimal Python HTTP evaluation runner

This standard-library script reads the JSONL cases shown above, sends each question as JSON to an application endpoint, expects a JSON response with an answer field, and checks the required phrase. Set APP_URL to an endpoint that accepts {"question":"..."}; adapt the request and response parsing if your application contract differs. Save it as eval.py and run APP_URL=https://your-app.example/evaluate python eval.py cases.jsonl.

import json
import os
import sys
import urllib.error
import urllib.request

app_url = os.environ["APP_URL"]
cases_path = sys.argv[1] if len(sys.argv) > 1 else "cases.jsonl"
results = []

with open(cases_path, encoding="utf-8") as cases_file:
    for line_number, line in enumerate(cases_file, 1):
        case = json.loads(line)
        payload = json.dumps({"question": case["question"]}).encode("utf-8")
        request = urllib.request.Request(
            app_url, data=payload,
            headers={"Content-Type": "application/json"}, method="POST"
        )
        try:
            with urllib.request.urlopen(request, timeout=60) as response:
                result = json.loads(response.read().decode("utf-8"))
            answer = result["answer"]
            passed = case["expected_contains"].casefold() in answer.casefold()
            results.append((case["id"], passed, answer))
            print(f"{'PASS' if passed else 'FAIL'} {case['id']}")
        except (urllib.error.URLError, TimeoutError, KeyError, ValueError) as error:
            results.append((case.get("id", f"line-{line_number}"), False, str(error)))
            print(f"ERROR line {line_number}: {error}")

passed_count = sum(passed for _, passed, _ in results)
print(f"Passed {passed_count}/{len(results)} cases")
sys.exit(0 if passed_count == len(results) else 1)

This is a deliberately narrow grader: it checks only whether a phrase occurs, not whether the answer is correct, grounded, safe, or appropriately qualified. Add checks that express your actual criteria, and preserve the per-case output so you can inspect a failure instead of relying on the aggregate count.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Make results interpretable and control evaluation cost

For each run, record the exact model identifier, prompt version, tools and tool settings, application or harness version, safeguards, dataset version, grader, and relevant run settings. For agent and RAG cases, retain the retrieved context, trace, and final outcome when possible. OpenAI’s playbook for trustworthy third-party evaluations emphasizes treating results as conditional evidence and checking validity threats such as shortcuts, contamination, refusals, and evaluation awareness.

Track cost in your own setup rather than assuming a universal price or performance ranking. Model-graded cases consume judge calls; repeated trials multiply runs; longer traces and human review add time. Keep the high-value regression set quick enough for frequent use, and run broader or repeated evaluations when the risk justifies the added expense. A score should be reported with the conditions that produced it, not generalized to different prompts, models, datasets, or users.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot common evaluation failures

  • The aggregate score looks healthy, but users still find serious errors. Break results down by task and risk category; add missing edge cases and review failures manually. Easy cases can conceal a weak area.
  • Results move between runs without a code change. Treat variability as evidence to characterize: repeat relevant cases, record settings and outputs, and report the range or frequency of failures rather than selecting the best run.
  • A model grader disagrees with reviewers. Tighten the rubric, provide concrete criteria, examine disputed examples, and validate the grader against human labels before relying on it at scale.
  • RAG answers are wrong despite a plausible response. Inspect retrieved passages first. Determine whether the evidence was absent, irrelevant, or present but not used; grade retrieval and answer grounding separately.
  • An agent claims success but the task did not happen. Check tool-call traces and verify the environment state instead of grading the final message alone.
  • Tests pass locally but fail in CI. Compare model, prompt, dependencies, credentials, data version, endpoint configuration, and timeouts. Log enough configuration to reproduce the run without exposing secrets.
  • Prompt changes improve one case and break another. Run the complete regression set and inspect case-level differences. Do not promote a change based on a handpicked example.

Or skip the browser setup

For browser-facing LLM features, ScreenshotNeo can capture a rendered page through one GET request when you need a screenshot or PDF rather than a custom browser harness. Its website screenshot API accepts a URL and returns PNG, JPEG, WebP, or PDF. For example, this cURL request saves a WebP capture; see the ScreenshotNeo API documentation for parameters and response details.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each of those steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for AI agents and MCP clients. The Free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Sign up for ScreenshotNeo to get 1,000 free screenshots a month with no card.

Further reading

For a broader treatment of evaluation alongside prompt engineering, RAG, and agents, see O’Reilly’s listing for AI Engineering by Chip Huyen. It can provide context, but your own versioned cases remain the evidence for your application.

Frequently Asked Questions

How is an LLM evaluation different from a benchmark?

A benchmark is a standardized comparison setup; an application evaluation is tailored to the task, data, tools, and failure costs of a particular system. A benchmark result alone does not establish that your deployed application behaves well.

Should every test have a single expected answer?

No. Exact references suit deterministic requirements, while open-ended tasks may need a rubric, multiple acceptable outcomes, or human review. Choose a grading method that reflects what counts as success for that case.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.