Free tools Windows power users keep installed
One-click scans. No signup required.
Compare small language models on the same representative, held-out cases using the instructions, schema, and output mode your application will actually use. Score whether each decision is correct separately from whether its output parses or meets the schema; for tool tasks, also measure tool choice, argument accuracy, and successful execution. There is no universal best model for an unspecified workload.
Define what a correct decision means
Before comparing models, specify the decision the system must make and how you will judge it. For classification, list the permitted labels; for extraction, define the target fields and acceptable values; for routing, name the available destinations and when each applies. For tool-oriented tasks, distinguish among calling a tool, declining to call one, asking for missing information, and choosing a different tool.
- Define the input the model receives and the information it may use.
- Specify the permitted actions or labels, required fields, and any abstention or clarification behavior.
- Write down a checkable success rule for each case, including what counts as a consequential error.
OpenAI’s evaluation guidance recommends checking instruction following, functional correctness, tool selection, data precision, and agent handoff where relevant. A precise success rule lets you assess those behaviors instead of relying on a vague overall impression.
Build a representative, held-out test set
Use examples that reflect the workload the model will face, not just clean demonstrations. Include ordinary inputs, ambiguous or incomplete requests, and consequential edge cases. Keep a held-out set for the final comparison so you do not judge a prompt or schema only on examples already used to tune it.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstall#1 Best Overall
Run the same cases against every candidate. The evaluation guidance supports application-specific tests but does not establish one universally adequate sample size. Choose a set with enough variety for the intended workload, and report its size and limitations rather than presenting its results as proof of performance on every possible input.
Hold the comparison conditions constant
For a fair comparison, keep the task instructions, schema, available tools, decoding settings, and retry policy consistent. If the application will use a provider’s constrained-output feature, test that feature as part of the system being compared. If you are considering prompt-only JSON or another decoder, evaluate those as separate configurations rather than attributing every difference to model weights.
Output modes matter. Function calling connects a model to tools or APIs; a structured response format shapes the model’s answer. OpenAI distinguishes JSON mode, which ensures valid JSON, from Structured Outputs, which are designed to ensure adherence to supported schemas and models. Confirm compatibility for the models and schemas in your application, and test the exact production path: a constraint or interface can affect the decision as well as its formatting.
Score correctness separately from formatting
A parseable object is not necessarily a correct decision, and valid JSON is not the same as adherence to a schema. Track the layers below independently so a formatting improvement cannot conceal a decline in task performance.
Rank #3
| Measure | What to check |
|---|---|
| Decision accuracy | Did the model choose the right label, route, extracted value, or action? |
| Parse and schema validity | Does the output parse, and does it satisfy the specified schema? Record these as distinct checks where applicable. |
| Semantic validity | Are the field values correct and mutually consistent, even if the object passes schema validation? |
| Tool behavior | Was the right tool selected, were its arguments precise, and did the model hand off, decline, or request information appropriately? When safe, execute calls in a test environment and check whether the intended task completed. |
| Robustness | Does performance hold across varied cases and repeated runs? Report instability that changes a decision, not just harmless wording variation. |
| Operational fit | Measure latency and cost under representative deployment conditions if they affect the decision. Set thresholds for the application; there is no universal acceptable delay or cost in the cited guidance. |
For a useful dashboard, report schema validity, answer accuracy, executable accuracy where applicable, and the wrong-valid-schema rate: the share of outputs that satisfy the schema while encoding an incorrect answer. Keep denominators and scoring rules clear so each rate can be interpreted.
Evaluate tool calls as actions, not just objects
For a model that selects tools or APIs, checking the returned JSON is only an intermediate test. Score whether the chosen tool matches the request, whether its arguments identify the right data or operation, and whether the call actually achieves the intended outcome. Also include cases where the correct behavior is to avoid a call or ask for missing details.
Jaideep Ray’s 2026 Constraint Tax paper illustrates why: on its deterministic calendar tool-call task using Qwen2.5-1.5B, prompt-only JSON achieved 91.5% executable accuracy, compared with 48.0% for the tested hard tool-call schema; both modes had 100.0% schema validity. This result concerns that model, task, and setup, not a general rule that prompt-only JSON is better. It shows why the production mode must be evaluated on task outcomes as well as validity.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Use public benchmarks as supporting evidence
Benchmarks can reveal strengths and limitations in specific capabilities, but they do not replace an application-specific test set.
Recommended Free Tools
Best Value
- Constrained output: The 2025 JSONSchemaBench paper evaluates constrained decoding for efficiency in generating compliant outputs, constraint coverage, and output quality. It describes 10,000 real-world JSON schemas paired with the official JSON Schema Test Suite. That evidence can help assess schema and decoder behavior; it does not establish whether a model makes the decision your application needs.
- Tool use and agents: Stanford HAI’s 2026 AI Index describes BFCL V4 as adding broader agentic and multiturn coverage. In its overall score, agentic tasks account for 40% and multiturn interactions for 30%, with the remainder split across live, nonlive, and hallucination categories. The report says the top 15 models span about 21 percentage points in overall accuracy as of early 2026. These are figures for that benchmark version and leaderboard, not a forecast for small models on your workload.
Check benchmark version, task mix, and scoring before comparing published scores. Results from different evaluation setups are not directly interchangeable.
Interpret constrained-output results cautiously
The 2026 Constraint Tax paper by Jaideep Ray reports 15,000 generations across Qwen2.5-0.5B, Qwen2.5-1.5B, and SmolLM2-1.7B in commodity-GPU experiments. In its tested hard answer-only schema decoding setup and model conditions, schema validity ranged from 61.5% to 100.0%, answer accuracy from 19.7% to 11.0%, and wrong-valid-schema outputs from 49.5% to 88.9%. These are paper-specific experimental results, not expected rates for other tasks or deployments.
The practical lesson is to report validity and correctness side by side. A constrained output can improve compliance without improving the underlying answer, so a parse-success-only metric is not an adequate measure of a structured decision system.
Select for the workload and deployment
Choose the candidate that meets your correctness and reliability requirements under the real output path and operating constraints. Compare the full cost of errors as well as model cost: a faster or cheaper model may require more human review or cause more failed actions. Likewise, a higher aggregate benchmark score does not by itself make a model the best fit for a narrow decision task.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
When you publish or share a comparison, include the test set description and size, output mode, schema, decoding configuration, retry policy, number of runs, and scoring rules. Without those details, readers cannot tell whether an apparent difference reflects the model, its constraints, or the evaluation setup.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




