Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Averaging agent scores across a mixed task pack produces one number that describes one particular blend of tasks, not the agent’s ability in general. To make a benchmark score meaningful, split the pack into declared strata, report results for each stratum, and state the weighting rule behind any overall figure.
Why should you stratify an agent benchmark before averaging its scores?
An agent benchmark is a collection of tasks, and collections are rarely uniform. A pack may mix bug fixes with feature work, short edits with multi-file refactors, or tasks that a strong agent solves almost every time with tasks that few agents solve. A single pass rate hides all of that. Two agents can post the same overall score while succeeding on entirely different parts of the pack, and a reader has no way to see the difference from the headline number alone.
Chris Ge, Daria Kryvosheieva, Daniel Fried, Uzay Girit, and Kaivalya Hariharan make this point directly in Agent psychometrics: Task-level performance prediction in agentic coding benchmarks (2026): single-number metrics obscure the diversity of tasks within a benchmark. Their work builds a task-level prediction framework using task features and an item-response-theory approach, which treats task difficulty as something to be modeled rather than ignored.
Choosing strata you can defend
A stratum is a group of tasks that share a meaningful property for the question you are asking. The right dimensions depend on the pack, so the categories should be declared and defined in the report rather than borrowed from a generic taxonomy. Common choices include:
#1 Best Overall
- Task family: the kind of work the task requires, such as debugging, test writing, feature implementation, or dependency migration.
- Difficulty: a band based on a measurable property, such as the historical pass rate of earlier runs, the number of files a correct fix touches, or the size of the required change.
- Source or construction method: tasks drawn from real issue trackers versus tasks written by benchmark authors, when the pack mixes both.
A good stratum is one where performance within it is more consistent than performance across the whole pack. If every task in a group produces the same kind of result, the grouping is informative. If a stratum still contains wildly different tasks, split it further or reconsider the dimension.
A reporting method you can apply
- State the evaluation question first. Decide whether you want to know how an agent performs on the pack as it exists, how it compares across kinds of work, or how it would do on a workload you care about. Each question calls for a different summary.
- Declare the strata and their definitions. List each category, the rule that assigns tasks to it, and the number of tasks it contains.
- Report results per stratum. Include the task count next to each figure so readers can judge how much evidence each category carries.
- If you publish one overall score, document the weighting. Say whether each task or each stratum counts equally, or whether weights reflect a target workload.
- Record the agent setup. Note the model, scaffold, tool access, and run configuration, because a result is tied to the system that produced it.
- Limit the claim to the pack. A result on one pack is evidence about that pack and setup, not a complete statement about all agent work.
These steps are a practical synthesis of the cited papers’ emphasis on task heterogeneity and task selection. Neither paper is a published standard that mandates this exact protocol.
Rank #2
- This guide is a perfect overview for the topics covered in introductory statistics courses.
What an overall score actually means
Once you average, you have chosen a weighting, whether or not you wrote it down. The weighting determines which mix of work the number describes.
| Weighting rule | What the overall number describes | When it fits | Main caveat |
|---|---|---|---|
| Task-weighted (every task counts equally) | Performance on the pack’s own mix of tasks | You want to describe the benchmark as built | Large strata dominate, so a small but important category barely moves the score |
| Stratum-weighted (every category counts equally) | A balanced view across categories regardless of their size | You want to compare how agents handle each kind of work | Small strata with few tasks can swing the average, so report their counts |
| Target-workload weighted (weights set to a declared external mix) | Performance on the mix you say you care about | You are making a claim about a specific real-world workload | Requires a justified, documented source for the target weights |
The cited papers identify task diversity as a concern but do not prescribe a universal weighting scheme. Choose the rule that matches your question, state it, and show enough per-category data that a reader can compute an alternative weighting if they disagree with yours.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteRank #3
What task selection can and cannot save
Running every task in a large pack is expensive, so it is natural to ask whether a smaller subset can give the same answer. Franck Ndzomga’s Efficient Benchmarking of AI Agents (2026) studies exactly that question.
The reported reduction and its conditions
In the setting that paper evaluated, selecting tasks with intermediate historical pass rates, between 30% and 70%, reduced the number of evaluation tasks by 44–70% while maintaining high rank fidelity. That is a result from a named study under its tested conditions. It is not a guarantee that every benchmark or agent will see the same saving, and it should not be quoted without those conditions attached.
Rank #4
Rank prediction is not absolute-score prediction
The same work reports that absolute score prediction degrades under scaffold-driven distribution shift, meaning when the agent’s scaffolding changes in ways the selection did not account for. A subset can preserve the ordering of agents while giving a poorer estimate of how high any one agent’s score would be on the full pack. When you report a reduced evaluation, say whether the claim concerns ranking or absolute performance, and keep the two separate.
Checklist for reading an agent benchmark report
- Are the strata named, defined, and sized?
- Are per-stratum results shown, not only an overall figure?
- Is the weighting rule for the overall score stated?
- Is the agent configuration, including scaffold and tool access, described?
- Does the claim concern rankings or absolute scores?
- If tasks were selected to reduce cost, are the selection rule and its tested conditions reported?
Limits of current guidance
No widely adopted standard currently specifies which task strata to use or which aggregation weights to apply to agent benchmarks. The guidance above rests on the two cited papers and on general measurement reasoning about heterogeneous tasks. Treat strata and weights as design decisions that must be justified for each pack, and expect that the right categories will differ between a bug-fixing suite and a feature-building suite.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Best Value
Stratification does not make a benchmark more accurate by itself. It makes the benchmark’s contents and the agent’s performance across them visible, so that readers can see what an average was hiding.
Frequently Asked Questions
Do I still need to stratify if my benchmark tasks look similar?
Probably less than for a mixed pack, but check before assuming. Similar-looking tasks can still differ in difficulty or in how many steps a correct solution requires. A quick look at per-task results usually shows whether one group of tasks behaves differently from the rest.
How many tasks should a stratum contain before I report it?
No threshold is established in the cited papers. Report the task count with every stratum so readers can judge how much evidence sits behind each figure, and flag small strata as indicative rather than conclusive.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Free tools Windows power users keep installed
One-click scans. No signup required.




