Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan Now×
Skip to content
World desk5 min

Stratify the Task Pack Before Averaging Agent Scores

A single average across a mixed agent task pack describes one blend of tasks. Here is how to declare strata, report them separately, and document the weighting behind any overall score.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Averaging agent scores across a mixed task pack produces one number that describes one particular blend of tasks, not the agent’s ability in general. To make a benchmark score meaningful, split the pack into declared strata, report results for each stratum, and state the weighting rule behind any overall figure.

Why should you stratify an agent benchmark before averaging its scores?

An agent benchmark is a collection of tasks, and collections are rarely uniform. A pack may mix bug fixes with feature work, short edits with multi-file refactors, or tasks that a strong agent solves almost every time with tasks that few agents solve. A single pass rate hides all of that. Two agents can post the same overall score while succeeding on entirely different parts of the pack, and a reader has no way to see the difference from the headline number alone.

Chris Ge, Daria Kryvosheieva, Daniel Fried, Uzay Girit, and Kaivalya Hariharan make this point directly in Agent psychometrics: Task-level performance prediction in agentic coding benchmarks (2026): single-number metrics obscure the diversity of tasks within a benchmark. Their work builds a task-level prediction framework using task features and an item-response-theory approach, which treats task difficulty as something to be modeled rather than ignored.

Choosing strata you can defend

A stratum is a group of tasks that share a meaningful property for the question you are asking. The right dimensions depend on the pack, so the categories should be declared and defined in the report rather than borrowed from a generic taxonomy. Common choices include:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
  • Task family: the kind of work the task requires, such as debugging, test writing, feature implementation, or dependency migration.
  • Difficulty: a band based on a measurable property, such as the historical pass rate of earlier runs, the number of files a correct fix touches, or the size of the required change.
  • Source or construction method: tasks drawn from real issue trackers versus tasks written by benchmark authors, when the pack mixes both.

A good stratum is one where performance within it is more consistent than performance across the whole pack. If every task in a group produces the same kind of result, the grouping is informative. If a stratum still contains wildly different tasks, split it further or reconsider the dimension.

A reporting method you can apply

  1. State the evaluation question first. Decide whether you want to know how an agent performs on the pack as it exists, how it compares across kinds of work, or how it would do on a workload you care about. Each question calls for a different summary.
  2. Declare the strata and their definitions. List each category, the rule that assigns tasks to it, and the number of tasks it contains.
  3. Report results per stratum. Include the task count next to each figure so readers can judge how much evidence each category carries.
  4. If you publish one overall score, document the weighting. Say whether each task or each stratum counts equally, or whether weights reflect a target workload.
  5. Record the agent setup. Note the model, scaffold, tool access, and run configuration, because a result is tied to the system that produced it.
  6. Limit the claim to the pack. A result on one pack is evidence about that pack and setup, not a complete statement about all agent work.

These steps are a practical synthesis of the cited papers’ emphasis on task heterogeneity and task selection. Neither paper is a published standard that mandates this exact protocol.

Rank #2
Sale
Statistics Laminate Reference Chart: Parameters, Variables, Intervals, Proportions (Quickstudy: Academic )
  • This guide is a perfect overview for the topics covered in introductory statistics courses.

What an overall score actually means

Once you average, you have chosen a weighting, whether or not you wrote it down. The weighting determines which mix of work the number describes.

Weighting rule What the overall number describes When it fits Main caveat
Task-weighted (every task counts equally) Performance on the pack’s own mix of tasks You want to describe the benchmark as built Large strata dominate, so a small but important category barely moves the score
Stratum-weighted (every category counts equally) A balanced view across categories regardless of their size You want to compare how agents handle each kind of work Small strata with few tasks can swing the average, so report their counts
Target-workload weighted (weights set to a declared external mix) Performance on the mix you say you care about You are making a claim about a specific real-world workload Requires a justified, documented source for the target weights

The cited papers identify task diversity as a concern but do not prescribe a universal weighting scheme. Choose the rule that matches your question, state it, and show enough per-category data that a reader can compute an alternative weighting if they disagree with yours.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3

What task selection can and cannot save

Running every task in a large pack is expensive, so it is natural to ask whether a smaller subset can give the same answer. Franck Ndzomga’s Efficient Benchmarking of AI Agents (2026) studies exactly that question.

The reported reduction and its conditions

In the setting that paper evaluated, selecting tasks with intermediate historical pass rates, between 30% and 70%, reduced the number of evaluation tasks by 44–70% while maintaining high rank fidelity. That is a result from a named study under its tested conditions. It is not a guarantee that every benchmark or agent will see the same saving, and it should not be quoted without those conditions attached.

Rank prediction is not absolute-score prediction

The same work reports that absolute score prediction degrades under scaffold-driven distribution shift, meaning when the agent’s scaffolding changes in ways the selection did not account for. A subset can preserve the ordering of agents while giving a poorer estimate of how high any one agent’s score would be on the full pack. When you report a reduced evaluation, say whether the claim concerns ranking or absolute performance, and keep the two separate.

Checklist for reading an agent benchmark report

  • Are the strata named, defined, and sized?
  • Are per-stratum results shown, not only an overall figure?
  • Is the weighting rule for the overall score stated?
  • Is the agent configuration, including scaffold and tool access, described?
  • Does the claim concern rankings or absolute scores?
  • If tasks were selected to reduce cost, are the selection rule and its tested conditions reported?
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Limits of current guidance

No widely adopted standard currently specifies which task strata to use or which aggregation weights to apply to agent benchmarks. The guidance above rests on the two cited papers and on general measurement reasoning about heterogeneous tasks. Treat strata and weights as design decisions that must be justified for each pack, and expect that the right categories will differ between a bug-fixing suite and a feature-building suite.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Stratification does not make a benchmark more accurate by itself. It makes the benchmark’s contents and the agent’s performance across them visible, so that readers can see what an average was hiding.

Frequently Asked Questions

Do I still need to stratify if my benchmark tasks look similar?

Probably less than for a mixed pack, but check before assuming. Similar-looking tasks can still differ in difficulty or in how many steps a correct solution requires. A quick look at per-task results usually shows whether one group of tasks behaves differently from the rest.

How many tasks should a stratum contain before I report it?

No threshold is established in the cited papers. Report the task count with every stratum so readers can judge how much evidence sits behind each figure, and flag small strata as indicative rather than conclusive.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.