An agent harness can improve by turning observed failures into small, testable changes to its prompts, tools, context handling, memory, and control flow. To show that those changes work beyond the tasks used to develop them, keep optimization separate from validation and final testing, hide held-out examples and scores from the optimizer, screen edits for benchmark-specific logic, and compare against simple baselines using matched compute budgets. The evidence is promising but mixed: some studies report held-out or cross-family gains, while another finds limited transfer and no consistent advantage over test-time scaling.
What an agent harness is—and what it means to improve one
An agent harness is the software around a language-model agent: it determines what information the model receives, which tools it can use, how context is managed, and how execution and task completion are controlled. Harness improvement changes that surrounding system rather than necessarily changing the underlying model. Several studies discussed below keep the base model fixed while changing the harness.
As an Amazon Associate I earn from qualifying purchases.
That distinction matters when interpreting a better pass rate. A result may reflect a more effective harness, but it may also reflect more search or inference compute, repeated exposure to evaluation feedback, or a change in the underlying model. A credible comparison needs to identify what stayed fixed and what changed.
Free tools Windows power users keep installed
One-click scans. No signup required.
What the reported results do—and do not—show
The figures below come from different papers, models, benchmark versions, splits, and methods. They are not a head-to-head comparison, and results reported by a paper’s authors are not independent replications. Read each number in its experimental setting rather than as a general estimate of how much harness evolution helps.
#1 Best Overall
- High-Performance AI Processor:The MS-02 Ultra features an Intel Core Ultra 9 285HX (24C/24T, up to 5.5 GHz, 13 TOPS NPU), delivering fast and efficient performance for AI inference, algorithm development, and media workloads. A PCIe x16 expansion slot supports desktop-class GPU upgrades for advanced model training and accelerated computing tasks. It's ideal for creators, engineers, and teams handling intensive parallel workloads.
- 4 × M.2 PCIe 4.0 + 4 × DDR5 SODIMM slots:Four DDR5 SODIMM slots support up to 256 GB of memory, while ECC helps maintain data integrity in mission-critical environments. Four PCIe 4.0 M.2 slots support up to 24 TB of storage, supporting RAID 0/1/5/10, combining high-speed performance with data protection. It allows for the creation of independent scratch disks, media libraries, and project drives, providing high-throughput for production workflows.
- PCIe & USB 4.0 v2: Up to three PCIe slots can be equipped, including a dual-slot x16 GPU. The main slot supports PCIe 5.0, meeting the needs of high-bandwidth creative and computing workloads. USB 4.0 v2 (80Gbps) supports high-bandwidth external storage and displays.
- Ultra-fast Networking: Wi-Fi 7 further enhances wireless performance with next-generation speeds and low-latency stability. Intelligent bandwidth switching optimizes throughput in different network environments, ensuring optimal performance for enterprise or local networks. Dual 25GbE ports (providing up to approximately 3.125 GB/s bandwidth, about 25 times faster than traditional 1GbE), enabling seamless large-scale file transfers and parallel computing. 10GbE and 2.5GbE ports, with support for Intel vPro technology, ensure enterprise-grade remote management and deployment flexibility.
- Server-grade thermal architecture: Utilizing a dedicated CPU/GPU airflow design, equipped with a 6-pipe dual-fan cooler, it maintains stable performance even under sustained loads, delivering up to 140W Turbo power while maintaining a 100W TDP, and operating with noise levels as low as 36 dB. An integrated 350W power supply ensures stable and reliable output for demanding computing tasks and fully loaded extended configurations.
| Study and setup | Reported result | What to take from it |
|---|---|---|
| Qiankai Xu, Self-Evolving Harness on Multiple Tasks with the Agent as Its Own Optimizer (September 2026). The same frozen model acts as solver and proposer; tasks cover five benchmarks, with training tasks separated from held-out tasks and evaluation on five additional out-of-distribution benchmarks not used during evolution. | After the first evolution stage, the authors report average improvements of 4.48 points on in-distribution benchmarks and 12.64 points on out-of-distribution benchmarks. | This setup tests transfer across tasks and includes out-of-distribution evaluation. The reported gains apply to this method and evaluation, not to harness evolution generally. |
| Self-Harness (2026), evaluated on Terminal-Bench 2.0 with held-out tasks. | Reported pass rates: MiniMax M2.5, 40.5% to 61.9%; Qwen3.5-35B-A3B, 23.8% to 38.1%; GLM-5, 42.9% to 57.1%. | The changes are attributed to a loop of trace-based weakness mining, minimal candidate edits, and regression-test validation. These rates are specific to the named models and benchmark. |
| Jiahang Lin and coauthors, Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses (latest version dated May 18, 2026), evaluated on Terminal-Bench 2. | The authors report pass@1 increasing from 69.7% to 77.0% over ten iterations, plus gains on three alternate model families without re-evolution. | The paper presents cross-family transfer for its method and setup. It is evidence worth testing, not proof that changes transfer reliably to other harnesses or tasks. |
| Microsoft Research’s June 2026 description of Retrospective Harness Optimization, applied to SWE-Bench Pro. | A 59% to 78% pass-rate change after one optimization round is reported by the method’s authors. | The method uses past trajectories, self-validation and self-consistency, and pairwise self-preference rather than external grading. Self-judged preference is not equivalent to an independently graded held-out test. |
| HarnessOpt-Bench, a four-task evaluation of optimizers using separate development, validation, and test partitions. | Optimizer performance varied by task and seed regime; the cited summary does not state a single overall score. | The benchmark is designed to evaluate the optimizer itself, using a trusted execution environment that hides held-out state, meters resource use, and versions candidates. |
| Rethinking the Evaluation of Harness Evolution for Agents, evaluated on Terminal-Bench 2.1. | The authors report that harness evolution did not consistently outperform matched-budget parallel sampling and sequential refinement, and showed only marginal improvements on held-out tasks. | This counterevidence makes budget-matched baselines and genuinely independent held-out evaluation essential to claims of improvement. |
How to improve a harness without training it to the test set
Use a controlled loop in which failure analysis informs a candidate change, while an evaluation set the proposer cannot see determines whether the change generalizes.
- Freeze the comparison. Record the base-model version, starting harness, task splits, and resource budget. Keep the model and starting conditions fixed when comparing harness variants so the source of a difference is identifiable.
- Collect traces with verifiable outcomes. Look for recurring failures in agent runs and tie each proposed change to a concrete failure. Keep edits small enough to test and roll back; broad rewrites make it harder to tell which change helped or caused a regression.
- Write down a hypothesis for each edit. Log the harness component changed, the intended effect, expected task outcomes, measured results, cost changes, and the decision to accept or reject the edit. This makes the evolution auditable and helps distinguish an evidence-based change from trial and error.
- Separate optimization, validation, and final testing. The optimizer may use development tasks and feedback, but should not receive held-out examples, labels, or scores. Reserve a final test set for the finished harness. If the claim is transfer beyond the original task family, add out-of-distribution tasks that were not used during evolution.
- Screen for benchmark-specific logic. Check candidate changes for special handling keyed to task names, entities, answers, or other benchmark particulars. Run regression tests, retain the history of accepted and rejected edits, and require an acceptance threshold that accounts for evaluation noise.
- Compare under matched budgets. Run simple alternatives such as parallel sampling or sequential refinement with comparable task feedback and inference budgets. Report resource use as well as success: extra search compute can explain an apparent gain even when the harness itself is not more efficient.
- Report enough detail to reproduce the claim. Name the model, harness version, benchmark version, split, optimization rounds, resource budget, and whether the result is held out or out of distribution. Scores from separate papers should not be ranked as if they came from a common test.
How to interpret self-improvement methods and their evidence
Trace-led edits can make changes testable
Self-Harness describes mining weaknesses from execution traces, proposing minimal edits, and validating candidates with regression tests. Agentic Harness Engineering makes harness components editable as files, distills trajectories into an evidence corpus, and pairs edits with predictions checked against later outcomes. These approaches share a useful discipline: connect each change to an observable failure and a prediction that can be checked, rather than accepting a change because it sounds plausible.
Rank #2
Self-supervision is not the same as independent evaluation
Retrospective Harness Optimization uses an agent’s past trajectories and its own validation, consistency checks, and pairwise preferences. That can provide a way to propose or rank changes without an external grader, but a system’s preference for one candidate does not establish that it performs better on tasks the system has never seen. Keep the final evaluation independent of the process that generated and selected edits.
Regularization can constrain the search
Google Research’s RRSI repository documents several controls: screen candidates for suite-specific logic, set an acceptance floor adjusted for evaluation noise, require measured gains to justify extra inference tokens, and prune components that no longer help. These are method-design principles, not evidence that one optimizer is best. Their value is in making the search less willing to accept brittle or unnecessarily expensive changes.
Rank #3
- Professional AI & Creator Workstation: AMD Radeon AI PRO R9700 GPU with 32GB GDDR6 is engineered for AI development, professional content creation, and compute-intensive workloads.
- Massive 32GB Memory Capacity: 32GB of GDDR6 memory on a 256-bit bus provides ample bandwidth for large AI models, 8K video editing, and complex 3D rendering.
- Advanced RDNA 4 with AI Accelerators: 64 Compute Units with 3rd Gen Ray Tracing and dedicated 2nd Gen AI Accelerators for groundbreaking AI performance and visual computing.
- Professional Blower Cooling: Efficient single blower design exhausts heat directly out of the chassis, ideal for multi-GPU workstation and server configurations.
- Enterprise-Grade Thermal Solution: Vapor chamber heatsink with industrial Honeywell PTM7950 thermal interface material ensures reliable cooling under sustained professional loads.
What a convincing claim of generalization requires
- Independent tasks: held-out examples and scores remain unavailable to the proposer, and out-of-distribution tests come from tasks not used to evolve the harness.
- Fair alternatives: simple search and test-time scaling baselines receive comparable feedback and compute.
- Cost and regressions: report inference or resource use alongside success, and show whether gains come with failures elsewhere or higher operating cost.
- Reproducibility: preserve versions, split boundaries, edit histories, and the conditions under which each score was measured.
- Scoped conclusions: distinguish a result on a named model and benchmark from a claim about harnesses in general.
Current reports do not establish a universally best method. Positive held-out and cross-family results coexist with evidence of limited transfer and no consistent advantage over matched-budget test-time scaling. The most defensible conclusion is conditional: harness evolution can produce useful gains, but benchmark memorization is ruled out only by a protected evaluation design, credible baselines, and transparent accounting of costs and uncertainty.
Quick Recap
Best Value
- 【High-Performance APU】The MS-S1 MAX features an AMD Ryzen AI Max+ 395 APU, integrating a Zen 5 architecture CPU (up to 5.1GHz, 16C/32T, 64M L3 Cache), an RDNA 3.5 GPU, and an NPU (50 TOPS). The total system output is 126 TOPS. It provides powerful parallel computing capabilities for demanding AI workflows. It is ideal for running local LLMs, multimodal models, and computationally intensive tasks
- 【128GB UMA Memory】Equipped with up to 128GB of LPDDR5x-8000MT/s unified memory, it enables the CPU and GPU to access a shared, high-bandwidth memory pool with extremely low latency. Ideal for large-scale AI inference, 3D workloads, and complex timelines in video editing. It eliminates traditional VRAM bottlenecks, ensuring smoother data transfer during high-intensity computations. The UMA design maximizes performance stability under high loads
- 【Flexible Expansion】The MS-S1 MAX features USB4 V2 (up to 80Gbps), dual 10GbE LAN, HDMI 2.1 (up to 8K60), a full-length PCIe x16 expansion slot, and dual M.2 slots supporting up to 16TB RAID 0/1. Wi-Fi 7 provides stronger signal coverage and a more stable wireless experience. The slide-out design facilitates upgrades and maintenance. It easily adapts to personal, studio, or rack-mount enterprise environments
- 【High-Efficiency Cooling System】Utilizing an aerospace-grade aluminum alloy chassis, copper base plate, six heat pipes, dual turbine fans, and advanced PCM thermal conductive material, it maintains stable cooling performance even under continuous load. This system supports 130W continuous power and 160W peak power operation, with a built-in 320W power supply. It boasts multiple global certifications including CCC, FCC, UL, CE, and UKCA, ensuring stable and reliable operation in various environments
- 【Cluster Design】Two MS-S1 MAX units can be configured as a dual-unit cluster to run a large 235B Q4 model locally, achieving an output speed of 10.87 tok/s. Supporting 2U rack deployment, multiple MS-S1 MAX units can be cascaded into a distributed cluster to create a high-efficiency AI computing center. A cluster of four MS-S1 MAX units successfully ran a DeepSeek-R1 671B Q4 large model. A reserved cluster power-on interface allows for unified start-up and shutdown
Rank #4
- FAST RUNS IN THE FAMILY — The 16-inch MacBook Pro with the M5 Pro or M5 Max chip brings next-generation speed and powerful on-device AI to personal, professional, and creative tasks. With all-day battery life, double the starting storage,* and a breathtaking Liquid Retina XDR display, it’s pro in every way.*
- BUCKLE UP — Along with a next-generation CPU, faster unified memory, and up to 2x faster SSD storage,* M5 Pro and M5 Max feature a more powerful GPU with a Neural Accelerator built into each core, delivering faster AI performance and on-device training capabilities. So you can blaze through demanding workloads at mind-bending speeds.
- BUILT FOR AI — Apple silicon, and every major component that powers it, is designed to run demanding on-device AI workloads like LLM inference and training. And Apple Intelligence helps you write, express yourself, and get things done effortlessly with groundbreaking privacy protections at every step.*
- ALL-DAY BATTERY LIFE — MacBook Pro delivers the same exceptional performance whether it’s running on battery or plugged in.*
- MACOS RUNS APPS FAST — All your go-to apps run lightning fast in macOS, including built-in apps like FaceTime and Messages. Plus, built-in virus protection and free software updates help keep your Mac running smoothly and securely.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




