A coding agent does not benefit from every harness feature in the same way. In a 2026 study across four models and two coding benchmarks, context management mattered most when the context window was tight; planning helped some models and hurt or barely changed others; and structured tools versus bash-only produced different results by model and task. The findings are evidence about one harness and these evaluations—not a universal recipe or a ranking of commercial coding agents.
What the study tested
Run-Ze Fan and eight coauthors’ An Empirical Study of Harness Design for Coding Agents was published on 17 September 2026. It reports 176 matched settings, varying three harness components while holding a lightweight ReAct-style execution loop fixed: persistent planning, the agent’s action interface, and context management. The models were Nemotron-3 30B, 120B, and 550B, plus Mistral-Medium-3.5-128B. The benchmarks were SWE-Bench Verified, with 500 tasks, and Terminal-Bench 2.1, with 89 tasks. Read the paper on arXiv.
The context-management comparisons covered nominal windows of 32k, 64k, 96k, and 128k tokens. Planning and action-space comparisons were narrower: they were run at the T4 context-management setting and 128k tokens. That distinction matters when applying the results: the paper does not establish how planning or interface choice interacts with smaller windows or other context policies.
When does context management help a coding agent?
Its clearest benefit appeared when the context window was small enough that unmanaged runs often overflowed. Across the study’s managed tiers, mean success on SWE-Bench Verified was 35.7 percentage points higher than no management at 32k tokens, but the advantage was 2.7 points at 128k. On Terminal-Bench 2.1, the corresponding advantages were 9.5 points at 32k and 2.8 points at 128k. These are averages across tested configurations, not guaranteed gains for a particular model or task.
#1 Best Overall
| Nominal context window | Benchmark | Managed-tier mean success advantage over no management | No-management overflow rate |
|---|---|---|---|
| 32k tokens | SWE-Bench Verified | 35.7 percentage points | 78.7% average |
| 32k tokens | Terminal-Bench 2.1 | 9.5 percentage points | 61.0% average |
| 128k tokens | SWE-Bench Verified | 2.7 percentage points | 8.7% average |
| 128k tokens | Terminal-Bench 2.1 | 2.8 percentage points | 12.1% average |
Every managed tier avoided overflow failures in the tested settings. The pattern suggests that context management primarily helped trajectories keep running when unmanaged context would fill, rather than improving each local decision. The authors tested policies ranging from no compaction (T0) through stale-output elision, optional recoverable external storage, LLM-generated summarization, and a staged policy: elision first, then selective summarization (T4).
Why T4 stood out on cost
T4 had the lowest average cost at every tested context budget and the lowest mean cost in seven of the eight model-benchmark combinations, while delivering broadly comparable success to the other managed tiers. That makes it the strongest efficiency profile among the context policies tested, not proof that the same ordering holds in other harnesses or workloads.
Rank #2
What the recall result does—and does not—show
Adding recoverable recall to elision did not produce a meaningful average accuracy advantage in this experiment. In 32 matched comparisons, T2 beat T1 in 15, lost in 14, and tied in three; the equal-weight mean difference was -0.36 percentage points. Across 64 T2 and T4 settings, 56.3% never invoked recall. Those observations suggest recall was often unnecessary for these runs, but they do not show that recall mechanisms are generally useless.
Does giving an AI coding agent a plan improve results?
Not consistently. The planning comparison contrasted a persistent task plan with no plan, and its impact varied by model. For Nemotron-3 30B, planning increased success by 11.6 percentage points on SWE-Bench and 4.5 points on Terminal-Bench, while raising cost on both. Without planning, its median SWE-Bench trajectory shrank from 40 turns to five, and runs ending without an edit rose from 27.8% to 68.6%. This supports the interpretation that a plan helped the weaker model persist long enough to make an edit.
Rank #3
For Nemotron-3 550B and Mistral-Medium-3.5-128B, planning instead reduced SWE-Bench inference cost by about 30% and 32%, respectively, with success-rate changes of -2.0 and -0.4 percentage points. The authors interpret this as planning helping stronger models avoid redundant verification. Nemotron-3 120B showed no consistent effect. These are model- and benchmark-specific outcomes, not evidence that planning is inherently a quality improvement or an efficiency penalty.
Do coding agents work better with structured tools or just bash?
The study compared a structured interface exposing file, search, web, and shell tools with a bash-only interface. It was not an isolated test of tool count: the designs also differed in interface instructions, file-state tracking, read-before-write enforcement, and automatic post-edit diagnostics. Results therefore describe the complete interfaces tested.
Rank #4
| Model and benchmark | Effect on success | Other reported effect |
|---|---|---|
| Nemotron-3 30B, SWE-Bench Verified | Structured tools: +15.0 percentage points versus bash-only | — |
| Nemotron-3 30B, Terminal-Bench 2.1 | Structured tools: +10.1 percentage points versus bash-only | With bash-only, 66% of trajectories ended after calls incompatible with the available interface |
| Nemotron-3 550B, SWE-Bench Verified | Bash-only: +3.6 percentage points versus structured tools | Bash-only cost was 53% lower |
| Nemotron-3 550B, Terminal-Bench 2.1 | Bash-only: +5.6 percentage points versus structured tools | Bash-only cost was 30% lower |
| Mistral-Medium-3.5-128B, SWE-Bench Verified | Structured tools: +23.2 percentage points versus bash-only | — |
| Mistral-Medium-3.5-128B, Terminal-Bench 2.1 | Bash-only: +6.7 percentage points versus structured tools | — |
The practical signal is conditional: structured tools helped Nemotron-3 30B and Mistral on repository repair, while bash-only sometimes gave stronger shell-capable models a cheaper or more successful path, especially on Terminal-Bench. Task structure matters alongside model capability and shell proficiency; this paper does not identify a universal crossover point.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How to apply the findings to a harness decision
Use the study as a guide to questions worth measuring in your own workload, rather than a prescriptive configuration. Compare choices along three axes:
Recommended Free Tools
Best Value
- Context pressure: if runs regularly approach the context limit, test compaction and track overflow alongside success. If windows are ample, the measured success advantage of management was much smaller.
- Model capability and shell proficiency: weaker models may need explicit affordances and guardrails, while a strong shell-capable model may benefit from a less restrictive interface.
- Task structure: repository issue repair and command-line-centric tasks can reward different interfaces. Evaluate them separately rather than averaging away meaningful differences.
For each comparison, track success rate, inference cost, overflow rate, and trajectory length. The study’s annotations were produced by LLM judges; the reported aggregate judge-human agreement was approximately 94.2%, with weighted mean Cohen’s kappa of 0.929. Those figures support the annotation process but do not remove uncertainty from the benchmark results.
What the results cannot establish
The study evaluates four models and two benchmarks; SWE-Bench Verified consists of Python repositories. Each task was run once per setting, so the findings do not measure run-to-run variability. Terminal-Bench has 89 tasks, and many of its contrasts did not reach significance under paired McNemar analysis. The reported differences should therefore be read as outcomes in these tested settings, not stable guarantees.
Most importantly, planning and action-space ablations were performed only with T4 context management at 128k. Their interaction with tight context windows or alternative context policies remains unresolved. The study is a useful component-level examination of one harness implementation, but it does not test all combinations or rank commercial coding-agent products.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →




