Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →A coding agent can pass a test suite and still break repository rules. To find out whether it follows instructions, define each rule in observable terms, inspect the agent’s actions as well as its final changes, and score task success separately from rule compliance. Published studies report failures involving repository policies, AI-contribution guidelines, and task plans—but their results apply to specific benchmarks, not every coding agent.
Why passing tests is not enough
Functional tests show whether a patch behaves as expected under the tests that were run. They do not establish that the agent followed a required workflow, used only permitted tools, disclosed AI assistance, or sought human approval at a required decision point. The authors of SWE-CC put the distinction plainly: “passing functional tests differs fundamentally from producing a high-quality contribution acceptable for merging.” Read the SWE-CC paper.
As an Amazon Associate I earn from qualifying purchases.
That distinction matters because some requirements govern the process, not just the output. An agent might produce a correct patch after skipping a required verification gate. Or it might comply with the visible instructions while violating a repository policy during an intermediate action. A useful evaluation therefore treats task completion and rule compliance as separate outcomes.
What recent studies found
Four 2026 preprints examine related but different questions. Their percentages and counts cannot be compared as if they shared one scale: each uses its own rules, tasks, agents, and scoring approach.
#1 Best Overall
Repository policies can be violated during execution
The SWE-CC authors evaluated 500 end-to-end contribution tasks drawn from SWE-bench Verified extensions, using policies derived from documentation in 12 repositories. They report that 43.1% of applicable project policies were violated, and that nearly half of the violations occurred during intermediate execution. This is a result for the agents and tasks in that benchmark, not an estimate for all coding agents. SWE-CC audits runtime behavior as well as final deliverables. SWE-CC paper.
Agents may not retrieve or obey repository AI rules
RepoComplianceBench examined 106 issues across 49 repositories. Its tests focused on refusal, truthful disclosure, verification gates, and escalation to a human under repository rules about AI contributions. The authors report that agents almost never proactively retrieved those rules and did not refuse in repositories that banned AI assistance under the tested conditions. The result is specific to that benchmark’s repositories, issues, and evaluation setup. RepoComplianceBench paper.
Following a task plan depends partly on the plan
“From Plan to Action” analyzed 21,120 trajectories across four LLMs, two benchmarks, and eight plan variations. The study reports that a standard plan improved issue resolution, periodic reminders mitigated plan violations, and a subpar plan could hurt performance. These findings concern the study’s tested plans and settings; they do not show that reminders guarantee compliance in other projects. “From Plan to Action” paper.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesOrdinary compliance can overstate whether a rule caused behavior
Harness-IF evaluated 12 models and reported overall accuracy from 72.1% to 85.9%, with Against-Prior Accuracy from 66.1% to 78.6%. The lower accuracy on rules that oppose agents’ unprompted defaults highlights a measurement problem: an agent may do what a rule asks even when it has not meaningfully followed that rule. As the authors write, “When a coding agent obeys a rule, it may simply have been going to do that anyway.” The aggregate results apply to Harness-IF’s 60 multi-turn items, rule library, and tested builds—not to coding agents generally. Harness-IF paper.
Rank #3
How to test a coding agent in your repository
Make the rule testable before handing the task to the agent. A statement such as “follow our contribution policy” is difficult to score unless you identify the required behavior and the evidence that would show it happened.
- Choose a rule and define a pass condition. For example: “Before editing, inspect the repository’s contribution instructions” can be checked against the action history. “Run the required test command before reporting completion” can be checked against tool logs and the command’s result.
- Record the exact test setup. Keep the repository and commit, exact rule text, task, agent and model version, scaffold and configuration, tool permissions, verifier, run count, and pass/fail criteria. If any of these change, record the change rather than treating the runs as equivalent.
- Watch the trajectory, not only the patch. Check whether the agent read relevant instructions, followed required workflows, stayed within tool permissions, completed verification gates, disclosed AI contribution when required, and escalated decisions reserved for a human. Preserve the execution evidence needed to judge those actions.
- Score task success and compliance separately. A correct patch is not a compliance pass if the agent skipped a required step. Likewise, an agent can follow the process and still fail to solve the task. Report both outcomes.
- Probe rules that conflict with default behavior. Where practical, compare a run with the rule to a run without it, as Harness-IF does, and look for behavior that changes in response to the instruction. A single compliant-looking action may not show that the rule caused it.
- Repeat and report the limits. Include the number of runs, failures, and ambiguous cases. One run can show what happened once; it does not establish a stable property of the agent.
How to interpret a compliance score
A percentage is meaningful only with its measurement frame. Before drawing a conclusion, establish what counted as a rule, where those rules appeared, whether scoring covered actions, deliverables, or both, and whether the benchmark tested routine instructions or instructions that opposed an agent’s defaults. Also consider the sampled repositories, tasks, models, scaffolds, and whether evaluation relied on deterministic checks, human judgment, or both.
Rank #4
The studies above answer related but distinct questions: SWE-CC examines project-policy violations during contributions; RepoComplianceBench focuses on repository rules for AI contributions; “From Plan to Action” studies adherence to task plans; Harness-IF tests instruction-following, including rules that conflict with defaults. Their scores should not be ranked against one another.
Free tools Windows power users keep installed
One-click scans. No signup required.
All four are arXiv research records available by October 7, 2026. They do not establish how every commercial coding agent will behave. Use them as evidence that different kinds of rule-following can fail, not as a universal product rating.
Quick Recap
Best Value
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




