Hardware FixRecommendedDevice not working? Your driver may be the problemCheck updates for common hardware issues.Fix DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan Now×
Skip to content
World desk4 min

Coding Agent Rule-Following: How to Test It Without Guessing

A coding agent can produce a working patch while breaking repository rules. Here’s how to evaluate its actions and interpret what recent studies do—and do not—show.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A coding agent can pass a test suite and still break repository rules. To find out whether it follows instructions, define each rule in observable terms, inspect the agent’s actions as well as its final changes, and score task success separately from rule compliance. Published studies report failures involving repository policies, AI-contribution guidelines, and task plans—but their results apply to specific benchmarks, not every coding agent.

Why passing tests is not enough

Functional tests show whether a patch behaves as expected under the tests that were run. They do not establish that the agent followed a required workflow, used only permitted tools, disclosed AI assistance, or sought human approval at a required decision point. The authors of SWE-CC put the distinction plainly: “passing functional tests differs fundamentally from producing a high-quality contribution acceptable for merging.” Read the SWE-CC paper.

As an Amazon Associate I earn from qualifying purchases.

That distinction matters because some requirements govern the process, not just the output. An agent might produce a correct patch after skipping a required verification gate. Or it might comply with the visible instructions while violating a repository policy during an intermediate action. A useful evaluation therefore treats task completion and rule compliance as separate outcomes.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What recent studies found

Four 2026 preprints examine related but different questions. Their percentages and counts cannot be compared as if they shared one scale: each uses its own rules, tasks, agents, and scoring approach.

Repository policies can be violated during execution

The SWE-CC authors evaluated 500 end-to-end contribution tasks drawn from SWE-bench Verified extensions, using policies derived from documentation in 12 repositories. They report that 43.1% of applicable project policies were violated, and that nearly half of the violations occurred during intermediate execution. This is a result for the agents and tasks in that benchmark, not an estimate for all coding agents. SWE-CC audits runtime behavior as well as final deliverables. SWE-CC paper.

Agents may not retrieve or obey repository AI rules

RepoComplianceBench examined 106 issues across 49 repositories. Its tests focused on refusal, truthful disclosure, verification gates, and escalation to a human under repository rules about AI contributions. The authors report that agents almost never proactively retrieved those rules and did not refuse in repositories that banned AI assistance under the tested conditions. The result is specific to that benchmark’s repositories, issues, and evaluation setup. RepoComplianceBench paper.

Following a task plan depends partly on the plan

“From Plan to Action” analyzed 21,120 trajectories across four LLMs, two benchmarks, and eight plan variations. The study reports that a standard plan improved issue resolution, periodic reminders mitigated plan violations, and a subpar plan could hurt performance. These findings concern the study’s tested plans and settings; they do not show that reminders guarantee compliance in other projects. “From Plan to Action” paper.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Ordinary compliance can overstate whether a rule caused behavior

Harness-IF evaluated 12 models and reported overall accuracy from 72.1% to 85.9%, with Against-Prior Accuracy from 66.1% to 78.6%. The lower accuracy on rules that oppose agents’ unprompted defaults highlights a measurement problem: an agent may do what a rule asks even when it has not meaningfully followed that rule. As the authors write, “When a coding agent obeys a rule, it may simply have been going to do that anyway.” The aggregate results apply to Harness-IF’s 60 multi-turn items, rule library, and tested builds—not to coding agents generally. Harness-IF paper.

How to test a coding agent in your repository

Make the rule testable before handing the task to the agent. A statement such as “follow our contribution policy” is difficult to score unless you identify the required behavior and the evidence that would show it happened.

  1. Choose a rule and define a pass condition. For example: “Before editing, inspect the repository’s contribution instructions” can be checked against the action history. “Run the required test command before reporting completion” can be checked against tool logs and the command’s result.
  2. Record the exact test setup. Keep the repository and commit, exact rule text, task, agent and model version, scaffold and configuration, tool permissions, verifier, run count, and pass/fail criteria. If any of these change, record the change rather than treating the runs as equivalent.
  3. Watch the trajectory, not only the patch. Check whether the agent read relevant instructions, followed required workflows, stayed within tool permissions, completed verification gates, disclosed AI contribution when required, and escalated decisions reserved for a human. Preserve the execution evidence needed to judge those actions.
  4. Score task success and compliance separately. A correct patch is not a compliance pass if the agent skipped a required step. Likewise, an agent can follow the process and still fail to solve the task. Report both outcomes.
  5. Probe rules that conflict with default behavior. Where practical, compare a run with the rule to a run without it, as Harness-IF does, and look for behavior that changes in response to the instruction. A single compliant-looking action may not show that the rule caused it.
  6. Repeat and report the limits. Include the number of runs, failures, and ambiguous cases. One run can show what happened once; it does not establish a stable property of the agent.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to interpret a compliance score

A percentage is meaningful only with its measurement frame. Before drawing a conclusion, establish what counted as a rule, where those rules appeared, whether scoring covered actions, deliverables, or both, and whether the benchmark tested routine instructions or instructions that opposed an agent’s defaults. Also consider the sampled repositories, tasks, models, scaffolds, and whether evaluation relied on deterministic checks, human judgment, or both.

The studies above answer related but distinct questions: SWE-CC examines project-policy violations during contributions; RepoComplianceBench focuses on repository rules for AI contributions; “From Plan to Action” studies adherence to task plans; Harness-IF tests instruction-following, including rules that conflict with defaults. Their scores should not be ranked against one another.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

All four are arXiv research records available by October 7, 2026. They do not establish how every commercial coding agent will behave. Use them as evidence that different kinds of rule-following can fail, not as a universal product rating.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.