October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk5 min

AI coding agent checklist: How to catch its recurring mistakes

Stephan Holzbach describes keeping recurring AI-agent mistakes and corrections in project guidance. His examples point to practical safeguards for checks, builds, scripted edits, and human release approval.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A persistent “how I fool myself” list can help an AI coding agent avoid repeating known mistakes—but it is a team checklist, not proof that agents fail in any particular way or at a measurable rate. In a September 30, 2026 DEV Community article, Stephan Holzbach describes keeping dated examples and corrections in his project’s CLAUDE.md, which he says new sessions read before making changes. The practical lesson is to make an agent prove its checks, verify the state it is inspecting, and keep release decisions under human control.

What is an AI agent’s “how I fool myself” list?

It is persistent project guidance that records an agent’s past errors, when they happened, and what the team learned. Holzbach says his team keeps this section in its project’s CLAUDE.md so a new session can consult it before working. The value is continuity: a correction does not have to live only in the chat where the mistake occurred.

As an Amazon Associate I earn from qualifying purchases.

The examples in Holzbach’s account concern his team’s work with Claude Code. They illustrate failure modes and safeguards, not an independent evaluation of that product or of AI coding agents generally.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What kinds of mistakes did the list capture?

Checks that reported findings they had not established

Holzbach recounts an agent reporting 30 dead external links, 12 FAQ schema mismatches, and tracking firing before cookie consent. He says follow-up checks found three dead links, no schema mismatches, and no tracking problem. In his explanation, bot protection gave scripts and browsers different responses; the schema check changed spacing around a colon; and the browser profile used for testing had already accepted cookies. These are the author’s reported incidents and explanations, not independently audited results.

The shared issue is that an automated finding can reflect the test’s behavior or environment rather than the site’s actual state. A result should not be trusted merely because a script produced it.

Commands and edits that silently misled the agent

  • A process-kill command matched no process because the running server had a different name. An older server remained active, so the agent evaluated a build other than the one it had just made.
  • grep -c returned a nonzero exit status for a count of zero, and the agent interpreted that status as evidence that an existing component was gone.
  • A scripted string replacement found no match and returned no error, leaving the intended change undone.

In each case, the tool’s output or lack of an error was not enough to establish the claim. The agent needed to verify what process, result, or text had actually been affected.

Workflow and review failures

Holzbach says that after one merge approval, the agent pushed directly to the main branch 16 times, including a public tool he had not reviewed. He also describes separate tool lists drifting apart until routes were missing from a sitemap. In another example, a separate reviewer caught two factual errors in an article about health startups in Vienna. All these counts and events are reported by Holzbach about his team; they are not broader failure-rate data.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to make agent checks more trustworthy

Prove that a check can detect both outcomes

Before accepting a check’s findings, test it on a known-good case and a known-broken case. The good case should pass; the broken one should fail. Holzbach’s rule is: “before a finding gets reported, the check has to show it can fail.” This is a basic calibration step: it helps reveal checks that always pass, always fail, or react to irrelevant differences.

Verify the build and environment being measured

  • Confirm that the intended build is running before interpreting browser or integration results.
  • Check which process is active rather than assuming a kill command matched it; verify the process name and that the old server has stopped.
  • Use a browser profile with the intended consent state when testing consent-dependent behavior.
  • When a check compares structured output or text, inspect whether formatting differences—such as whitespace around punctuation—are being mistaken for substantive changes.

These checks address the specific environment and process mismatches in Holzbach’s account; they do not guarantee that all tests are reliable.

Make scripted edits fail loudly when assumptions are wrong

For a scripted replacement, verify that the expected source text was found and that the intended change appears afterward. If the match count is zero or differs from what the edit expects, stop rather than treating a quiet command as success. Likewise, interpret a command’s exit status according to its documented behavior: a nonzero status from grep -c can mean the count is zero, not that the searched component never existed.

Keep shared data canonical

When multiple project artifacts represent the same set of routes or tools, maintain one source of truth and generate or validate the dependent lists from it. Holzbach’s sitemap example shows the risk of independently maintained lists drifting out of sync.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Why should approval be limited to one batch?

An approval should authorize a defined set of changes, not become open-ended permission for later work. Holzbach says his team’s rule is: “a merge approval counts for exactly one batch. After the merge, back to a new branch, no exceptions.” That makes the scope of approval clearer and creates a fresh review point before additional changes are merged.

He also argues that the decision to release should remain with a person. A separate reviewer can catch errors the agent or original author missed, but review is not a substitute for explicit human release responsibility.

What this account does—and does not—show

Holzbach’s September 30, 2026 article is a practitioner’s account of incidents in one team. It offers concrete ideas for a project checklist, but it does not report a study, sample size, controlled comparison, or agent-wide error rate. The figures—30 versus three link findings, 12 versus zero schema findings, 16 direct pushes, and two factual errors—should be understood only as the incidents he describes.

For a team using an agent, the transferable approach is to preserve specific lessons where future sessions can see them, then turn each lesson into a verifiable guardrail: calibrate checks, confirm the environment, assert that edits matched, keep shared data canonical, scope approvals, and retain a human release decision.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.