A transcript shows what an AI coding-agent session said; it does not show that another engineer can reproduce the change. Harper Zhu’s proposed replay ritual tests that claim directly: preserve a failing characterization test and a frozen starting commit, save the change as a patch, and use a script to apply it and run the test without the original chat or agent.
What a replay proves—and what it does not
Zhu’s framing is pointed: “A chat transcript is a memoir of intent, not a receipt that another machine can cash.” The useful evidence is not a polished account of the session but a runnable artifact that survives outside it. If someone can take the permitted files to a clean worktree or second host at the same base commit and reproduce the test result without the original session, the change has passed this particular replay check.
That is a test of reproducibility, not a guarantee of correctness. A patch can apply cleanly while changing the wrong behavior; a test can pass while failing to cover the important case. The proposed ritual therefore separates patch applicability, test behavior, and review of the resulting diff. Zhu presents it as a narrow, time-boxed spike—not as a validated engineering standard or a demonstrated improvement in outcomes.
Set up the experiment before prompting the agent
Freeze the starting point
Prepare a clean worktree at a recorded commit before the coding agent starts. Write down the commit identifier and the repository paths the agent may change. A frozen base makes it possible to distinguish the agent’s contribution from unrelated local edits or later changes.
#1 Best Overall
Write one narrow hypothesis
State the behavior you expect the change to establish, and define what result would count as failure. Use a characterization test that already exists and fails on the frozen base; otherwise, a later passing test may not demonstrate that the patch fixed a previously observable problem. Zhu’s example uses a pagination test fixture, but it is illustrative rather than a report of an actual incident.
Record the allowed artifacts and stop condition
The proposed spike folder can hold a hypothesis, clock settings, a test fixture, and the replay outputs, such as replay.sh, patch.diff, and RESULT.json. Record a source-tree fingerprint and specify what the operator will stop doing when the time box expires. The sample SPIKE_MINUTES=90 is an operator-chosen example, not a vendor or model limit, a measured recommendation, or a reported benchmark.
Rank #2
Build a replay that separates the checks
The template’s sequence is deliberately simple: enter the repository, require the patch, check that it applies, apply it, run the existing fixture with pytest, and record the outcome. Use the exact test command the project expects; a replay is less useful if another person has to infer how the test was run.
#!/bin/sh
set -eu
cd "$(git rev-parse --show-toplevel)"
test -f patch.diff
git apply --check patch.diff
git apply patch.diff
pytest -q tests/test_pagination.py
This is an illustrative template, not an executed script. Adjust the fixture path and invocation to match the project, and make sure the script can find its inputs from the location where it will be run. Git’s official git-apply documentation says --check checks whether a patch is applicable and turns off applying it. A successful check establishes only that the patch can be applied in that context; it says nothing by itself about test adequacy or correctness.
After the test runs, record its exit status and enough context to interpret the result—for example, the base commit, patch identity, command, and pass or fail status. Keep the record factual: a result file is evidence of what that run reported, not an independent review of whether the change is sound.
Run the second-worktree or second-host check
- Check out the frozen base. Use the recorded commit in a separate worktree or on a second host, not the original agent’s modified working directory.
- Copy only the permitted replay artifacts. Bring over the patch, script, fixture if it is not already in the repository, and any explicitly allowed inputs. Do not rely on the old chat, notes, unsaved editor buffers, or an unrecorded local modification.
- Run the replay without the agent. Execute the script using the documented test environment and observe whether it applies the patch and reproduces the expected test result.
- Inspect the resulting change. Review the diff separately from the test outcome. Passing tests and an applicable patch do not replace code review.
If the replay needs access to the original conversation, agent assistance, an unsaved buffer, or hidden local state, the proposed hypothesis has failed: the saved artifacts were not sufficient to reproduce the change under the stated conditions.
Rank #4
Keep the limits of the ritual visible
Zhu describes the workflow as a small countermeasure for teams that have been treating session narratives as proof. It is a poor fit for exploratory product design and incidents that require live production traffic. Where every agent patch already passes a trusted CI gate, a separate manual replay may add little.
The source also sketches a watchdog and an inventory checker, but these are unexecuted examples, not tested safeguards. A small helper script cannot be assumed to inventory every hidden dependency or guarantee that the replay is isolated. The method’s claim is narrower: it asks whether the explicitly permitted artifacts work from a known base without the original session.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsBest Value
When this is worth doing
Use the replay approach when a change is bounded, the starting state and expected behavior can be fixed in advance, and independent reproduction is more valuable than speed alone. Before starting, make sure the team can answer these questions:
- Is the starting commit recorded and held fixed?
- Does a relevant characterization test already fail on that base?
- Are permitted file paths and external inputs explicit?
- Can the replay run the same test command without the agent or its conversation?
- Will someone review the resulting diff in addition to checking the test result?
These are workflow considerations, not a validated scoring system. The point is to make the evidence portable enough that another person can check it independently.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




