Freeze an intermittent agent-patch test only for the specific invariant and fixture bytes behind its recorded outcomes—not indefinitely by its test name. Finley Zhou’s proposed protocol pairs a stable property ID with a SHA-256 fixture digest and a record of independent reruns, so fixture edits invalidate old freeze evidence. It is a practitioner proposal, not an independently validated standard or proof that a patch is correct.
What a test freeze applies to
A test’s display name identifies implementation, not necessarily the behavior or inputs that past runs exercised. Zhou’s proposal instead identifies the thing being evaluated with three fields: a stable property_id for the invariant, a fixture_digest for the bytes that property actually read, and an evidence_window of independent executions with pass and failure counts.
As an Amazon Associate I earn from qualifying purchases.
This distinction matters when an agent patch changes a fixture without changing a test function name. The new bytes produce a different digest, so evidence gathered against the previous input should no longer authorize a freeze. If the property itself is renamed or changed, Zhou’s proposal treats it as a new property with no inherited evidence.
How the gate classifies outcomes
The proposed decision logic separates residual intermittency from stable defects, fixture drift, and an unreliable runner. A matching digest is necessary for prior evidence to apply, but the outcome pattern determines what the gate should do:
| Current property and fixture result | Proposed action |
|---|---|
| Matching digest; property holds on every recorded run | Merge-ok for that property. |
| Matching digest; outcomes are mixed and a failure signature recurs | Eligible to become a freeze candidate after the evidence window; candidacy does not itself skip the test. |
| Matching digest; the same violation occurs every time | Block as a stable violation, not a flake. |
| Matching digest; failure signatures vary widely | Do not freeze; investigate runner isolation or shared state. |
| Digest differs from the recorded digest | Classify as fixture drift and discard the old freeze, regardless of test name. |
| Catalogued fixture is missing | Block because the property catalog is broken. |
A freeze is therefore narrow technical debt for residual timing noise, attached to an owner and a precise input identity—not a permanent exemption for a test label.
Workflow for collecting freeze evidence
- Define the invariant. Create the stable property ID before evaluating the patch. Keep it tied to the behavior being asserted rather than a pytest node ID or function name.
- Lock the inputs. Identify the fixture bytes the property consumes and compute their SHA-256 digest. Re-hash on each patch so edits are detected.
- Run an evidence window. Execute the property independently, recording pass/fail counts and failure signatures. Zhou suggests seven isolated runs as a starting budget, not as a validated statistical threshold.
- Classify before recording. Use the outcome pattern to distinguish a recurring intermittent signature from a stable violation, divergent runner failures, fixture drift, or a broken catalog.
- Write the ledger entry. The proposed JSONL record includes the property ID, fixture path and digest, run/pass/fail counts, failure signatures, status, and reason. The article’s sample values are illustrative, not observed production results.
- Apply only a matching freeze. A proposed pytest collection hook skips only when a ledger record is marked
frozenand both the property ID and current digest match. Incomplete records should block or run, not silently skip.
Zhou positions these repeated property checks in a separate, inexpensive worker lane rather than the integration-test pool. The lane sits alongside the full suite before merge; it does not replace that suite. The example classifier uses subprocesses, hashes fixtures, counts outcomes, and emits a ledger entry. The hook and classifier are proposed examples, not a published plugin or independently production-tested tools.
What the digest does—and does not—establish
The digest guards against applying old evidence to different fixture bytes. It does not show that the fixture represents every relevant production condition, prove that an agent patch is correct, or guarantee detection of regressions. As Zhou puts it, “N independent runs are not a confidence interval.” Seven runs are only a starting budget; one failure in seven may still signal a real race.
The proposal assumes byte-stable fixture inputs and properties deterministic on those inputs. It is not designed to measure performance, network retries, or UI flakiness. Generated timestamps can make digests change on every run; shared clocks, live clocks, and unordered network mocks can also produce divergent failure signatures. The sample uses subprocess isolation, not container isolation, so native extensions or shared temporary directories may need stronger boundaries. Teams that cannot run isolated subprocesses may get little value from this approach.
- Do not use a frozen property as an oracle for security-sensitive or money-handling paths.
- Do not let a freeze conceal a changed I/O contract.
- Investigate diverse failure signatures as a possible isolation or shared-state problem instead of treating them as ordinary flakiness.
Source and status of the proposal
Finley Zhou described this workflow in a DEV Community article published September 3, 2026. The article presents a worked proposal with Python, pytest, JSON fixtures, SHA-256, subprocesses, and a JSONL ledger; it does not report independent validation or adoption as a standard. Its central rule is concise: “A flaky freeze is valid for one fixture digest only.”
Quick Recap
Best Value
Rank #4
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




