The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →When unfamiliar code has unclear behavior, first write tests that record what it actually does for selected inputs. Check that those tests fail when behavior is deliberately changed, then make the smallest scoped edit and review what moved. Characterization tests preserve existing behavior—including bugs—so they are a baseline, not proof that the behavior is correct.
What characterization tests are for
Characterization tests make a system’s existing, observable behavior explicit before you change it. They are especially useful when documentation is incomplete, tests are missing, or callers depend on behavior nobody intended to specify.
The goal is not to bless every current result. It is to create a record against which you can distinguish behavior that should remain stable from behavior you intend to change. As Dakota Huang puts it in “Characterization Tests First, Then the Smallest Safe Change”, “A snapshot is not a truth claim.”
How to characterize behavior before editing
1. Turn the ticket into an observable question
Replace a vague request such as “clean up billing” with a question a test can answer: for this input, which result is returned, which exception is raised, and under what condition? Start with the narrow behavior connected to the change rather than attempting to specify an entire system at once.
2. Identify inputs that can change the result
Trace the relevant code and its callers. Note explicit inputs as well as hidden ones such as the system clock, environment variables, network responses, random seeds, and thread scheduling. Control those inputs where practical; otherwise, repeated test runs may observe different behavior.
Huang’s Python example freezes the system date and the PLAN environment variable. It patches names where the code under test looks them up, and notes that a different import style calls for a different patch point. That is a Python-specific example, not a universal mocking rule: use the conventions of the language and test framework in your own code.
3. Record representative results, then inspect them
Run tests with representative inputs and save the observed output as a snapshot or explicit expectation. Review that output before accepting it. A generated result can capture an accidental assumption just as easily as a real behavior.
Choose cases that reveal meaningful distinctions in the code, such as ordinary inputs and relevant boundaries. Avoid both extremes: a single happy-path example may miss important behavior, while an enormous snapshot can be difficult to understand and brittle to harmless changes. Exact comparisons also deserve care for values such as floating-point numbers, where tiny representation differences may not signal a meaningful behavior change.
4. Pin important branches and errors
Snapshots are not always the clearest way to specify an important path. Add direct assertions for errors, boundary conditions, and branches that matter to callers. In Huang’s billing example, one separate assertion covers an unknown-plan error, including a case where the input rows are empty but the plan lookup still happens. That case is useful because it captures an observable ordering detail—not because every billing system should behave that way.
Choose edge cases by inspecting the code and its callers. Do not copy another project’s examples without checking whether those branches exist in yours.
5. Check that the test harness can detect a change
A passing test suite is not persuasive if it would also pass after the behavior under test changed. Make a deliberate, temporary mutation in a local copy or controlled test setup—for example, change a clamp so negative days are handled incorrectly—and verify that the relevant test fails. Restore the original code afterward.
This is a practical check that the test can notice that particular change. It does not prove that every meaningful change will be detected, or that the suite covers the whole system.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Make one small change and evaluate what moved
Once the baseline is useful, change one focused thing at a time. Run the relevant tests and examine differences in outputs, exceptions, and other observable effects. The key question is whether each difference was intended.
Rank #4
Martin Fowler describes refactoring as a controlled sequence of small, behavior-preserving transformations. On the second-edition page for Refactoring: Improving the Design of Existing Code, he writes, “By doing them in small steps you reduce the risk of introducing errors.” That principle applies to refactoring; it should not be confused with an intentional bug fix, which changes behavior by design.
When the edit is a refactor
For a behavior-preserving change, keep the selected inputs producing the same relevant outputs and errors. A rename or helper extraction should not silently alter the behavior your tests have recorded. If a test changes, investigate whether the change is actually necessary or exposes a defect in the edit.
When the edit intentionally changes behavior
Write or update an expectation for the new behavior, and identify which old assertion is being replaced. Keep unrelated observations pinned. In Huang’s example, changing a strict dictionary lookup to use a fallback changes what happens for an unknown plan; the error expectation is deliberately replaced rather than treated as an unexplained test failure.
Best Value
Huang offers a change ladder ranging from local renames and guard changes through helper extraction, behavior changes, module moves, and rewrites. Treat it as an author’s heuristic for thinking about scope, not a formal standard or a universal rule about how many lines an edit may contain.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.When to narrow the task or stop
Characterization only works when you can reproduce and observe the behavior you are trying to preserve. If important inputs cannot be controlled, reduce the test’s scope to the part you can reliably observe or defer the change until you can establish a dependable baseline.
- Clock or environment: freeze the relevant value or set it explicitly when the code allows it.
- Network calls: use a controlled response if the behavior under test does not require a live service.
- Randomness: fix a seed or inject a predictable source when feasible.
- Concurrency: avoid claiming stable characterization if thread interleaving makes the observed result irreproducible.
- Unclear expected behavior: separate recording what happens from deciding what ought to happen; consult the responsible product or domain owner when that decision is needed.
If you cannot make the observation reproducible, do not present a flaky snapshot as a reliable safety net. Narrow the change, improve control over the relevant inputs, or postpone the edit.
Further reading on unfamiliar legacy code
Michael Feathers’s Working Effectively with Legacy Code is a related resource on testing and changing existing systems. O’Reilly lists the book as published in September 2004, with Pearson as publisher and ISBN 0131177052. It is useful background on legacy-code techniques; that connection does not mean this exact workflow originates in Feathers’s book.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




