To learn distributed systems by breaking them, start with a guarantee the system claims to make, run operations that exercise it, inject a fault, and check the recorded history against that guarantee. For example, an illustrative test question is: should a write acknowledged before a node or network failure still be visible afterward? The answer depends on the system’s stated contract; the test is meaningful only when that contract is explicit.
What a failure test is meant to establish
A healthy-cluster demonstration shows that a system can work under favorable conditions. A failure test asks whether it preserves a stated property when conditions change. Jepsen describes a process of characterizing a system’s design and claims, generating operations, introducing faults, and checking the resulting history against a model. Jepsen’s analyses show how this approach is applied to real systems.
| # | Preview | Product | Price | |
|---|---|---|---|---|
| 1 |
|
Distributed Systems | $32.68 | Buy on Amazon |
| 2 |
|
Understanding Distributed Systems, Second Edition: What every developer should know about large... | $32.41 | Buy on Amazon |
| 3 |
|
Distributed Systems | $35.00 | Buy on Amazon |
| 4 |
|
Foundations of Scalable Systems: Designing Distributed Architectures | $42.49 | Buy on Amazon |
| 5 |
|
Distributed Systems: Concepts and Design | $255.63 | Buy on Amazon |
As an Amazon Associate I earn from qualifying purchases.
Think of the test as a chain: a claim defines what should happen; a workload creates operations; fault injection stresses assumptions; and a checker evaluates the observed outcomes. If any link is missing, the result can be misleading. A crash test without relevant reads and writes, for example, says little about whether acknowledged writes survive.
Define the property before choosing a fault
Turn a broad promise such as “reliable” or “consistent” into a checkable invariant. An invariant describes what must remain true across the operation history, including when requests overlap or a node becomes unreachable. The exact invariant should come from the system’s documented guarantee, not from an assumption that all distributed databases behave alike.
#1 Best Overall
- Identify the operation types to test, such as writes, reads, transactions, or lock acquisition.
- Specify what outcomes are permitted and forbidden, including what counts as an acknowledged operation.
- Decide how the checker will assess the complete history, rather than judging isolated responses.
Jepsen’s method pairs generated operations with history checking. The workload and checker therefore matter as much as the injected fault: a test only evaluates the properties its operations exercise and its model recognizes.
Build a test from operations, faults, and observed history
- Characterize the system. Record its stated guarantees, relevant configuration, version, and deployment assumptions.
- Run a workload. Issue operations that put the chosen guarantee under pressure, recording invocation and completion outcomes.
- Inject a fault. Disrupt a process, network path, clock, power source, or disk, depending on the hypothesis being tested.
- Check the history. Compare the operations and results—including concurrent activity and recovery—with the invariant or model.
- Report the scope. State the software version, environment, fault conditions, and workload so readers know what the result does and does not cover.
Increase failure complexity deliberately
Start with a single failure class, then add interactions. This makes a surprising history easier to interpret and helps distinguish a specific weakness from a compound scenario.
Rank #2
Process crash or pause
Stop or pause a process while operations are in progress, then observe what clients receive and what happens when the process returns. A crash and a pause are not interchangeable: a paused process may later resume with old state or delayed work, while a crashed process must restart.
Network partition and latency
Separate nodes or introduce delay to test behavior when communication is restricted or slow. A majority/minority split can expose different behavior on each side, but the exact outcome depends on the system’s design and the operations issued during the partition.
Rank #3
Clock errors
Skew or otherwise perturb clocks when the system’s behavior depends on timestamps or leases. Observe both the operation history and any recovery behavior; a clock-related test should not be generalized to systems or configurations that were not evaluated.
Power and disk failures
Test power loss or disk errors when durability is part of the claim. These conditions differ from a process crash because they can affect persisted state. A successful process-restart test alone does not establish how the system behaves after storage damage or abrupt power loss.
Overlapping failures
Once individual cases are understood, combine faults—for example, a node pause during a partition, or a restart while other nodes are recovering. Compound tests can expose interactions that isolated tests miss, but their results are harder to attribute and should be described with precise conditions.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Separate safety, availability, and recovery
Do not collapse every outcome into “the system worked” or “the system failed.” These are different questions:
Best Value
- Safety: Did the observed history violate the specified invariant, such as an acknowledged write disappearing or a read returning an impermissible value?
- Availability: Did operations complete during the fault, and which clients or nodes could make progress?
- Recovery: After the fault ended, did the system return to service, and what state did clients observe?
A system can preserve safety by refusing or delaying operations, so availability must be assessed separately. Likewise, successful recovery does not erase an invalid history that occurred during the fault.
Read test results as bounded evidence
Jepsen describes its tests as opaque-box tests of real systems: they can uncover implementation behavior and bugs, but they are nondeterministic and cannot prove correctness. Its ethics discussion also notes limits from bounded search and the possibility of harness errors. Jepsen’s ethics page explains those constraints.
Every result is tied to what was actually tested: a particular release, setup, workload, fault schedule, and checker. For example, Jepsen’s Capela analysis describes tests on three-to-five-node Debian clusters and specifies the versions and failure conditions evaluated. That scope should not be stretched into a claim about all Capela releases or deployments. The Capela analysis is an example of why test context belongs beside the finding.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsEmpirical tests sample executions; they do not cover every possible schedule or future implementation. They complement design review, formal reasoning, and other verification methods by showing how a concrete implementation behaves under selected conditions. Jepsen’s stated aim is: “We want to teach everyone how to analyze their own systems, and for the industry as a whole to produce software which is resilient to common failure modes.” Jepsen also lists talks, training, and consulting for readers seeking further instruction.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




