Use a control chart to see whether repeated performance-test results remain consistent or show a statistical signal that warrants investigation. Choose a meaningful measure, collect comparable observations in time order, establish limits from a historical baseline, and select a chart suited to the data. A signal is not a diagnosis—and statistical stability does not mean the result meets a performance target.
What a control chart tells you
A control chart plots measurements in time or sample order against a center line and upper and lower control limits. The limits describe the observed process behavior when it is stable; a point beyond a limit or a nonrandom pattern can indicate that something changed. NIST describes the chart as a tool for assessing process stability, and identifies execution time as one software activity to which control charts can be applied.
For performance testing, a chart can help distinguish ordinary variation between runs from a result that merits investigation. It does not tell you whether a change came from the application, workload, environment, test procedure, or measurement system.
Choose the performance measure and define one observation
Start with the operational question. NIST’s NML performance-testing documentation gives examples including maximum and average read/write time, average CPU time for a read/write operation, throughput (new messages received per second in that program’s context), and latency (average time between a write returning and the corresponding message being received by a read). These are examples from a particular NML testing context, not a universal list or prescription.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
Decide what each plotted point represents before collecting data: one test run, or a summary of a subgroup of runs. Keep the measure and test conditions sufficiently comparable for the sequence to represent the process you want to monitor. Record relevant context such as workload, software build, hardware or cloud environment, and test procedure alongside each result. Preserve chronological order; sorting results by value destroys the time sequence the chart is meant to examine.
Measurement quality matters. NIST notes that clock resolution can affect maximum-time measurements. If the instrument cannot reliably distinguish the changes you care about, the chart cannot recover that information.
Establish a baseline before monitoring
NIST describes control-chart use in two phases. In Phase I, use historical observations to calculate initial limits and investigate points outside them for assignable causes. If a result reflects a documented, understood cause and the process is reassessed, limits may be recalculated. Once the baseline is judged suitable, carry those limits forward into Phase II monitoring.
Rank #2
- Used Book in Good Condition
- Gather historical results: use observations produced by a sufficiently consistent test and measurement process.
- Calculate initial limits: select the chart appropriate to the data and collection design.
- Investigate unusual points or patterns: examine changes in the build, workload, environment, instrumentation, and procedure; do not remove data merely because it is inconvenient.
- Document justified changes: record any assignable cause, action taken, and reason for revising the baseline.
- Monitor prospectively: apply the established limits to new comparable observations in Phase II.
Do not silently reset limits after an unfavorable result. A baseline change should reflect a real, documented change in the process, not an attempt to make a regression disappear.
Choose a chart that fits the data
| Observation design or goal | Chart family to consider | What it monitors |
|---|---|---|
| Continuous measurements collected in subgroups | X-bar chart, commonly paired with an R or S chart | X-bar monitors subgroup means; R or S monitors within-subgroup variation. |
| Continuous individual observations without subgroups | Moving average, moving range, or moving standard deviation chart | Individual-result behavior and variation over a moving sequence. |
| Relatively small shifts in a process mean matter | CUSUM or EWMA | Methods developed to detect small shifts in location. |
| Proportions or counts | P/NP or C/U chart, depending on the count setup | Proportion/count data under the corresponding binomial or Poisson setup. |
These are selection cues, not an automatic prescription. Match the chart to whether observations are grouped, the measurement’s statistical form, and the size and kind of change that matters. NIST’s general documentation notes that approximate normality is assumed for several standard charts for continuous data; skewed or discrete measurements require appropriate chart selection and assumption checks.
Keep unlike units—such as latency and CPU time—on separate ordinary univariate charts. NIST distinguishes univariate from multivariate charts; combining unrelated measures on one univariate plot makes its limits difficult to interpret.
Rank #3
Read signals without overclaiming
A point above the upper limit or below the lower limit is a reason to investigate, not proof of a particular cause. A run or other systematic, nonrandom pattern can also be a signal even when every point lies between the limits. NIST characterizes an in-control process as having points within limits and a random pattern.
Signals entail a false-alarm tradeoff. For a normal process observed with a Shewhart X-bar chart using three-sigma limits, NIST’s handbook gives an illustrative per-point probability of 0.0027 outside the limits and an average run length of about 371 points before a false alarm if the process has not changed. This figure applies to those stated conditions; it is not a guaranteed false-alarm rate for every performance chart. Additional run rules can alter both detection and false-alarm behavior.
Recommended Free Tools
When a signal appears, check the test record and investigate plausible changes in software, workload, environment, instrumentation, or procedure. Log what you find and any corrective action. Recalculate limits only when the process has materially changed and the new baseline is justified.
Rank #4
Control limits are not performance requirements
Control limits estimate the behavior of a process; specification limits, service-level objectives, and engineering acceptance criteria state what is acceptable. A stable process can consistently miss a latency target, while a process that usually meets its target can still be unstable. Use the chart to assess stability and compare measurements separately with the relevant performance requirement.
Practical workflow
- State the question. For example: “Has the median run-level response time shifted since the last release?” Choose a measure that answers it.
- Make the test repeatable. Define workload, environment, software version, warm-up and measurement procedure, and the point represented by each observation. These are implementation choices; the appropriate design depends on the system under test.
- Log results in order. Preserve timestamps or sequence numbers and the context needed to interpret each run.
- Perform Phase I analysis. Establish initial limits from historical data, investigate unusual points and patterns, and document any justified baseline revision.
- Select the chart family. Use subgroup charts for subgrouped continuous values, a moving chart for individual continuous values, CUSUM/EWMA where small shifts are important, or a suitable count/proportion chart for count data.
- Monitor new comparable results. Plot each result chronologically against the Phase II limits and inspect both limit crossings and nonrandom sequences.
- Investigate and report. Record the signal, contextual checks, cause if established, and corrective action. Assess target compliance separately.
Capture browser-based performance evidence
If your performance test uses a browser, the timing and capture procedure should be consistent: the same page state, wait condition, viewport, and relevant browser settings help make repeated observations comparable. A screenshot is useful as visual evidence of what loaded, but it is not a substitute for recording the performance metric and test context used in the chart.
Or skip the browser setup
For screenshot evidence in an automated workflow, ScreenshotNeo provides a website screenshot API and MCP server. A GET request can return a screenshot or PDF; the API accepts parameters used by other screenshot APIs as well. Example using cURL:
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteBest Value
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for parameters and response details. ScreenshotNeo accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf tools for AI agents and MCP clients.
The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Sign up for ScreenshotNeo and start with 1,000 free screenshots a month, no card required.
Troubleshooting chart results
- Limits change sharply between baselines: check whether the test conditions or measurement system changed, and confirm Phase I observations represent a consistent process before adopting new limits.
- Every run looks noisy: verify that the workload and environment are comparable, the observation definition is consistent, and the measurement resolution is adequate. Do not interpret measurement noise as a software regression without checking the measurement process.
- All points are inside the limits, but performance seems to drift: inspect chronological patterns and applicable run rules; limit crossings are not the only potential signal.
- The chart is stable, but the target is missed: treat this as a performance-acceptance issue, not evidence that control limits should be moved.
- Data are skewed or discrete: do not assume a standard continuous-data chart is appropriate; check its assumptions and use a chart suited to the measurement form.
Frequently Asked Questions
Does a control chart replace a performance test’s pass/fail threshold?
No. It assesses process stability; evaluate pass/fail or target compliance against the separately defined requirement.
Free tools Windows power users keep installed
One-click scans. No signup required.
Can I chart performance results from different workloads together?
Only if those conditions represent the same process for the question you are monitoring. Otherwise, separate or explicitly account for the differing test conditions rather than treating unlike observations as one comparable sequence.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




