Validate a learned quantum state by predicting the outcomes of the experiment’s measurement settings and comparing those predictions with data that were not used to fit the model. Then check whether the inferred state is physically valid, whether the measurements can identify the claimed state, and whether the result is stable under plausible statistical and experimental uncertainties. A close fit alone does not establish that the state is unique or that the measurement model is correct.
Start by defining what the learner predicts
Before calculating a score, write down the experiment and model being evaluated. Record the measurement settings, observed counts or expectation values, shot counts where applicable, calibration assumptions, and every preprocessing step. State whether the learner outputs a density matrix, outcome probabilities, or predicted expectation values; these are related representations, but they are not interchangeable validation targets.
Also identify which observations were used to train or tune the learner. Re-scoring the same observations measures in-sample fit, not independent predictive performance. If possible, reserve measurement settings or data for evaluation before training. If the experiment is too small to hold data out, disclose that limitation rather than describing the fit as an independent test.
- Measurement model: Specify the measurement operators and any assumptions about state preparation and measurement calibration.
- Data: Report counts or measured expectations, the number of shots when known, and how uncertainty is represented.
- Learning constraints: Record whether the model enforces positivity, trace one, purity, rank, or another structural prior.
- Evaluation data: Mark which measurements were used for fitting, model selection, and final evaluation.
Predict the observations and test the fit
For each measured setting, use the learned state and the stated measurement operators to calculate the outcome probabilities or expected values. For a density matrix ρ and an outcome operator E, the predicted probability is p = Tr(ρE). Compare those predictions with the corresponding observed frequencies or expectation values—not with a different quantity merely because it is easier to compute.
#1 Best Overall
Choose a comparison that matches the data-generating noise model. For outcome counts, a likelihood based on the predicted probabilities can assess how plausible the observed counts are under the model. For measured expectation values, residuals can show the size and pattern of prediction errors; their interpretation should account for the uncertainty of each measurement. A single aggregate score can conceal a setting that is badly predicted, so inspect setting-by-setting discrepancies as well.
Set the acceptance criterion before interpreting the result and explain how it was chosen. There is no universal cutoff in the cited evidence: an appropriate threshold depends on the measurement design, finite-sample noise, calibration, and purpose of the validation. The 2019 four-qubit NMR study describes predicting local measurements from its learned state and comparing them with measured values against an acceptable error bound. That is a useful validation pattern, not a threshold transferable to other experiments.
Check whether the estimated state is physically valid
If the learner returns a density matrix, check Hermiticity, unit trace, and positive semidefiniteness. These are separate from agreement with data: a matrix can fit measured values while failing to be a valid quantum state. Conversely, a physically valid matrix can still predict the observations poorly.
Be explicit about constraints used during learning. In a two-photon experiment, imposing physical-state constraints improved reconstruction quality under noise, but assuming purity without justification could bias the estimate. The authors of the 2020 Physical Review A paper Neural-network quantum state tomography in a two-qubit experiment cautioned: “Including additional, possibly unjustified, constraints, such as assuming pure states, facilitates learning, but also biases the estimator.” A constraint is a modeling assumption, not evidence that the real state satisfies it.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteRank #2
Take care when computing fidelity. A raw matrix from linear inversion can fail positivity, so it may not be a physical density matrix suitable for a fidelity comparison. State how any nonphysical estimate was handled and distinguish fidelity to a trusted target from a fit to measured data.
Ask whether the measurements identify the state
A good fit cannot show that the learned state is the only state compatible with the data. Determine whether the measurement design is informationally complete for the state you claim to reconstruct. If it is incomplete, the observations may permit multiple states; the learner may select one because of its architecture, training procedure, or prior constraints rather than because the data uniquely determine it.
When uniqueness is not established, describe the result as a state consistent with the measurements under the stated model. Where feasible, report bounds over compatible states for the property that matters, or explain how the estimate depends on the restricted model class. A 2018 Physical Review A paper by Adam C. Keith, Charles H. Baldwin, Scott C. Glancy, and Emanuel H. Knill on joint state-and-measurement tomography notes that some procedures do not enable unique state estimation. That limitation should shape the claim, not be hidden by a low residual.
Use the data to check experimental stability
Validation also depends on whether the experiment behaved as assumed. State preparation or measurement can drift, and a learner can absorb those effects into an apparently good state estimate. Examine the tomography data for instability rather than treating every observation as a sample from one unchanged process.
Cross-validated tomography was proposed as a way to test assumptions about preparation and measurement using data already collected. Overcomplete measurement designs—with more distinct measurements than the minimum needed for reconstruction—are generally easier to validate this way than minimal designs, because they provide additional consistency checks. A stability check does not automatically diagnose or correct every calibration error; report what was tested and what remained assumed.
Choose an independent comparison that fits the experiment
When a trusted target is available, compare the learned state with it using a clearly defined fidelity or other task-relevant measure. Synthetic data with a known generating state provide one kind of check; a separately reconstructed experimental reference or held-out measurement settings provide another. Explain how the reference was obtained, since a reference built from the same unexamined assumptions is not fully independent.
Methods answer different questions, so choose based on the experiment rather than treating them as interchangeable:
| Validation approach | What it can check | Main limitation to report |
|---|---|---|
| Held-out measurement settings or observations | Whether the learned model predicts data not used in fitting | Requires evaluation data that are meaningfully separate from training and model selection |
| Fidelity to a known target or separate reference | Closeness to a target state when a trustworthy reference exists | Depends on the reference being valid and sufficiently independent |
| Cross-validated tomography | Consistency with assumptions about preparation and measurement; potential instability or drift | Validation is easier with overcomplete than with minimal measurement designs |
| Joint state-and-measurement estimation | Coupled uncertainty in the state and the measurement apparatus | May leave non-unique state estimates when measurements are incomplete |
| Direct fidelity-learning methods | Can reduce measurement requirements for the fidelity task | Results depend on the method’s trained domain and calibration |
Interpret published performance figures narrowly
Published numbers illustrate results under particular conditions; they are not general acceptance thresholds. In a 2019 four-qubit NMR experiment with 20 experimental instances, the authors reported 98.8% average fidelity between learned reconstructions and experimental tomography states, and 98.7% average test-set fidelity for their four-qubit neural-network estimates. Their reported seven-qubit simulated case had 97.9% average test-set fidelity under that paper’s generated data and assumptions. None of these figures predicts the accuracy of a different apparatus, learner, or measurement design.
Free tools Windows power users keep installed
One-click scans. No signup required.
Rank #4
A 2020 experimental neural-network tomography paper reported average reconstruction-fidelity enhancements of 10% and 27% relative to two specified alternatives. Those comparisons describe that paper’s protocol and baselines; they are not a general performance promise. In every case, identify what the metric compared, the data and setup involved, and whether it was a test result or a comparison against a particular reference.
Report enough detail for someone else to assess the claim
A validation report should make it possible to distinguish predictive agreement from physical validity and uniqueness. Include:
- The measurement settings, outcomes, count or shot information, and calibration assumptions.
- The learner’s output form, preprocessing, training/evaluation split, and all imposed constraints.
- The predicted-versus-observed comparison, statistical model, uncertainty method, and acceptance bound chosen before interpretation.
- Physicality checks and how any nonphysical linear-inversion estimate was treated.
- Whether the measurement design supports unique identification of the claimed state; if not, describe compatible-state bounds or model dependence.
- Stability checks for preparation and measurement, including limitations of the available design.
- Any reference state used, how it was obtained, and how independent it is from the learned estimate.
The defensible conclusion is specific: the learned state predicts these measured data within this stated criterion, under these assumptions, with these limits on physicality, stability, and identifiability. Do not turn that conclusion into a claim of universal accuracy or unique reconstruction unless the experiment supports it.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problems




