Yes. In Ertuğrul Mutlu’s 2026 preprint, small convolutional neural networks reached similar predictive performance after shared training while still showing measurable differences in selected internal representations. The result is evidence that training order can remain detectable in this particular MNIST-based setup—not proof that neural networks generally retain permanent memories of their training history.
What does “behavioral convergence” mean here?
Mutlu’s preprint uses “behavioral convergence” operationally: two networks satisfy the study’s predeclared criterion for matched predictive performance. It does not mean they make identical predictions on every possible input or that their functions are interchangeable in every setting.
As an Amazon Associate I earn from qualifying purchases.
The central question is whether matching behavior after a shared training phase also means the networks have similar internal representations. The reported results indicate that these are separate questions: similar accuracy does not, by itself, establish identical internal representations.
How did the experiment test training-history dependence?
Two different task orders, then a common training phase
The repository describes the principal split as MNIST digits 0–4 (task A) and digits 5–9 (task B). The paired convolutional networks began from identical initial weights. One trained on A and then B; the other trained on B and then A. Both then trained on a common, balanced 0–9 distribution (task C).
#1 Best Overall
During this common-relaxation phase, the pair received the same deterministic batch sequence and checkpoint schedule. That design makes training order the intended contrast: after the differing histories, did a shared later experience erase measurable differences?
How the study measured representations
The repository’s primary representation-history summary is H_repr = 1 - mean(CKA_conv2, CKA_fc1). CKA, or centered kernel alignment, is a way to compare representations from selected network layers. In this summary, a larger score means lower similarity under the chosen comparisons. The primary score uses Conv2 and fc1; it excludes logits and Conv1.
Rank #2
This is a specific measurement, not a complete test of model identity or function. Mutlu’s “representational convergence” refers to similarity under the selected layers and CKA-based summary, not an absolute state in which all internal features must match.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsWhat did Mutlu report?
The figures below are from the preprint abstract and the author’s repository. They describe the tested CNN/MNIST-derived protocols, not neural networks as a whole.
Rank #3
| Experiment or result | Reported finding | What it supports |
|---|---|---|
| Main paired-run experiment | 16 of 20 paired runs met the behavioral-matching criterion. Mean representation-history score: 0.139 (95% bootstrap CI 0.127–0.153); prediction disagreement: about 3.1%. | Matched predictive performance under the study’s criterion coexisted with measurable differences under its selected representation comparisons. |
| Long common-relaxation stress test | After 50,000 common optimizer updates, five paired seeds had a mean representation-history score of 0.190 (95% bootstrap CI 0.161–0.219) and a mean accuracy gap of 0.18 percentage points. | A measurable residue remained over this tested horizon. Five pairs do not establish that it would persist indefinitely. |
| Same-label rotated-MNIST control | Across five paired seeds, behavioral matching was reached while the mean representation-history score was 0.162. | The reported pattern also appeared in this control, which changes the input domain while retaining the labels. |
| Matched-learning-rate activation control | A ReLU/LeakyReLU control reduced the 50,000-update representation residue by about 0.040 across five paired seeds. | This is directional evidence that activation-mediated plasticity may contribute; it does not establish a causal mechanism. |
| Fresh linear probes | The abstract reports practically equivalent linearly accessible class information with sufficient labeled data. The repository specifies a ±0.5 percentage-point equivalence margin for the 500-examples-per-class endpoint. | Linear readout performance can be practically equivalent without the measured representations being identical. This does not rule out differences in low-data readout settings. |
Mutlu reports these as results from the study’s protocols; see the arXiv preprint and the reproducibility repository.
Do different representations mean one network learned less?
Not on the evidence reported. Fresh linear probes with sufficient labeled examples showed practically equivalent linearly accessible class information between the two histories. That narrows the interpretation: the measured representational differences did not translate into a reported practical disadvantage for that readout and data regime.
Rank #4
It does not show that every downstream task will be unaffected. The probe result concerns linear readouts with sufficient labeled data, including the repository’s stated endpoint of 500 examples per class and its ±0.5 percentage-point equivalence margin. It does not settle performance with very little labeled data, nonlinear readouts, other tasks, or other ways of comparing representations.
What can—and cannot—be concluded about training history
What the result supports
- Predictive similarity and internal representational similarity are distinct measurements.
- In these paired experiments, reversing task order before a shared training phase left a measurable difference under the selected CKA comparisons.
- The long-horizon result shows a residue after 50,000 common updates in five pairs, alongside a small mean accuracy gap.
What it does not establish
- It does not show that all architectures, datasets, or training procedures retain such differences.
- It does not prove permanent memory or persistence beyond the tested training horizon.
- It does not identify a causal mechanism. The activation-function control is suggestive, not decisive.
- It does not prove that the models occupy disconnected optimization basins. The repository cautions that its raw weight interpolation was not permutation-aligned, so a linear barrier is not sufficient evidence of full basin disconnection.
The repository also notes that AB-versus-BA effects can overlap with catastrophic forgetting and ordinary last-task effects. The available results therefore do not isolate a universal form of “hysteresis” from every alternative explanation.
Best Value
How to judge whether the result generalizes
The evidence centers on a small CNN and MNIST-derived protocols. The sources do not establish independent replication or results for transformers or large models. To assess generality, further comparisons would need to vary several parts of the design:
- Architecture and scale: test whether the effect holds beyond the reported small convolutional network.
- Dataset and task history: distinguish histories that change labels from those that change the input domain.
- Shared-training schedule: vary the duration and schedule of the common phase.
- Representation measurement: compare other layers and similarity metrics, not just Conv2 and fc1 under the reported CKA summary.
- Downstream readout: test different readout methods and amounts of labeled data.
- Replication and uncertainty: use more seeds and report uncertainty across runs.
Where to inspect the paper and reproduction materials
The arXiv record lists Ertuğrul Mutlu’s preprint as submitted on 29 September 2026, version 1, in Computer Science / Machine Learning. The author links code and reproducibility artifacts. The repository describes a Python virtual-environment setup, dependency installation from requirements.txt, paired training configurations, and validation; training downloads MNIST if it is not already present.
The repository notes that hardware, PyTorch, and CUDA differences can affect reproduction and that environment metadata is recorded when available. It presents generated experiment outputs as the underlying source of truth, so reproductions should be interpreted alongside their recorded environment and outputs.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




