Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run Scan×
Skip to content
World desk5 min

Can Training Order Shape Neural Network Representations?

A 2026 preprint reports that small CNNs trained in opposite task orders could match behavior after shared MNIST training while retaining measurable differences under selected CKA comparisons.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Yes. In Ertuğrul Mutlu’s 2026 preprint, small convolutional neural networks reached similar predictive performance after shared training while still showing measurable differences in selected internal representations. The result is evidence that training order can remain detectable in this particular MNIST-based setup—not proof that neural networks generally retain permanent memories of their training history.

What does “behavioral convergence” mean here?

Mutlu’s preprint uses “behavioral convergence” operationally: two networks satisfy the study’s predeclared criterion for matched predictive performance. It does not mean they make identical predictions on every possible input or that their functions are interchangeable in every setting.

As an Amazon Associate I earn from qualifying purchases.

The central question is whether matching behavior after a shared training phase also means the networks have similar internal representations. The reported results indicate that these are separate questions: similar accuracy does not, by itself, establish identical internal representations.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How did the experiment test training-history dependence?

Two different task orders, then a common training phase

The repository describes the principal split as MNIST digits 0–4 (task A) and digits 5–9 (task B). The paired convolutional networks began from identical initial weights. One trained on A and then B; the other trained on B and then A. Both then trained on a common, balanced 0–9 distribution (task C).

During this common-relaxation phase, the pair received the same deterministic batch sequence and checkpoint schedule. That design makes training order the intended contrast: after the differing histories, did a shared later experience erase measurable differences?

How the study measured representations

The repository’s primary representation-history summary is H_repr = 1 - mean(CKA_conv2, CKA_fc1). CKA, or centered kernel alignment, is a way to compare representations from selected network layers. In this summary, a larger score means lower similarity under the chosen comparisons. The primary score uses Conv2 and fc1; it excludes logits and Conv1.

This is a specific measurement, not a complete test of model identity or function. Mutlu’s “representational convergence” refers to similarity under the selected layers and CKA-based summary, not an absolute state in which all internal features must match.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What did Mutlu report?

The figures below are from the preprint abstract and the author’s repository. They describe the tested CNN/MNIST-derived protocols, not neural networks as a whole.

Experiment or result Reported finding What it supports
Main paired-run experiment 16 of 20 paired runs met the behavioral-matching criterion. Mean representation-history score: 0.139 (95% bootstrap CI 0.127–0.153); prediction disagreement: about 3.1%. Matched predictive performance under the study’s criterion coexisted with measurable differences under its selected representation comparisons.
Long common-relaxation stress test After 50,000 common optimizer updates, five paired seeds had a mean representation-history score of 0.190 (95% bootstrap CI 0.161–0.219) and a mean accuracy gap of 0.18 percentage points. A measurable residue remained over this tested horizon. Five pairs do not establish that it would persist indefinitely.
Same-label rotated-MNIST control Across five paired seeds, behavioral matching was reached while the mean representation-history score was 0.162. The reported pattern also appeared in this control, which changes the input domain while retaining the labels.
Matched-learning-rate activation control A ReLU/LeakyReLU control reduced the 50,000-update representation residue by about 0.040 across five paired seeds. This is directional evidence that activation-mediated plasticity may contribute; it does not establish a causal mechanism.
Fresh linear probes The abstract reports practically equivalent linearly accessible class information with sufficient labeled data. The repository specifies a ±0.5 percentage-point equivalence margin for the 500-examples-per-class endpoint. Linear readout performance can be practically equivalent without the measured representations being identical. This does not rule out differences in low-data readout settings.

Mutlu reports these as results from the study’s protocols; see the arXiv preprint and the reproducibility repository.

Do different representations mean one network learned less?

Not on the evidence reported. Fresh linear probes with sufficient labeled examples showed practically equivalent linearly accessible class information between the two histories. That narrows the interpretation: the measured representational differences did not translate into a reported practical disadvantage for that readout and data regime.

It does not show that every downstream task will be unaffected. The probe result concerns linear readouts with sufficient labeled data, including the repository’s stated endpoint of 500 examples per class and its ±0.5 percentage-point equivalence margin. It does not settle performance with very little labeled data, nonlinear readouts, other tasks, or other ways of comparing representations.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What can—and cannot—be concluded about training history

What the result supports

  • Predictive similarity and internal representational similarity are distinct measurements.
  • In these paired experiments, reversing task order before a shared training phase left a measurable difference under the selected CKA comparisons.
  • The long-horizon result shows a residue after 50,000 common updates in five pairs, alongside a small mean accuracy gap.

What it does not establish

  • It does not show that all architectures, datasets, or training procedures retain such differences.
  • It does not prove permanent memory or persistence beyond the tested training horizon.
  • It does not identify a causal mechanism. The activation-function control is suggestive, not decisive.
  • It does not prove that the models occupy disconnected optimization basins. The repository cautions that its raw weight interpolation was not permutation-aligned, so a linear barrier is not sufficient evidence of full basin disconnection.

The repository also notes that AB-versus-BA effects can overlap with catastrophic forgetting and ordinary last-task effects. The available results therefore do not isolate a universal form of “hysteresis” from every alternative explanation.

Best Value
Sale
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to judge whether the result generalizes

The evidence centers on a small CNN and MNIST-derived protocols. The sources do not establish independent replication or results for transformers or large models. To assess generality, further comparisons would need to vary several parts of the design:

  • Architecture and scale: test whether the effect holds beyond the reported small convolutional network.
  • Dataset and task history: distinguish histories that change labels from those that change the input domain.
  • Shared-training schedule: vary the duration and schedule of the common phase.
  • Representation measurement: compare other layers and similarity metrics, not just Conv2 and fc1 under the reported CKA summary.
  • Downstream readout: test different readout methods and amounts of labeled data.
  • Replication and uncertainty: use more seeds and report uncertainty across runs.

Where to inspect the paper and reproduction materials

The arXiv record lists Ertuğrul Mutlu’s preprint as submitted on 29 September 2026, version 1, in Computer Science / Machine Learning. The author links code and reproducibility artifacts. The repository describes a Python virtual-environment setup, dependency installation from requirements.txt, paired training configurations, and validation; training downloads MNIST if it is not already present.

The repository notes that hardware, PyTorch, and CUDA differences can affect reproduction and that environment metadata is recorded when available. It presents generated experiment outputs as the underlying source of truth, so reproductions should be interpreted alongside their recorded environment and outputs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.