October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk8 min

Do AI Models Really Need Rebuilding Every Time They’re Updated? The Continual-Learning Problem Explained

Full retraining is common for major AI updates, but it is not inevitable. Here is why neural networks forget, what loss of plasticity means, how retrieval and continual-learning methods work, and when a costly rebuild is still justified.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

No. AI systems do not have to be rebuilt from zero for every update. However, substantial new data, new tasks or a major change in safety goals often trigger a full or near-full retraining run because ordinary neural networks can overwrite earlier skills and may gradually lose the ability to learn. Smaller updates can use fine-tuning, replay, retrieval, model editing or modular components, each with important trade-offs.

Why full retraining became the default

A neural network stores knowledge in shared numerical parameters, commonly called weights. Training changes those weights through gradient updates. When new examples require different internal representations, the same parameters that support the new behavior may also support older skills. Updating one capability can therefore interfere with another.

For a major data refresh, the safest conventional approach is to start from a checkpoint and train on the old and new material together—or discard the old checkpoint and train a new model from scratch. This gives the training process access to a broad distribution instead of asking it to remember everything while seeing only the newest data.

The authors of the 2024 Nature paper Loss of plasticity in deep continual learning describe the common practice this way: “In practice, the most common strategy for incorporating substantial new data has been simply to discard the old network and train a new one from scratch on the old and new data together.”

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

That approach is expensive, but it avoids having to preserve every earlier behavior through special mechanisms. It also creates a clean point for evaluation, safety testing and rollback.

Two different problems are often confused

Catastrophic forgetting

Catastrophic forgetting means that performance on earlier tasks or examples falls after the model learns new ones. The old data may not appear again during training, so updates optimized for the new distribution can overwrite useful representations. A model can look better on the latest benchmark while quietly becoming worse at older capabilities.

Loss of plasticity

Loss of plasticity is different. It is the decline in the network’s ability to learn new information at all after repeated updates. The Nature experiments, using continual-learning versions of ImageNet and CIFAR-100, found that standard deep-learning methods can become progressively less adaptable as new classes arrive. A model can therefore suffer both problems: it may forget old skills and become less capable of acquiring later ones.

Keeping these terms separate matters. A method that protects old accuracy may still leave the network hard to update, while a method that preserves learning capacity may not retain every previous skill.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning

What can replace a complete rebuild?

There is no universal substitute. The right method depends on whether historical data are available, how broad the change is, how quickly it must ship and how much auditability the deployment requires.

Fine-tuning

Fine-tuning continues training from an existing checkpoint on new examples. It is usually faster and cheaper than pretraining from zero, and it can specialize a general model for a domain or task. Without protection for earlier behavior, however, aggressive fine-tuning can cause catastrophic forgetting or unwanted shifts in tone, facts and safety behavior.

Replay of older examples

Replay mixes selected historical data with new data during the update. It directly reminds the model of prior tasks and can preserve accuracy better than training only on the newest material. The price is access to representative old data, additional storage and the privacy, licensing and governance work that comes with retaining it. If the original data cannot be recovered, replay may be limited or impossible.

Regularization and consolidation

These methods penalize changes to parameters judged important for earlier tasks. They reduce interference without storing every old example, but the importance estimates are imperfect. Protecting too many parameters can make the model rigid; protecting too few can leave old capabilities vulnerable.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Knowledge distillation

Distillation trains an updated model to match the outputs or behavior of an older model while learning the new task. It can preserve a teacher’s responses even when the original training set is unavailable. The teacher’s mistakes and biases are also transferred, and the new model may still need representative prompts to cover the old behavior.

Amazon Science’s 2021 work on continual learning for new natural-language tasks describes the practical motivation: “adding a new task to an existing MTL model usually requires retraining the model from scratch on all the tasks and this can be time-consuming and computationally expensive.” Its distillation approach updates an existing multi-task model while reducing forgetting.

Targeted model editing

Model editing changes a narrow set of facts or behaviors rather than running a broad training cycle. Microsoft Research has described techniques that cache and selectively retrieve new transformations between layers, allowing specific answers to be changed without treating every correction as full pretraining.

Editing is useful for a limited correction, but it is not a general replacement for learning a new domain. A patch can interact with nearby facts, fail under paraphrase or create a growing collection of exceptions. Each edit therefore needs regression tests and a reliable undo path.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Retrieval and external memory

Retrieval-augmented systems keep changing information outside the base model. At query time, a search or database component supplies documents, records or policies that the model can use to compose an answer. Updating the index is generally faster than changing billions of weights, and deleted records can be removed from the source.

Retrieval does not rewrite the model’s underlying reasoning or general knowledge. It depends on document quality, ranking, access controls and the model’s ability to use the retrieved context. It is a strong fit for current policies, catalogs and private documents, but not every new capability can be reduced to lookup.

Modular architectures

Adapters, task-specific modules, routed experts and other modular designs isolate some updates from the rest of the system. A new module can be trained or replaced while a shared base remains stable. This can improve rollback and ownership boundaries, although routing, compatibility and evaluation become more complex.

How the update choices compare

Strategy Retention of old capabilities Quality on broad new data or tasks Compute and memory Historical-data needs Deployment speed Audit, rollback and unlearning
Full retraining Strong when old and new data are trained together Broadest reset for a changed distribution, architecture or safety objective Highest Usually requires the old corpus or an equivalent replacement Slowest Clear new version, but expensive to reproduce; removal requires another controlled training process
Fine-tuning Variable; can forget older skills Good for a focused domain or task Lower than pretraining, but can still be substantial New data alone may be enough, though replay improves retention Relatively fast Simple checkpoint rollback; behavior changes need broad regression testing
Replay Often strong if the replay set represents prior uses Good for incremental tasks More training and storage than fine-tuning Requires retained, permitted historical examples Moderate Replay-set provenance and privacy must be documented
Regularization or consolidation Protects parameters judged important to old tasks Moderate; protection can limit adaptation Extra calculations, usually below a full rebuild May work with summaries or parameter statistics instead of all old data Moderate Importance estimates are harder to explain and reverse
Distillation Can preserve the older model’s observable behavior Good for adding tasks while retaining a teacher’s responses Requires running the teacher and training a student Old model access is essential; old examples improve coverage Moderate Teacher snapshots support rollback, but inherited errors must be tracked
Targeted editing Strong for the edited fact when the edit generalizes Narrow; not a substitute for broad learning Low relative to retraining New fact or transformation is enough for the patch Fast Can be reverted if edits are logged; interactions among many edits are difficult to audit
Retrieval or external memory Leaves base capabilities unchanged Excellent for current, document-grounded information; limited for new internal skills Low model-training cost, plus index and serving cost Needs authoritative, accessible source documents Fast refresh Source-level deletion and versioning are practical; answer behavior still requires monitoring
Modular updates Can isolate old modules from new ones Good when tasks or domains can be separated Varies with number of modules and routing Module-specific data can be enough Fast to moderate Clear module ownership and rollback, with added routing complexity
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Why “just teach ChatGPT a new fact” is harder than it sounds

A single factual correction is not the same as teaching a model a reliable concept. The fact may need to work across wording, languages, context windows and follow-up questions. It may also conflict with related facts already encoded in the weights.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a frequently changing source—such as a company policy, product inventory or legal document—retrieval is usually the cleaner update path. For a narrow, high-confidence correction that must appear without a lookup, model editing may be appropriate. For a new reasoning skill, style, modality or broad domain, fine-tuning or continued pretraining is more plausible, with replay, distillation or a fresh training run used to protect existing behavior.

What retraining a large model costs

The 2024 Nature paper states: “When the network is a large language model and the data are a substantial portion of the internet, then each retraining may cost millions of dollars in computation.” That is an order-of-magnitude warning, not a universal price for a named commercial model.

The actual bill depends on parameter count, token volume, accelerator type and rental rate, training duration, energy, networking, checkpoint storage, evaluation, safety testing and engineering time. A company may also need multiple failed or filtered runs before release. Because those variables differ, a precise per-model estimate cannot be inferred from the millions-of-dollars statement alone.

Does making models larger solve the update problem?

Not by itself. Google Research reports that larger pretrained ResNets and Transformers are more resistant to catastrophic forgetting than randomly initialized models trained from scratch, and that resistance improves with model and pretraining-data scale. Larger, better-pretrained models therefore provide a stronger starting point for continual learning.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Scale does not guarantee retention, preserve every safety property or prevent loss of plasticity. A model can still face distribution shifts, missing historical data, incompatible objectives and the need to remove information. Bigger models also make a full rebuild more expensive.

A practical way to choose an update path

  1. Define the change. Classify it as a changing reference document, a narrow factual correction, a focused task, a broad domain expansion or a change to architecture or safety objectives.
  2. Set the retention requirement. Identify which old capabilities, refusal behaviors, languages and edge cases must remain unchanged, and create tests for them before updating.
  3. Check data rights and availability. Determine whether representative historical examples can legally and safely be retained for replay or evaluation.
  4. Choose the smallest reliable mechanism. Use retrieval for frequently changing documents, editing for narrow corrections, adapters or fine-tuning for focused skills, and replay or distillation when old behavior must be preserved.
  5. Escalate when the foundation changed. A major distribution shift, new modality, architecture change or new safety objective can justify broad retraining rather than stacking patches.
  6. Evaluate and keep a rollback. Test new performance and old capabilities separately, record the exact data and checkpoint used, and retain a version that can be restored if regressions appear.

The bottom line

AI models do not have to be entirely rebuilt every time they are updated. Full retraining remains common because shared weights create interference, repeated updates can cause catastrophic forgetting and loss of plasticity, and training on old plus new data is a dependable way to reset the problem. Continual-learning methods can make updates cheaper and faster, but each exchanges something—historical data, compute, breadth, simplicity or auditability—for that convenience.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. Shenzhen desk3 min
    HONOR Expands Beyond Smartphones With Humanoid Robot RevealHONOR said it unveiled its first humanoid robot at MWC 2026 and named shopping assistance, workplace inspections, and supportive companionship as intended uses. Later Robotics D1 claims and a reported…
  2. Cupertino desk5 min
    Apple Unveils AirPods Max 2: The Upgrade That Should Have Happened Years AgoAirPods Max 2 adds H2-powered audio features and Apple claims up to 1.5× more effective ANC, but its design, Smart Case, and 20-hour battery rating are unchanged. Wired lossless audio…
  3. Cupertino desk4 min
    Apple’s OLED Touch MacBooks Are Coming—but the Dynamic Island Is the Real GambleApple has not announced an OLED touchscreen MacBook, but reports point to high-end models arriving in late 2026 or early 2027. The reported Mac Dynamic Island could be useful, but…
Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.