October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run ScanOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk3 min

LoRA and DoRA Explained: Math, Memory, and Trade-offs

LoRA trains compact low-rank updates while freezing pretrained weights; DoRA adds a separate learned magnitude component. Their memory and quality trade-offs depend on the model, task, and implementation.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

LoRA freezes a pretrained model’s weights and learns a compact, low-rank update. DoRA uses a similar low-rank update for weight direction, but also learns a separate magnitude component. That distinction changes the parameterization and training behavior; it does not make either method universally faster, more accurate, or lighter on memory. The practical choice depends on the model, task, implementation, and deployment path.

What LoRA changes in a pretrained layer

Consider a pretrained linear-layer weight matrix W0 with dimensions d × k. Full fine-tuning can update all dk entries. LoRA instead freezes W0 and represents the learned change as a low-rank update:

As an Amazon Associate I earn from qualifying purchases.

W = W0 + ΔW,   ΔW = BA

Here, B has dimensions d × r, A has dimensions r × k, and the rank r is chosen to be small relative to the layer dimensions. The factors contain r(d + k) trainable entries rather than dk entries for a dense update. Implementations commonly apply a scaling factor to the update. In the standard initialization described by the LoRA paper, the initial update is zero, so the adapted layer starts out matching the pretrained layer.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The savings come from constraining the update to a low-rank form and training the factors rather than the full matrix. That is an inductive constraint, not a guarantee that the resulting model will match full fine-tuning on every task. It also does not remove the need to keep the base model available or account for activations, optimizer state, precision, and other training costs.

#1 Best Overall
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning

How DoRA adds magnitude and direction

DoRA draws on weight normalization to separate a weight’s magnitude from its direction. In the paper’s normalized formulation, the adapted weight is:

W′ = m(V + BA) / ||V + BA||c

V begins as the pretrained weight matrix and is frozen in the described formulation. The magnitude component m and the low-rank factors A and B are learned. The low-rank factors adjust direction; the magnitude component provides a separate way to change scale.

The DoRA authors’ motivation is that ordinary LoRA couples magnitude and direction changes through its low-rank update, while separating them may better resemble full fine-tuning. Their paper reports an analysis in which magnitude-direction correlation was −0.62 for full fine-tuning, −0.31 for DoRA, and +0.83 for LoRA. Those values describe the paper’s selected analysis, not a general score of model quality or a prediction for another workload.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the papers’ memory figures actually compare

Memory numbers are meaningful only with their comparison setup attached. The LoRA paper reports 10,000 times fewer trainable parameters and a threefold reduction in GPU memory compared with full fine-tuning GPT-3 175B using Adam. These are results for that paper’s setup, not universal reduction factors for other model sizes, optimizers, sequence lengths, or implementations. Fewer trainable parameters do not mean that total training memory falls in the same proportion.

DoRA’s authors identify an additional training-memory cost from the changed backpropagation path. They propose treating the normalization denominator as constant during backpropagation while recalculating it dynamically. In the experiments reported, this modification reduced training memory by approximately 24.4% on LLaMA and 12.4% on VL-BART. The paper reports a 0.2 accuracy difference for LLaMA and unchanged accuracy for VL-BART in those experiments. These figures concern the authors’ modification and experimental comparisons; they are not a general DoRA-versus-LoRA memory guarantee.

How to choose and compare them for a workload

The LoRA and DoRA papers evaluate specific models, tasks, ranks, and implementations. Their results establish useful evidence about those experiments, but not a universal winner. Compare the methods on the setup you intend to run:

  • Task quality: evaluate both methods on the intended data and task rather than assuming paper results transfer.
  • Rank and target layers: rank and the choice of adapted modules affect trainable parameter count, adapter size, and capacity.
  • Training memory and throughput: account for optimizer state, activations, precision or quantization, sequence length, and implementation—not just the number of adapter parameters.
  • Inference and merging: both papers describe merging learned weights for inference without additional adapter latency in their method framing. Check that the framework and deployment path you use support the intended behavior.
  • Compatibility and maintenance: verify model architecture, layer types, quantization path, and current library versions before committing to an implementation.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Implementation support and license checks

Microsoft’s LoRA repository describes a PyTorch implementation, loralib, and notes support through Hugging Face PEFT. NVIDIA’s DoRA repository reports PEFT support for Linear, Conv1d, and Conv2d layers, as well as bitsandbytes-quantized linear layers. These are repository statements and can change; check current documentation and versions against your model and stack. The NVIDIA DoRA repository also identifies its NVIDIA Source Code License-NC, which should be reviewed before use.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Primary papers

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.