LoRA freezes a pretrained model’s weights and learns a compact, low-rank update. DoRA uses a similar low-rank update for weight direction, but also learns a separate magnitude component. That distinction changes the parameterization and training behavior; it does not make either method universally faster, more accurate, or lighter on memory. The practical choice depends on the model, task, implementation, and deployment path.
What LoRA changes in a pretrained layer
Consider a pretrained linear-layer weight matrix W0 with dimensions d × k. Full fine-tuning can update all dk entries. LoRA instead freezes W0 and represents the learned change as a low-rank update:
As an Amazon Associate I earn from qualifying purchases.
W = W0 + ΔW, ΔW = BA
Here, B has dimensions d × r, A has dimensions r × k, and the rank r is chosen to be small relative to the layer dimensions. The factors contain r(d + k) trainable entries rather than dk entries for a dense update. Implementations commonly apply a scaling factor to the update. In the standard initialization described by the LoRA paper, the initial update is zero, so the adapted layer starts out matching the pretrained layer.
The savings come from constraining the update to a low-rank form and training the factors rather than the full matrix. That is an inductive constraint, not a guarantee that the resulting model will match full fine-tuning on every task. It also does not remove the need to keep the base model available or account for activations, optimizer state, precision, and other training costs.
#1 Best Overall
- Use scikit-learn to track an example ML project end to end
- Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
- Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
- Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
- Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
How DoRA adds magnitude and direction
DoRA draws on weight normalization to separate a weight’s magnitude from its direction. In the paper’s normalized formulation, the adapted weight is:
W′ = m(V + BA) / ||V + BA||c
V begins as the pretrained weight matrix and is frozen in the described formulation. The magnitude component m and the low-rank factors A and B are learned. The low-rank factors adjust direction; the magnitude component provides a separate way to change scale.
Rank #2
The DoRA authors’ motivation is that ordinary LoRA couples magnitude and direction changes through its low-rank update, while separating them may better resemble full fine-tuning. Their paper reports an analysis in which magnitude-direction correlation was −0.62 for full fine-tuning, −0.31 for DoRA, and +0.83 for LoRA. Those values describe the paper’s selected analysis, not a general score of model quality or a prediction for another workload.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
What the papers’ memory figures actually compare
Memory numbers are meaningful only with their comparison setup attached. The LoRA paper reports 10,000 times fewer trainable parameters and a threefold reduction in GPU memory compared with full fine-tuning GPT-3 175B using Adam. These are results for that paper’s setup, not universal reduction factors for other model sizes, optimizers, sequence lengths, or implementations. Fewer trainable parameters do not mean that total training memory falls in the same proportion.
DoRA’s authors identify an additional training-memory cost from the changed backpropagation path. They propose treating the normalization denominator as constant during backpropagation while recalculating it dynamically. In the experiments reported, this modification reduced training memory by approximately 24.4% on LLaMA and 12.4% on VL-BART. The paper reports a 0.2 accuracy difference for LLaMA and unchanged accuracy for VL-BART in those experiments. These figures concern the authors’ modification and experimental comparisons; they are not a general DoRA-versus-LoRA memory guarantee.
How to choose and compare them for a workload
The LoRA and DoRA papers evaluate specific models, tasks, ranks, and implementations. Their results establish useful evidence about those experiments, but not a universal winner. Compare the methods on the setup you intend to run:
Rank #4
- Task quality: evaluate both methods on the intended data and task rather than assuming paper results transfer.
- Rank and target layers: rank and the choice of adapted modules affect trainable parameter count, adapter size, and capacity.
- Training memory and throughput: account for optimizer state, activations, precision or quantization, sequence length, and implementation—not just the number of adapter parameters.
- Inference and merging: both papers describe merging learned weights for inference without additional adapter latency in their method framing. Check that the framework and deployment path you use support the intended behavior.
- Compatibility and maintenance: verify model architecture, layer types, quantization path, and current library versions before committing to an implementation.
Implementation support and license checks
Microsoft’s LoRA repository describes a PyTorch implementation, loralib, and notes support through Hugging Face PEFT. NVIDIA’s DoRA repository reports PEFT support for Linear, Conv1d, and Conv2d layers, as well as bitsandbytes-quantized linear layers. These are repository statements and can change; check current documentation and versions against your model and stack. The NVIDIA DoRA repository also identifies its NVIDIA Source Code License-NC, which should be reviewed before use.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Quick Recap
Best Value
Primary papers
- Hu et al., “LoRA: Low-Rank Adaptation of Large Language Models” (2021 preprint; published at ICLR 2022).
- Liu et al., “DoRA: Weight-Decomposed Low-Rank Adaptation” (2024; ICML 2024).
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




