Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsBoth versions below compute the same affine function, y = x @ weight + bias. The difference is how PyTorch discovers and manages the model’s state: with raw tensors, you pass and organize that state yourself; with nn.Module, parameters and child modules are registered so framework tools can traverse, move, optimize, save, and restore them.
What changes when you use nn.Module?
PyTorch’s API documentation calls torch.nn.Module the “Base class for all neural network modules.” It is an organizing and registration abstraction, not a requirement for tensor arithmetic or automatic differentiation. Autograd can compute gradients through tensor operations either way. The practical distinction is that a module provides a standard place for model state and a common interface for working with it.
As an Amazon Associate I earn from qualifying purchases.
The examples use PyTorch 2.14 documentation as their version context. They intentionally keep the calculation identical so the difference in state management is clear.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallThe same affine model in two forms
Direct tensor operations
In a raw-tensor implementation, keep references to the learnable tensors and use them directly:
#1 Best Overall
import torch
weight = torch.randn(3, 2, requires_grad=True)
bias = torch.randn(2, requires_grad=True)
def predict(x):
return x @ weight + bias
x = torch.randn(4, 3)
y = predict(x)
Here, weight and bias are ordinary tensor variables. Setting requires_grad=True lets autograd track operations involving them, but it does not register them with a model object. You must pass the intended tensors to an optimizer yourself, for example torch.optim.SGD([weight, bias], lr=0.01), and decide how to organize and save them.
Subclassing nn.Module
A module puts the same state and calculation behind PyTorch’s standard model interface:
Rank #2
import torch
from torch import nn
class Affine(nn.Module):
def __init__(self):
super().__init__()
self.weight = nn.Parameter(torch.randn(3, 2))
self.bias = nn.Parameter(torch.randn(2))
def forward(self, x):
return x @ self.weight + self.bias
model = Affine()
x = torch.randn(4, 3)
y = model(x)
The call super().__init__() initializes the module machinery before attributes are assigned. Wrapping each learnable tensor in nn.Parameter marks it for registration. The arithmetic remains the same, but the optimizer can now obtain the parameters through model.parameters(), or their names through model.named_parameters().
Recommended Free Tools
How the two versions differ in everyday use
| Task | Raw tensors | nn.Module |
|---|---|---|
| Where weight and bias live | In variables or another structure you manage. | As registered nn.Parameter attributes. |
| Give learnable values to an optimizer | Pass the chosen tensors explicitly, such as [weight, bias]. |
Pass model.parameters(). |
| Find nested components | Track and traverse your own structures. | Assign child modules as attributes; the parent registers them recursively. |
| Apply device or dtype changes | Move or convert each relevant tensor yourself. | Use module operations such as model.to(...) to act on registered parameters and buffers. |
| Save and restore model state | Choose what to save and how to map it back yourself. | Use state_dict() and load_state_dict() for registered state. |
This registration is especially useful as a model grows. For example, if a parent module assigns a layer to self.layer, that child is discoverable through the parent’s module hierarchy. Parent-level parameter iteration, state inspection, and device conversion can then include its registered contents without separately listing every layer.
Rank #3
Parameters, buffers, and what a checkpoint contains
Parameters are learnable state
A module does not treat every tensor attribute as a parameter. Use nn.Parameter for a learnable tensor that should appear in parameter iteration, or use a built-in module such as nn.Linear. A plain tensor assigned as an attribute is not automatically equivalent for parameters() enumeration.
Buffers are state that is not optimized
Some module state is needed for computation but is not a learnable parameter. BatchNorm running statistics are a common example. Register such state as a buffer: persistent buffers are included in state_dict(), while non-persistent buffers are deliberately left out. Both kinds of registered buffers are affected by module-wide device and dtype changes through to().
A state dictionary is state, not the model definition
A module’s state_dict() contains its parameters and persistent buffers, keyed by their names in the module hierarchy. PyTorch documents the returned dictionary as a shallow copy whose values reference the module’s parameters and buffers; by default, the returned tensors are detached from autograd. It is useful for saving and loading state, but it does not contain the executable Python architecture itself.
Free tools Windows power users keep installed
One-click scans. No signup required.
To restore a model from a state dictionary, construct a compatible module and load the saved state with load_state_dict(). With strict loading enabled, checkpoint keys must match the keys expected by the module. Loading copies the state into the module hierarchy; it is not a replacement for defining that hierarchy.
When should you choose each approach?
Use raw tensors for a deliberately small calculation
Direct tensor code can be appropriate for a compact experiment, a one-off operation, or an example focused on the math. It keeps the computation explicit, but you take responsibility for parameter lists, nested state, device and dtype handling, and checkpoint organization as those needs arise.
Use nn.Module for reusable or composed models
Subclass nn.Module when you want learnable state and components to participate in PyTorch’s standard model workflows. Define state in __init__, call super().__init__() first, and implement the calculation in forward. Registered parameters and child modules make optimizer setup and model-state traversal more systematic.
Neither form changes the function being computed in the example, and the module abstraction should not be assumed to make computation faster. Any performance difference would need to be measured for the specific implementations and workload.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




