For RuntimeError: mat1 and mat2 shapes cannot be multiplied in a call to nn.Linear, compare the input tensor’s last dimension with that layer’s in_features. They must match. The correct fix depends on the tensor immediately before the failing layer: change the layer’s feature count only if the layer is configured incorrectly; reshape or reorder the input only if its feature axes are wrong.
What shapes does nn.Linear expect?
PyTorch’s Linear API reference defines the operation as y = xA^T + b. The input shape is (*, H_in), where the final dimension H_in must equal in_features. The output shape is (*, H_out), with the same leading dimensions and a final dimension equal to out_features.
That means nn.Linear is not limited to two-dimensional batches. It applies the transformation along the last axis of a vector, a batch, or a tensor with additional leading dimensions such as sequence positions. It does not automatically flatten those dimensions.
layer = torch.nn.Linear(in_features=20, out_features=30)
x = torch.randn(128, 20)
y = layer(x)
# x: (128, 20)
# layer.weight: (30, 20)
# y: (128, 30)
In this example, each of the 128 rows has 20 input features and is transformed into 30 output features. The weight’s shape is (out_features, in_features); if bias is enabled, its shape is (out_features,). These parameter shapes follow from the API contract, even though the expression uses the transposed weight in the multiplication.
#1 Best Overall
How to find the mismatch
- Locate the failing call. Read the traceback to find the exact
nn.Linearinvocation that raises the error. A model may have several linear layers, and the error text alone does not identify which one failed. - Inspect the input directly before it. Check the shape of the tensor passed to that layer—not just the shape of the model’s original input.
- Compare the last dimension. For an input shaped
(batch, features), comparefeatureswithlayer.in_features. For an input shaped(batch, sequence, features), compare the finalfeaturesdimension. The leading dimensions are preserved. - Choose the fix based on what the axes mean. If the final dimension contains the intended features but the layer expects the wrong count, configure the layer with the actual feature count. If the features are on a different axis, correct the upstream flattening, reshape, transpose, or permutation so the last axis represents the features the layer should consume.
Community examples on the PyTorch Forums illustrate several causes: an activation with more flattened features than the first linear layer expects, axes arranged incorrectly, or a layer configured for the wrong feature count. Those examples are not universal recipes; use the shape at your own failing call.
Should you change in_features, transpose, or flatten?
| What you find | Likely correction | Why |
|---|---|---|
The input’s final dimension is the intended feature count, but differs from in_features. |
Set in_features to that actual feature count, if the model design calls for those features. |
The layer’s input contract is determined by the last dimension. |
| The intended features exist, but are on an axis other than the last one. | Correct the upstream axis order or reshape. Transpose or permute only when it matches the meaning and layout of the data. | A blind transpose can swap batch, sequence, channel, or feature axes and change what each example represents. |
| A CNN activation still has spatial or channel dimensions before a fully connected layer. | Flatten the intended per-example feature dimensions while preserving the batch dimension; configure the linear layer for the resulting feature count. | The linear layer consumes the final feature dimension, not an implicit flattening of the whole activation. |
Do not add a transpose just because the error mentions matrices. A transpose can be right for a particular two-dimensional layout, but it is not a general fix for nn.Linear. Likewise, changing in_features merely to silence the exception is wrong if the input axes are arranged incorrectly.
Rank #2
What to do in a CNN-to-linear transition
After convolution and pooling, determine the activation shape immediately before the fully connected layer. Preserve the batch dimension and flatten the intended channel and spatial dimensions into one per-example feature dimension. Then make the first linear layer’s in_features equal to that flattened dimension.
The number of features depends on the actual activation shape after the preceding operations, so do not copy a dimension from an unrelated model or forum example. Confirm it from the tensor at the failing call and from the intended model layout.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Rank #3
Is this a shape error or a dtype error?
mat1 and mat2 shapes cannot be multiplied indicates incompatible dimensions for the operation. A dtype mismatch—such as input and parameters using incompatible floating-point types—is a separate problem and should not be addressed by changing in_features. Diagnose the error actually reported rather than treating all matrix-operation exceptions as the same issue.
Quick checks before rerunning
- The traceback points to the specific failing linear layer.
- You checked the tensor shape immediately before that call.
- Its last dimension matches the layer’s
in_features. - Any flattening preserves the batch dimension and groups each example’s features as intended.
- Any axis permutation reflects the data layout rather than being a trial-and-error transpose.
For a broader foundation in tensors and neural-network workflows, PyTorch’s Learn the Basics tutorial introduces tensors, neural networks, and an image-classification example.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




