To calculate a PyTorch nn.Conv2d output shape, keep the batch and channel dimensions in view, then apply the documented height and width formulas using kernel size, stride, padding, and dilation. The output is (N, out_channels, H_out, W_out) for a batched input; each spatial result is rounded down when the division by stride is not exact.
What nn.Conv2d takes as input and returns
PyTorch’s Conv2d API documentation describes the module as applying a 2D convolution over an input signal composed of several input planes. The operation is implemented as valid 2D cross-correlation, with a bias added for each output channel.
As an Amazon Associate I earn from qualifying purchases.
A batched input has shape (N, C_in, H_in, W_in); its output has shape (N, C_out, H_out, W_out). An unbatched input is also supported: (C_in, H_in, W_in) produces (C_out, H_out, W_out). Here, N is batch size, C_in must equal the layer’s in_channels, and C_out is the configured out_channels.
Recommended Free Tools
How to calculate the spatial output shape
For height and width parameters supplied as pairs in (height, width) order, calculate each output dimension separately:
#1 Best Overall
H_out = floor((H_in + 2*padding[0] - dilation[0]*(kernel_size[0] - 1) - 1) / stride[0] + 1)
W_out = floor((W_in + 2*padding[1] - dilation[1]*(kernel_size[1] - 1) - 1) / stride[1] + 1)
If a spatial argument is a single integer, that value applies to both axes. The floor operation means any fractional result rounds down; do not round to the nearest integer.
Worked example
For an input shaped (20, 16, 50, 100) and a layer configured as nn.Conv2d(16, 33, (3, 5), stride=(2, 1), padding=(4, 2), dilation=(3, 1)), the height calculation is floor((50 + 2*4 - 3*(3-1) - 1)/2 + 1) = 27. The width calculation is floor((100 + 2*2 - 1*(5-1) - 1)/1 + 1) = 100. The resulting shape is (20, 33, 27, 100).
Rank #2
This Python example uses the documented configuration; the expected shape is derived from the documented formula:
import torch
from torch import nn
layer = nn.Conv2d(
in_channels=16,
out_channels=33,
kernel_size=(3, 5),
stride=(2, 1),
padding=(4, 2),
dilation=(3, 1),
)
x = torch.randn(20, 16, 50, 100)
y = layer(x)
print(y.shape) # expected: (20, 33, 27, 100)
What each Conv2d parameter controls
The documented module signature is:
nn.Conv2d(
in_channels,
out_channels,
kernel_size,
stride=1,
padding=0,
dilation=1,
groups=1,
bias=True,
padding_mode="zeros",
device=None,
dtype=None,
)
| Parameter | What it controls | Effect to check |
|---|---|---|
in_channels |
Number of channels in the input. | Must match the input’s channel dimension. |
out_channels |
Number of output channels. | Sets C_out and the first dimension of the weight tensor. |
kernel_size |
Height and width of the convolution window. | A larger effective kernel generally changes the spatial result; rectangular kernels can produce different height and width behavior. |
stride |
Step between window positions. | Values above 1 reduce the number of positions and can trigger floor rounding. |
padding |
Implicit padding on each side of each spatial axis. | Numeric values contribute twice to the corresponding dimension because they apply on both sides. |
dilation |
Spacing between kernel points. | Increases the effective spatial reach of the kernel without changing the stored kernel dimensions. |
groups |
Controls which input channels connect to which output channels. | Must divide both channel counts; it also changes the number of weights. |
bias |
Whether to learn one bias value per output channel. | When disabled, no bias values are included in the parameter count. |
padding_mode |
How numeric padding values are filled. | Documented modes are zeros, reflect, replicate, and circular. |
device, dtype |
Device and data type for the module’s parameters. | Use these when the layer’s placement or type needs to be specified at construction. |
For kernel_size, stride, padding, and dilation, an integer repeats across height and width; a pair specifies the height value first and width value second.
Rank #3
Padding choices and spatial-size effects
padding=0adds no numeric padding. Withpadding='valid', the documented behavior is no padding.- A numeric padding amount applies to both sides of its spatial axis. For example, a height padding of 4 adds 8 to the height term before the kernel and stride are applied.
padding='same'keeps output height and width equal to the input dimensions, but only when stride is 1. The API does not support this mode with other stride values.
Padding changes where the window can be placed; it does not change out_channels. For non-square inputs or settings, calculate height and width independently rather than assuming a square output.
Groups, depthwise convolution, and channel connections
With the default groups=1, every input channel connects to every output channel. Setting groups=2 partitions the operation into two channel groups. Both in_channels and out_channels must be divisible by groups.
The API calls a configuration depthwise convolution when groups == in_channels and out_channels == K * in_channels, where K is a positive integer. Grouping reduces the number of input channels connected to each output channel, and therefore reduces the weight count for a fixed kernel and output-channel count.
Weight shape and learnable parameter count
The weight tensor shape is (out_channels, in_channels/groups, kernel_height, kernel_width). If bias is enabled, its shape is (out_channels,). The resulting learnable parameter count is:
out_channels * (in_channels / groups) * kernel_height * kernel_width
+ (out_channels if bias else 0)
For Conv2d(16, 33, 3, stride=2), the defaults are groups=1 and bias=True, so the count is 33 * 16 * 3 * 3 + 33 = 4,785 parameters. This is a calculation from the documented tensor shapes.
The Conv2d documentation also describes a uniform initialization distribution whose bound depends on channel count, groups, and kernel area. Initialization is random; the documentation does not imply that each construction produces identical parameter values.
Why an output shape may differ from expectation
- Wrong dimension order: Conv2d expects channel-first input,
(N, C, H, W)or unbatched(C, H, W), not channel-last. - Channel mismatch: The input channel count must equal
in_channels; the returned channel count isout_channels. - Forgotten dilation: The formula uses
dilation * (kernel_size - 1), not simply the kernel size minus one. - Padding counted only once: Numeric padding is on both sides, so the formula includes
2 * padding. - Missing floor: When the intermediate division by stride is fractional, the resulting dimension rounds down.
- Unequal axis settings: Tuple entries are height then width, and the two output dimensions can differ.
- Unsupported same-padding stride:
padding='same'is supported only with stride 1.
Implementation notes
The Conv2d reference documents support for TensorFloat32 and complex data types. It also notes that on certain ROCm devices, float16 inputs use different precision for backward computation. These are backend-specific qualifications, not statements that every device or run behaves identically.
The functional conv2d reference notes that some CUDA/CuDNN situations may select a nondeterministic algorithm for performance. If deterministic behavior is preferred, it points to torch.backends.cudnn.deterministic = True; this can come with a performance cost.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




