Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
A convolutional neural network (CNN), also called a ConvNet, is a neural network that learns patterns in grid-shaped data. For images, it applies small learned filters across local groups of pixels, then combines their responses into features that help predict what the image contains. Reusing each filter across the image makes CNNs especially effective at finding patterns that can appear in different locations.
How a convolutional network works
A color image is a grid of pixel values, commonly represented as height × width × 3 color channels. A CNN processes that grid in stages: early layers often respond to simple patterns such as edges or color changes; later layers combine those responses into more complex visual features. A classification network then uses the resulting representation to produce scores for possible labels.
| # | Preview | Product | Price | |
|---|---|---|---|---|
| 1 |
|
Deep Learning (Adaptive Computation and Machine Learning series) | $51.51 | Buy on Amazon |
| 2 |
|
Deep Learning: Foundations and Concepts | $49.55 | Buy on Amazon |
| 3 |
|
Understanding Deep Learning | $97.15 | Buy on Amazon |
| 4 |
|
Deep Learning (The MIT Press Essential Knowledge series) | $11.36 | Buy on Amazon |
| 5 |
|
Deep Learning: A Visual Approach | $55.86 | Buy on Amazon |
A basic pipeline looks like this:
image → convolution → activation → optional downsampling
→ more feature-extracting layers → prediction
This is a teaching pattern, not a required architecture. Modern CNNs may use residual connections, normalization, strided convolutions, global average pooling, or attention, and pooling is optional.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsWhat a filter and feature map do
A filter, or kernel, is a small array of learned weights. As it moves over an input, it examines a local patch, multiplies the patch values by corresponding filter weights, sums the products, and typically adds a bias. The resulting value records how strongly that filter responded at that position. Applying the filter across the input produces a feature map, also called an activation map.
#1 Best Overall
- Language Published: English
- Binding: hardcover
- It ensures you get the best usage for a longer period
For example, a layer with 16 filters produces 16 output channels—one map per filter. With a 32 × 32 × 3 input, sixteen 3 × 3 filters, stride 1, and padding that preserves the spatial dimensions, the output is 32 × 32 × 16. The final 16 is the number of learned feature maps, not the number of image colors.
Filters are generally learned during training rather than manually assigned meanings. Early filters may respond to edge-like patterns, color transitions, or textures; deeper layers can combine responses into more complex patterns. Descriptions such as “edge detector” are useful interpretations, not guaranteed labels for what a filter represents. Stanford’s CS231n convolutional-network notes explain this progression from local responses to more complex learned features.
Why CNNs share weights
Two design choices make CNNs efficient for images:
- Local connectivity: Each filter initially uses a small neighborhood instead of connecting every output to every pixel.
- Weight sharing: The same filter is applied at many positions, so it can respond to a pattern wherever that pattern occurs.
For a 224 × 224 RGB image, a fully connected layer from all 150,528 input values to 1,000 neurons would require more than 150 million weights, before biases. A convolution with 3 input channels, 16 output channels, a 3 × 3 kernel, and one bias per output channel has (3 × 3 × 3 × 16) + 16 = 448 parameters. These are not equivalent networks, but the comparison illustrates how filter reuse avoids a separate set of weights for every image location.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Rank #2
Weight sharing lets CNNs reuse a learned pattern detector across positions; it does not make them perfectly invariant to translation, rotation, scale, lighting, or viewpoint. Their assumptions—that nearby values matter and patterns may recur in different locations—are useful for images and other grid-like inputs, but are not ideal for every kind of data. The Deep Learning textbook’s chapter on convolutional networks describes convolutional layers as a way to exploit structure in grid-shaped data.
Stride, padding, activation, and pooling
Stride and padding set the output size
Stride is the number of input positions the filter moves between calculations. Stride 1 moves one position at a time; stride 2 skips positions and generally reduces the output’s height and width.
Padding adds values around the input border, usually zeros. With stride 1, “same” padding keeps height and width unchanged; “valid” padding adds none, so the output becomes smaller. For one spatial dimension, the output size is:
Rank #3
floor((input + 2 × padding − dilation × (kernel − 1) − 1) / stride + 1)
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchFor a 32 × 32 input and a 3 × 3 kernel at stride 1, padding 1 preserves a 32 × 32 output; no padding produces 30 × 30. Padding, stride, dilation, and output shapes are documented in the PyTorch Conv2d reference.
Activation adds nonlinearity
A convolution is a linear operation. CNNs commonly follow it with an activation such as ReLU, defined as max(0, x): positive values pass through and negative values become zero. This nonlinearity lets stacked layers represent more complex relationships than repeated linear operations alone.
Pooling and other downsampling reduce resolution
Max pooling keeps the largest value in a small region—for example, one value from each 2 × 2 region. Downsampling reduces spatial resolution and computation, and can make a representation less sensitive to small shifts. It can also discard detail, which matters when small features are important. Pooling is one option; strided convolutions and other methods can also reduce resolution.
How an image becomes a prediction
- Input: The network receives pixel values, usually arranged as height, width, and color channels (or in a framework-specific tensor order).
- Feature extraction: Convolutional layers compute local responses, and activation functions add nonlinearity.
- Downsampling: Pooling or strided operations may reduce spatial dimensions while later layers build broader representations.
- Classification: A prediction head turns the learned representation into scores for classes such as “cat,” “dog,” or “car.” A softmax is one common way to convert multiclass scores into probabilities.
The network is not given a fixed list of concepts such as “eye” or “wheel.” Those are human descriptions of patterns that may be useful to its training objective.
How the network learns
- The model processes a training example and makes a prediction.
- A loss function measures how far that prediction is from the correct target.
- Backpropagation calculates how the model’s parameters contributed to the error.
- An optimizer adjusts the filters and other parameters to reduce the loss.
- The process repeats across batches and epochs, while performance is checked on separate validation data.
One epoch is one pass through the training dataset; a batch is the group of examples processed together. The learning rate controls the size of optimizer updates. Data augmentation, such as cropping or flipping images, can expose the model to useful variations, while validation data helps reveal overfitting—when training performance is strong but performance on new examples is poor.
Best Value
Where CNNs are useful—and where they are not
CNNs are a natural option when local neighborhoods and repeated patterns matter. Uses include image classification, object detection, segmentation, optical character recognition, medical-image analysis, industrial inspection, video frames, audio spectrograms, time series, and 3D scans. They are not limited to photographs: the key is useful grid-like structure, not whether the data looks like an ordinary image.
A CNN can be a good fit when patterns may appear in different locations and efficient processing of a large grid matters. It may be less suitable when the task depends primarily on long-range relationships, when data is scarce, or when strict interpretability is essential. Transformers can model relationships between distant regions through attention, while fully connected networks are general-purpose but often inefficient for high-resolution images. Classical computer-vision methods may be preferable in constrained settings. No architecture is best for every task; data, compute, latency, and deployment requirements matter.
Common pitfalls and limitations
- Distribution shift: A model can fail when production images differ from training data in camera, lighting, geography, or other conditions.
- Shortcut learning: It may rely on background or texture correlations instead of the intended object feature.
- Lost detail: Repeated pooling or stride can erase small but important features.
- Preprocessing mismatch: Different pixel scaling, normalization, RGB/BGR order, or channel layout between training and inference can degrade predictions or cause shape errors.
- Dataset problems: Class imbalance can hide poor performance on rare categories; leakage between training and evaluation data can make results misleading.
- Compute and interpretability: CNNs can be computationally demanding, and their learned representations may be difficult to explain.
Weight sharing and downsampling can improve tolerance to some positional changes, but they do not guarantee robust behavior on unusual inputs. A CNN learns statistical patterns useful for its objective, not human-like understanding.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →A small PyTorch example
This layer accepts batches in PyTorch’s (N, C, H, W) layout: batch, channels, height, width. With 3 input channels, 16 filters, a 3 × 3 kernel, stride 1, and padding 1, it preserves the 32 × 32 spatial dimensions:
import torch
from torch import nn
layer = nn.Conv2d(
in_channels=3,
out_channels=16,
kernel_size=3,
stride=1,
padding=1
)
x = torch.randn(8, 3, 32, 32)
y = layer(x)
print(y.shape) # torch.Size([8, 16, 32, 32])
Framework terminology has a technical wrinkle: despite the familiar name, many deep-learning libraries implement cross-correlation, which slides the kernel without reversing it as strict mathematical convolution does. Because the weights are learned, this distinction usually does not change the practical explanation. See the PyTorch layer documentation and TensorFlow’s conv2d documentation.
Where the architecture came from
Modern convolutional networks have multiple historical predecessors; they were not invented in one moment. LeNet was an influential early practical architecture, and the 1998 work by LeCun, Bottou, Bengio, and Haffner is a major milestone. Stanford’s 2026 CS231n lecture identifies that work among the formative developments in computer vision.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Free tools Windows power users keep installed
One-click scans. No signup required.

