Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

A convolutional neural network (CNN), also called a ConvNet, is a neural network that learns patterns in grid-shaped data. For images, it applies small learned filters across local groups of pixels, then combines their responses into features that help predict what the image contains. Reusing each filter across the image makes CNNs especially effective at finding patterns that can appear in different locations.

How a convolutional network works

A color image is a grid of pixel values, commonly represented as height × width × 3 color channels. A CNN processes that grid in stages: early layers often respond to simple patterns such as edges or color changes; later layers combine those responses into more complex visual features. A classification network then uses the resulting representation to produce scores for possible labels.

A basic pipeline looks like this:

image → convolution → activation → optional downsampling
       → more feature-extracting layers → prediction

This is a teaching pattern, not a required architecture. Modern CNNs may use residual connections, normalization, strided convolutions, global average pooling, or attention, and pooling is optional.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What a filter and feature map do

A filter, or kernel, is a small array of learned weights. As it moves over an input, it examines a local patch, multiplies the patch values by corresponding filter weights, sums the products, and typically adds a bias. The resulting value records how strongly that filter responded at that position. Applying the filter across the input produces a feature map, also called an activation map.

#1 Best Overall
Sale
Deep Learning (Adaptive Computation and Machine Learning series)
  • Language Published: English
  • Binding: hardcover
  • It ensures you get the best usage for a longer period

For example, a layer with 16 filters produces 16 output channels—one map per filter. With a 32 × 32 × 3 input, sixteen 3 × 3 filters, stride 1, and padding that preserves the spatial dimensions, the output is 32 × 32 × 16. The final 16 is the number of learned feature maps, not the number of image colors.

Filters are generally learned during training rather than manually assigned meanings. Early filters may respond to edge-like patterns, color transitions, or textures; deeper layers can combine responses into more complex patterns. Descriptions such as “edge detector” are useful interpretations, not guaranteed labels for what a filter represents. Stanford’s CS231n convolutional-network notes explain this progression from local responses to more complex learned features.

Why CNNs share weights

Two design choices make CNNs efficient for images:

  • Local connectivity: Each filter initially uses a small neighborhood instead of connecting every output to every pixel.
  • Weight sharing: The same filter is applied at many positions, so it can respond to a pattern wherever that pattern occurs.

For a 224 × 224 RGB image, a fully connected layer from all 150,528 input values to 1,000 neurons would require more than 150 million weights, before biases. A convolution with 3 input channels, 16 output channels, a 3 × 3 kernel, and one bias per output channel has (3 × 3 × 3 × 16) + 16 = 448 parameters. These are not equivalent networks, but the comparison illustrates how filter reuse avoids a separate set of weights for every image location.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Weight sharing lets CNNs reuse a learned pattern detector across positions; it does not make them perfectly invariant to translation, rotation, scale, lighting, or viewpoint. Their assumptions—that nearby values matter and patterns may recur in different locations—are useful for images and other grid-like inputs, but are not ideal for every kind of data. The Deep Learning textbook’s chapter on convolutional networks describes convolutional layers as a way to exploit structure in grid-shaped data.

Stride, padding, activation, and pooling

Stride and padding set the output size

Stride is the number of input positions the filter moves between calculations. Stride 1 moves one position at a time; stride 2 skips positions and generally reduces the output’s height and width.

Padding adds values around the input border, usually zeros. With stride 1, “same” padding keeps height and width unchanged; “valid” padding adds none, so the output becomes smaller. For one spatial dimension, the output size is:

floor((input + 2 × padding − dilation × (kernel − 1) − 1) / stride + 1)

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a 32 × 32 input and a 3 × 3 kernel at stride 1, padding 1 preserves a 32 × 32 output; no padding produces 30 × 30. Padding, stride, dilation, and output shapes are documented in the PyTorch Conv2d reference.

Activation adds nonlinearity

A convolution is a linear operation. CNNs commonly follow it with an activation such as ReLU, defined as max(0, x): positive values pass through and negative values become zero. This nonlinearity lets stacked layers represent more complex relationships than repeated linear operations alone.

Pooling and other downsampling reduce resolution

Max pooling keeps the largest value in a small region—for example, one value from each 2 × 2 region. Downsampling reduces spatial resolution and computation, and can make a representation less sensitive to small shifts. It can also discard detail, which matters when small features are important. Pooling is one option; strided convolutions and other methods can also reduce resolution.

How an image becomes a prediction

  1. Input: The network receives pixel values, usually arranged as height, width, and color channels (or in a framework-specific tensor order).
  2. Feature extraction: Convolutional layers compute local responses, and activation functions add nonlinearity.
  3. Downsampling: Pooling or strided operations may reduce spatial dimensions while later layers build broader representations.
  4. Classification: A prediction head turns the learned representation into scores for classes such as “cat,” “dog,” or “car.” A softmax is one common way to convert multiclass scores into probabilities.

The network is not given a fixed list of concepts such as “eye” or “wheel.” Those are human descriptions of patterns that may be useful to its training objective.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How the network learns

  1. The model processes a training example and makes a prediction.
  2. A loss function measures how far that prediction is from the correct target.
  3. Backpropagation calculates how the model’s parameters contributed to the error.
  4. An optimizer adjusts the filters and other parameters to reduce the loss.
  5. The process repeats across batches and epochs, while performance is checked on separate validation data.

One epoch is one pass through the training dataset; a batch is the group of examples processed together. The learning rate controls the size of optimizer updates. Data augmentation, such as cropping or flipping images, can expose the model to useful variations, while validation data helps reveal overfitting—when training performance is strong but performance on new examples is poor.

Best Value
Sale
Deep Learning: A Visual Approach
  • Deep Learning: A Visual Approach
  • No Starch Press
  • ABIS BOOK
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Where CNNs are useful—and where they are not

CNNs are a natural option when local neighborhoods and repeated patterns matter. Uses include image classification, object detection, segmentation, optical character recognition, medical-image analysis, industrial inspection, video frames, audio spectrograms, time series, and 3D scans. They are not limited to photographs: the key is useful grid-like structure, not whether the data looks like an ordinary image.

A CNN can be a good fit when patterns may appear in different locations and efficient processing of a large grid matters. It may be less suitable when the task depends primarily on long-range relationships, when data is scarce, or when strict interpretability is essential. Transformers can model relationships between distant regions through attention, while fully connected networks are general-purpose but often inefficient for high-resolution images. Classical computer-vision methods may be preferable in constrained settings. No architecture is best for every task; data, compute, latency, and deployment requirements matter.

Common pitfalls and limitations

  • Distribution shift: A model can fail when production images differ from training data in camera, lighting, geography, or other conditions.
  • Shortcut learning: It may rely on background or texture correlations instead of the intended object feature.
  • Lost detail: Repeated pooling or stride can erase small but important features.
  • Preprocessing mismatch: Different pixel scaling, normalization, RGB/BGR order, or channel layout between training and inference can degrade predictions or cause shape errors.
  • Dataset problems: Class imbalance can hide poor performance on rare categories; leakage between training and evaluation data can make results misleading.
  • Compute and interpretability: CNNs can be computationally demanding, and their learned representations may be difficult to explain.

Weight sharing and downsampling can improve tolerance to some positional changes, but they do not guarantee robust behavior on unusual inputs. A CNN learns statistical patterns useful for its objective, not human-like understanding.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A small PyTorch example

This layer accepts batches in PyTorch’s (N, C, H, W) layout: batch, channels, height, width. With 3 input channels, 16 filters, a 3 × 3 kernel, stride 1, and padding 1, it preserves the 32 × 32 spatial dimensions:

import torch
from torch import nn

layer = nn.Conv2d(
    in_channels=3,
    out_channels=16,
    kernel_size=3,
    stride=1,
    padding=1
)

x = torch.randn(8, 3, 32, 32)
y = layer(x)
print(y.shape)  # torch.Size([8, 16, 32, 32])

Framework terminology has a technical wrinkle: despite the familiar name, many deep-learning libraries implement cross-correlation, which slides the kernel without reversing it as strict mathematical convolution does. Because the weights are learned, this distinction usually does not change the practical explanation. See the PyTorch layer documentation and TensorFlow’s conv2d documentation.

Where the architecture came from

Modern convolutional networks have multiple historical predecessors; they were not invented in one moment. LeNet was an influential early practical architecture, and the 1998 work by LeCun, Bottou, Bengio, and Haffner is a major milestone. Stanford’s 2026 CS231n lecture identifies that work among the formative developments in computer vision.

Quick Recap

SaleBestseller No. 1
Deep Learning (Adaptive Computation and Machine Learning series)
Deep Learning (Adaptive Computation and Machine Learning series)
Language Published: English; Binding: hardcover; It ensures you get the best usage for a longer period
$51.51
SaleBestseller No. 2
Bestseller No. 3
SaleBestseller No. 5
Deep Learning: A Visual Approach
Deep Learning: A Visual Approach
Deep Learning: A Visual Approach; No Starch Press; ABIS BOOK
$55.86

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.