October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk8 min

The Mathematics of Forward and Backpropagation

A rigorous, practical guide to forward propagation and backpropagation: derive the equations, understand transposes and outer products, implement a small network, and diagnose gradient failures.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Forward propagation computes a neural network’s prediction; backpropagation computes how the resulting loss changes with respect to every parameter. An optimizer such as SGD or Adam then uses those gradients to update weights and biases. Keeping these three stages separate makes the equations, implementations, and debugging much easier to understand.

What problem does backpropagation solve?

A network may contain millions of parameters, yet training needs the gradient of one scalar loss with respect to all of them. Differentiating each parameter independently would repeat the same downstream calculations. Backpropagation reuses intermediate derivatives by traversing the computation graph in reverse, accumulating the complete gradient efficiently.

Backpropagation is a gradient-computation algorithm. It does not select a learning rate, choose an optimizer, or update parameters by itself:

  1. Forward pass: evaluate activations, prediction, and loss.
  2. Backward pass: compute derivatives of the loss with respect to activations, weights, and biases.
  3. Optimizer step: apply an update such as W ← W - η∇W.

Rumelhart, Hinton, and Williams popularized a practical multilayer-network formulation in 1986, although related automatic-differentiation and gradient-propagation work predates that paper (Nature, 1986; historical context at Harvard CS249r).

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
EXPO Dry Erase Markers Kit, Chisel Tip, Assorted Colors, Eraser, Spray Cleaner, 6 Count - Whiteboard, Calendar, Office Essentials, School, Classroom, Teacher Supplies
  • Dry erase markers with the most vibrant ink yet from EXPO
  • Vibrant ink makes it easier to read information from a distance
  • Made for the whiteboard and beyond, writing pops on most non-porous surfaces like glass, acrylic, and more!
  • Easily and cleanly erases with included EXPO eraser and cleaner spray
  • Versatile chisel tip creates multiple line widths

Minimum mathematics: derivatives, gradients, and Jacobians

Scalar derivative

For y = f(x), dy/dx measures the local change in y caused by a change in x.

Gradient

For a scalar loss L(w) and parameter vector w ∈ Rn,

∇wL = [∂L/∂w1, …, ∂L/∂wn]T.

The gradient points toward steepest local increase; its negative is the direction used by basic gradient descent.

Jacobian and the vector chain rule

For y = f(x), where x ∈ Rn and y ∈ Rm, the Jacobian Jf is an m × n matrix. Composition obeys

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Jg○f(x) = Jg(f(x))Jf(x).

Backpropagation applies this rule without explicitly constructing every large Jacobian. In differential notation, dL = (∂L/∂x)Tdx; this form makes transpose placement and tensor shapes easier to check.

One neuron: the chain rule in full

A neuron first forms an affine value and then applies an activation:

Rank #2
Sale
EXPO Dry Erase Markers, Low Odor Ink, Assorted Colors, Chisel Tip, 12 Count
  • Dry erase markers with the most vibrant ink yet from EXPO
  • Vibrant ink makes it easier to read information from a distance
  • Made for the whiteboard and beyond, writing pops on most non-porous surfaces like glass, acrylic, and more!
  • Easily and cleanly erases with an EXPO eraser or dry cloth
  • Versatile chisel tip creates multiple line widths

z = wTx + b
a = f(z).

For one weight wi and loss L(a,y),

∂L/∂wi = (∂L/∂a) f′(z) xi.

Likewise, ∂L/∂b = (∂L/∂a)f′(z). A weight’s gradient is therefore the product of three factors: downstream loss sensitivity, the neuron’s local activation slope, and the input connected to that weight.

Forward propagation through layers

Use column vectors for one example. For layer l:

z(l) = W(l)a(l-1) + b(l)
a(l) = f(l)(z(l)), with a(0) = x.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Quantity Shape Meaning
W(l) nl × nl-1 Weights from the previous layer
a(l-1) nl-1 Previous activations
b(l), z(l), a(l) nl Bias, preactivation, and postactivation

For a batch stored as rows, a common convention is A(l-1) ∈ Rm×nl-1 and Z(l) = A(l-1)(W(l))T + b(l), with bias broadcasting across rows. Row-batch and column-example conventions are both valid; apparent transpose disagreements usually come from switching conventions.

Loss functions determine the starting gradient

Mean-squared error

For one scalar prediction, L = ½(ŷ-y)2, so ∂L/∂ŷ = ŷ-y. The factor one-half cancels the derivative’s factor of two.

Sigmoid plus binary cross-entropy

With ŷ = σ(z) = 1/(1+e-z) and L = -[y log ŷ + (1-y)log(1-ŷ)], the combined derivative is

∂L/∂z = ŷ-y.

This simplification is why frameworks provide fused sigmoid-cross-entropy losses.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
EXPO Dry Erase Markers Kit, Fine and Chisel Tip Markers, Assorted Colors, Eraser, Spray Cleaner, 14 Count
  • EXPO kit comes with everything you need to start marking and keep your surfaces clean
  • Consistent, skip-free writing, vibrant color options and low-odor ink make the kit perfect for classrooms and offices
  • Versatile chisel tip allows for broad and fine writing. Fine tip is great for details
  • Spray and Expo eraser help you erase cleanly and easily while also extending whiteboard life
  • 14-piece set includes fine and chisel tip markers in Black, Red, Blue, Green, Orange, Brown, Purple & Lime plus an 8 oz. bottle of Expo white board cleaning spray & an Expo eraser

Softmax plus multiclass cross-entropy

softmax(z)i = ezi/∑jezj. For one-hot target y and L = -∑iyilog ŷi,

∂L/∂z = ŷ-y.

Naively exponentiating large logits can overflow, so use a numerically stabilized softmax or a framework loss that accepts logits directly.

Deriving the backpropagation recurrence

Define the error signal (also called an adjoint)

δ(l) = ∂L/∂z(l).

Output layer

δ(N) = (∂L/∂a(N)) ⊙ f′(N)(z(N)). For sigmoid with binary cross-entropy and softmax with cross-entropy, this often reduces to prediction minus target.

Hidden layers

Because z(l+1) = W(l+1)a(l) + b(l+1),

δ(l) = ((W(l+1))Tδ(l+1)) ⊙ f′(l)(z(l)).

  • The transposed matrix distributes downstream sensitivities to earlier activations.
  • The activation derivative applies each neuron’s local slope.
  • ⊙ is elementwise multiplication.

Weight and bias gradients

For z(l) = W(l)a(l-1) + b(l),

∂L/∂W(l) = δ(l)(a(l-1))T
∂L/∂b(l) = δ(l).

Elementwise, ∂L/∂Wij(l) = δi(l)aj(l-1), which is exactly the outer product. The bias derivative has no input multiplier because ∂zi/∂bi = 1.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Computational graphs, branches, and transposes

The network is a directed acyclic graph: x → z(1) → a(1) → … → prediction → L. Backward propagation starts with ∂L/∂L = 1. For a node v = g(u1,…,uk),

∂L/∂ui = (∂L/∂v)(∂v/∂ui).

If several paths reach the same variable, contributions add: ∂L/∂u = ∑r(∂L/∂vr)(∂vr/∂u). This matters for residual connections, branches, shared parameters, and recurrent unrolling. PyTorch describes reverse automatic differentiation over a dynamically built graph (autograd notes); TensorFlow records operations with GradientTape (API reference).

Rank #4
Sale
EXPO Dry Erase Markers, Low Odor Ink, Assorted Colors, Chisel Tip, 16 Count - Whiteboard, Calendar, Organization, Back to School, Teacher Supplies
  • Dry erase markers with the most vibrant ink yet from EXPO
  • Vibrant ink makes it easier to read information from a distance
  • Made for the whiteboard and beyond, writing pops on most non-porous surfaces like glass, acrylic, and more!
  • Easily and cleanly erases with an EXPO eraser or dry cloth
  • Versatile chisel tip creates multiple line widths

Worked numerical example

Take x=[1,2]T, a ReLU hidden layer, and a linear output with target y=1:

W1=[[0.1,0.2],[0.3,0.4]], b1=[0.1,0.1]T, W2=[0.5,-0.4], b2=0.2.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Hidden affine value: z1 = [0.6,1.2]T.
  2. ReLU: a1=[0.6,1.2]T.
  3. Output: z2 = 0.5(0.6)-0.4(1.2)+0.2 = 0.02.
  4. Half-squared loss: L=½(0.02-1)2=0.4802.
  5. Output error: δ2=z2-y=-0.98.
  6. Output gradients: dW2=[-0.588,-1.176], db2=-0.98.
  7. Hidden error: δ1 = [0.5,-0.4]T(-0.98)=[-0.49,0.392]T; both ReLU inputs are positive, so their derivative is one.
  8. First-layer gradients: dW1=[[-0.49,-0.98],[0.392,0.784]], db1=[-0.49,0.392]T.

With learning rate η=0.1, one gradient-descent step gives W2=[0.5588,-0.2824], b2=0.298, W1=[[0.149,0.298],[0.2608,0.3216]], and b1=[0.149,0.0608]T. This update is an optimizer operation performed after backpropagation, not part of the backward recurrence itself.

Manual algorithm

forward:
    a[0] = x
    for l = 1 ... N:
        z[l] = W[l] @ a[l-1] + b[l]
        a[l] = activation[l](z[l])
    loss = loss_function(a[N], y)

backward:
    delta[N] = dloss/da[N] * activation_prime[N](z[N])
    for l = N ... 1:
        dW[l] = delta[l] @ a[l-1].T
        db[l] = delta[l]
        if l > 1:
            delta[l-1] = (W[l].T @ delta[l]) * activation_prime[l-1](z[l-1])

update:
    W[l] -= learning_rate * dW[l]
    b[l] -= learning_rate * db[l]

For a batch, sum or average the per-example deltas according to the loss reduction. If Lbatch=(1/m)∑rLr, its gradient is the average of individual gradients; a summed loss is larger by a factor of m.

Activation derivatives and gradient flow

Activation Derivative Typical issue
Sigmoid σ(z)(1-σ(z)) Saturates; gradients can vanish
Tanh 1-tanh2(z) Saturates at large magnitude
ReLU 1 for z>0, 0 for z<0 Dead units; undefined classical derivative at zero
Leaky ReLU 1 or α Reduces, but does not eliminate, dead-unit risk
Softmax Coupled Jacobian Not an elementwise scalar derivative

Across many layers, derivatives multiply. Repeated factors below one produce vanishing gradients; factors above one can produce exploding gradients. Initialization, normalization, architecture, activation choice, and sometimes gradient clipping affect these dynamics. Clipping changes the effective gradient, so it is not a universal cure.

Backpropagation versus automatic differentiation

Automatic differentiation (AD) applies the chain rule to programs built from differentiable operations. Forward-mode AD propagates tangents from inputs to outputs; reverse-mode AD propagates sensitivities from outputs back to inputs. Backpropagation is reverse-mode AD specialized to neural-network loss functions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Sale
EXPO Dry Erase Markers, Low Odor Ink, Assorted Fashion Colors, Chisel Tip, 36 Count - Easily Erases, Ideal for Classroom, Home, Office, Back to School, Teacher Supplies
  • Versatile Chisel Tip: For broad, medium, or fine lines
  • Low-Odor Ink: Ideal for classrooms, offices, and home use
  • Multipurpose: Suitable for use on whiteboards and most non-porous surfaces
  • Vivid & Quick Drying: Bold color that is easy to erase and see from a distance
  • Pack Includes: 36 assorted color dry erase markers
Method Best use Limitation
Forward-mode AD Few inputs, many outputs Expensive when parameters greatly outnumber outputs
Reverse-mode AD Scalar loss, many parameters Stores or recomputes intermediates
Backpropagation Training differentiable networks Memory and differentiability constraints
Finite differences Gradient checking Slow and step-size sensitive
Symbolic differentiation Algebraic analysis Expression growth and poor scalability

AD avoids finite-difference approximation for supported elementary operations, but floating-point arithmetic still introduces numerical error. PyTorch documents reverse-mode autograd and a forward-mode API (autograd reference).

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Why the forward pass uses memory

Backward computation commonly needs inputs, preactivations, activations, ReLU masks, branch choices, normalization statistics, or other operation-specific state saved during the forward pass. Saving more values makes backward faster but increases memory. Recomputing or checkpointing selected activations saves memory at extra compute cost. TensorFlow documents tape-held intermediates and persistent-tape memory behavior (autodiff guide); PyTorch documents saved tensors and storage hooks (autograd reference).

Framework implementations

PyTorch

import torch

x = torch.tensor([[1.0, 2.0]])
y = torch.tensor([[1.0]])
W1 = torch.tensor([[0.1, 0.2], [0.3, 0.4]], requires_grad=True)
b1 = torch.tensor([0.1, 0.1], requires_grad=True)
W2 = torch.tensor([[0.5, -0.4]], requires_grad=True)
b2 = torch.tensor([0.2], requires_grad=True)

z1 = x @ W1.T + b1
a1 = torch.relu(z1)
z2 = a1 @ W2.T + b2
loss = 0.5 * (z2 - y).pow(2).mean()
loss.backward()
print(loss.item(), W1.grad, b1.grad, W2.grad, b2.grad)

Floating-point or complex tensors are required for ordinary gradients. PyTorch accumulates gradients when backward() runs, so clear them (for example with an optimizer’s zero_grad()) before the next step. Avoid in-place operations that overwrite values needed by backward; use torch.autograd.grad() for explicit gradients and torch.autograd.gradcheck() for numerical checks.

TensorFlow

import tensorflow as tf

x = tf.constant([[1.0, 2.0]])
y = tf.constant([[1.0]])
W1 = tf.Variable([[0.1, 0.2], [0.3, 0.4]])
b1 = tf.Variable([0.1, 0.1])
W2 = tf.Variable([[0.5, -0.4]])
b2 = tf.Variable([0.2])

with tf.GradientTape() as tape:
    z1 = tf.matmul(x, W1, transpose_b=True) + b1
    a1 = tf.nn.relu(z1)
    z2 = tf.matmul(a1, W2, transpose_b=True) + b2
    loss = 0.5 * tf.reduce_mean(tf.square(z2 - y))
grads = tape.gradient(loss, [W1, b1, W2, b2])

GradientTape watches trainable variables used inside its context. Use tape.watch() for ordinary tensors, and persistent=True only when multiple gradient calls are needed. NumPy conversions, integer or string tensors, stateful assignments, and nonscalar targets can interrupt or change differentiation; use jacobian() or an explicit upstream gradient for vector targets (advanced autodiff).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Gradient checking and troubleshooting

A central finite-difference check estimates one derivative as [L(θ+h)-L(θ-h)]/(2h). Compare it with backpropagation in double precision, choose a small but not tiny h, and test away from ReLU’s zero kink. Matching values within a sensible relative tolerance validates implementation details; finite differences are not a practical training method.

Symptom Likely cause Check or fix
Shape error or transposed gradient Mixed row/column conventions Write every tensor dimension before each multiplication.
Missing bias gradient Incorrectly multiplying bias by input Use db=δ per example; reduce across the batch.
Unexpected scale difference One loss is summed and the other averaged Check mean, sum, or unreduced loss.
All-zero ReLU gradient Unit is negative or at the kink Inspect preactivations and activation convention.
Exploding or vanishing values Repeated Jacobian factors, saturation, or poor initialization Inspect gradient norms; review initialization, normalization, activations, and learning rate.
Backward error after mutation In-place operation overwrote a saved value Use out-of-place operations and inspect the framework’s anomaly diagnostics.
No gradient Integer/discrete operation, detached tensor, or NumPy conversion Keep the path differentiable and verify tracking.
Unstable softmax loss Naive exponentials or logarithms Use stabilized logits-based cross-entropy.

Beyond dense feed-forward networks

The same local rule applies to convolutional layers, normalization, residual additions, and shared parameters; only the operation-specific derivative changes. In branches, gradients add at merges. Recurrent networks apply the same backward pass through an unrolled time graph, often with truncated backpropagation through time. Higher-order derivatives differentiate the backward computation itself, while custom operations require a correctly defined backward rule.

The complete picture

The essential pipeline is

forward pass → loss → backward pass → gradients → optimizer update.

Forward equations tell you what the network computes. The chain rule explains how each parameter influences the loss. Transposes and outer products follow from matrix dimensions, not memorization. Once loss reduction, activation derivatives, saved intermediates, and tensor shapes are explicit, manual derivations and framework autograd results can be checked against one another.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Quick Recap

Bestseller No. 1
EXPO Dry Erase Markers Kit, Chisel Tip, Assorted Colors, Eraser, Spray Cleaner, 6 Count - Whiteboard, Calendar, Office Essentials, School, Classroom, Teacher Supplies
EXPO Dry Erase Markers Kit, Chisel Tip, Assorted Colors, Eraser, Spray Cleaner, 6 Count - Whiteboard, Calendar, Office Essentials, School, Classroom, Teacher Supplies
Dry erase markers with the most vibrant ink yet from EXPO; Vibrant ink makes it easier to read information from a distance
$7.57
SaleBestseller No. 2
EXPO Dry Erase Markers, Low Odor Ink, Assorted Colors, Chisel Tip, 12 Count
EXPO Dry Erase Markers, Low Odor Ink, Assorted Colors, Chisel Tip, 12 Count
Dry erase markers with the most vibrant ink yet from EXPO; Vibrant ink makes it easier to read information from a distance
$8.52
Bestseller No. 3
EXPO Dry Erase Markers Kit, Fine and Chisel Tip Markers, Assorted Colors, Eraser, Spray Cleaner, 14 Count
EXPO Dry Erase Markers Kit, Fine and Chisel Tip Markers, Assorted Colors, Eraser, Spray Cleaner, 14 Count
EXPO kit comes with everything you need to start marking and keep your surfaces clean; Versatile chisel tip allows for broad and fine writing. Fine tip is great for details
$18.38
SaleBestseller No. 4
EXPO Dry Erase Markers, Low Odor Ink, Assorted Colors, Chisel Tip, 16 Count - Whiteboard, Calendar, Organization, Back to School, Teacher Supplies
EXPO Dry Erase Markers, Low Odor Ink, Assorted Colors, Chisel Tip, 16 Count - Whiteboard, Calendar, Organization, Back to School, Teacher Supplies
Dry erase markers with the most vibrant ink yet from EXPO; Vibrant ink makes it easier to read information from a distance
$8.99
SaleBestseller No. 5
EXPO Dry Erase Markers, Low Odor Ink, Assorted Fashion Colors, Chisel Tip, 36 Count - Easily Erases, Ideal for Classroom, Home, Office, Back to School, Teacher Supplies
EXPO Dry Erase Markers, Low Odor Ink, Assorted Fashion Colors, Chisel Tip, 36 Count - Easily Erases, Ideal for Classroom, Home, Office, Back to School, Teacher Supplies
Versatile Chisel Tip: For broad, medium, or fine lines; Low-Odor Ink: Ideal for classrooms, offices, and home use
$22.49

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. Shenzhen desk3 min
    HONOR Expands Beyond Smartphones With Humanoid Robot RevealHONOR said it unveiled its first humanoid robot at MWC 2026 and named shopping assistance, workplace inspections, and supportive companionship as intended uses. Later Robotics D1 claims and a reported…
  2. Cupertino desk5 min
    Apple Unveils AirPods Max 2: The Upgrade That Should Have Happened Years AgoAirPods Max 2 adds H2-powered audio features and Apple claims up to 1.5× more effective ANC, but its design, Smart Case, and 20-hour battery rating are unchanged. Wired lossless audio…
  3. Cupertino desk4 min
    Apple’s OLED Touch MacBooks Are Coming—but the Dynamic Island Is the Real GambleApple has not announced an OLED touchscreen MacBook, but reports point to high-end models arriving in late 2026 or early 2027. The reported Mac Dynamic Island could be useful, but…
Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.