Forward propagation computes a neural network’s prediction; backpropagation computes how the resulting loss changes with respect to every parameter. An optimizer such as SGD or Adam then uses those gradients to update weights and biases. Keeping these three stages separate makes the equations, implementations, and debugging much easier to understand.
What problem does backpropagation solve?
A network may contain millions of parameters, yet training needs the gradient of one scalar loss with respect to all of them. Differentiating each parameter independently would repeat the same downstream calculations. Backpropagation reuses intermediate derivatives by traversing the computation graph in reverse, accumulating the complete gradient efficiently.
Backpropagation is a gradient-computation algorithm. It does not select a learning rate, choose an optimizer, or update parameters by itself:
- Forward pass: evaluate activations, prediction, and loss.
- Backward pass: compute derivatives of the loss with respect to activations, weights, and biases.
- Optimizer step: apply an update such as
W ← W - η∇W.
Rumelhart, Hinton, and Williams popularized a practical multilayer-network formulation in 1986, although related automatic-differentiation and gradient-propagation work predates that paper (Nature, 1986; historical context at Harvard CS249r).
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
- Dry erase markers with the most vibrant ink yet from EXPO
- Vibrant ink makes it easier to read information from a distance
- Made for the whiteboard and beyond, writing pops on most non-porous surfaces like glass, acrylic, and more!
- Easily and cleanly erases with included EXPO eraser and cleaner spray
- Versatile chisel tip creates multiple line widths
Minimum mathematics: derivatives, gradients, and Jacobians
Scalar derivative
For y = f(x), dy/dx measures the local change in y caused by a change in x.
Gradient
For a scalar loss L(w) and parameter vector w ∈ Rn,
∇wL = [∂L/∂w1, …, ∂L/∂wn]T.
The gradient points toward steepest local increase; its negative is the direction used by basic gradient descent.
Jacobian and the vector chain rule
For y = f(x), where x ∈ Rn and y ∈ Rm, the Jacobian Jf is an m × n matrix. Composition obeys
Jg○f(x) = Jg(f(x))Jf(x).
Backpropagation applies this rule without explicitly constructing every large Jacobian. In differential notation, dL = (∂L/∂x)Tdx; this form makes transpose placement and tensor shapes easier to check.
One neuron: the chain rule in full
A neuron first forms an affine value and then applies an activation:
Rank #2
- Dry erase markers with the most vibrant ink yet from EXPO
- Vibrant ink makes it easier to read information from a distance
- Made for the whiteboard and beyond, writing pops on most non-porous surfaces like glass, acrylic, and more!
- Easily and cleanly erases with an EXPO eraser or dry cloth
- Versatile chisel tip creates multiple line widths
z = wTx + ba = f(z).
For one weight wi and loss L(a,y),
∂L/∂wi = (∂L/∂a) f′(z) xi.
Likewise, ∂L/∂b = (∂L/∂a)f′(z). A weight’s gradient is therefore the product of three factors: downstream loss sensitivity, the neuron’s local activation slope, and the input connected to that weight.
Forward propagation through layers
Use column vectors for one example. For layer l:
z(l) = W(l)a(l-1) + b(l)a(l) = f(l)(z(l)), with a(0) = x.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →| Quantity | Shape | Meaning |
|---|---|---|
W(l) |
nl × nl-1 |
Weights from the previous layer |
a(l-1) |
nl-1 |
Previous activations |
b(l), z(l), a(l) |
nl |
Bias, preactivation, and postactivation |
For a batch stored as rows, a common convention is A(l-1) ∈ Rm×nl-1 and Z(l) = A(l-1)(W(l))T + b(l), with bias broadcasting across rows. Row-batch and column-example conventions are both valid; apparent transpose disagreements usually come from switching conventions.
Loss functions determine the starting gradient
Mean-squared error
For one scalar prediction, L = ½(ŷ-y)2, so ∂L/∂ŷ = ŷ-y. The factor one-half cancels the derivative’s factor of two.
Sigmoid plus binary cross-entropy
With ŷ = σ(z) = 1/(1+e-z) and L = -[y log ŷ + (1-y)log(1-ŷ)], the combined derivative is
∂L/∂z = ŷ-y.
This simplification is why frameworks provide fused sigmoid-cross-entropy losses.
Rank #3
- EXPO kit comes with everything you need to start marking and keep your surfaces clean
- Consistent, skip-free writing, vibrant color options and low-odor ink make the kit perfect for classrooms and offices
- Versatile chisel tip allows for broad and fine writing. Fine tip is great for details
- Spray and Expo eraser help you erase cleanly and easily while also extending whiteboard life
- 14-piece set includes fine and chisel tip markers in Black, Red, Blue, Green, Orange, Brown, Purple & Lime plus an 8 oz. bottle of Expo white board cleaning spray & an Expo eraser
Softmax plus multiclass cross-entropy
softmax(z)i = ezi/∑jezj. For one-hot target y and L = -∑iyilog ŷi,
∂L/∂z = ŷ-y.
Naively exponentiating large logits can overflow, so use a numerically stabilized softmax or a framework loss that accepts logits directly.
Deriving the backpropagation recurrence
Define the error signal (also called an adjoint)
δ(l) = ∂L/∂z(l).
Output layer
δ(N) = (∂L/∂a(N)) ⊙ f′(N)(z(N)). For sigmoid with binary cross-entropy and softmax with cross-entropy, this often reduces to prediction minus target.
Hidden layers
Because z(l+1) = W(l+1)a(l) + b(l+1),
δ(l) = ((W(l+1))Tδ(l+1)) ⊙ f′(l)(z(l)).
- The transposed matrix distributes downstream sensitivities to earlier activations.
- The activation derivative applies each neuron’s local slope.
⊙is elementwise multiplication.
Weight and bias gradients
For z(l) = W(l)a(l-1) + b(l),
∂L/∂W(l) = δ(l)(a(l-1))T∂L/∂b(l) = δ(l).
Elementwise, ∂L/∂Wij(l) = δi(l)aj(l-1), which is exactly the outer product. The bias derivative has no input multiplier because ∂zi/∂bi = 1.
Computational graphs, branches, and transposes
The network is a directed acyclic graph: x → z(1) → a(1) → … → prediction → L. Backward propagation starts with ∂L/∂L = 1. For a node v = g(u1,…,uk),
∂L/∂ui = (∂L/∂v)(∂v/∂ui).
If several paths reach the same variable, contributions add: ∂L/∂u = ∑r(∂L/∂vr)(∂vr/∂u). This matters for residual connections, branches, shared parameters, and recurrent unrolling. PyTorch describes reverse automatic differentiation over a dynamically built graph (autograd notes); TensorFlow records operations with GradientTape (API reference).
Rank #4
- Dry erase markers with the most vibrant ink yet from EXPO
- Vibrant ink makes it easier to read information from a distance
- Made for the whiteboard and beyond, writing pops on most non-porous surfaces like glass, acrylic, and more!
- Easily and cleanly erases with an EXPO eraser or dry cloth
- Versatile chisel tip creates multiple line widths
Worked numerical example
Take x=[1,2]T, a ReLU hidden layer, and a linear output with target y=1:
W1=[[0.1,0.2],[0.3,0.4]], b1=[0.1,0.1]T, W2=[0.5,-0.4], b2=0.2.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problems- Hidden affine value:
z1 = [0.6,1.2]T. - ReLU:
a1=[0.6,1.2]T. - Output:
z2 = 0.5(0.6)-0.4(1.2)+0.2 = 0.02. - Half-squared loss:
L=½(0.02-1)2=0.4802. - Output error:
δ2=z2-y=-0.98. - Output gradients:
dW2=[-0.588,-1.176],db2=-0.98. - Hidden error:
δ1 = [0.5,-0.4]T(-0.98)=[-0.49,0.392]T; both ReLU inputs are positive, so their derivative is one. - First-layer gradients:
dW1=[[-0.49,-0.98],[0.392,0.784]],db1=[-0.49,0.392]T.
With learning rate η=0.1, one gradient-descent step gives W2=[0.5588,-0.2824], b2=0.298, W1=[[0.149,0.298],[0.2608,0.3216]], and b1=[0.149,0.0608]T. This update is an optimizer operation performed after backpropagation, not part of the backward recurrence itself.
Manual algorithm
forward:
a[0] = x
for l = 1 ... N:
z[l] = W[l] @ a[l-1] + b[l]
a[l] = activation[l](z[l])
loss = loss_function(a[N], y)
backward:
delta[N] = dloss/da[N] * activation_prime[N](z[N])
for l = N ... 1:
dW[l] = delta[l] @ a[l-1].T
db[l] = delta[l]
if l > 1:
delta[l-1] = (W[l].T @ delta[l]) * activation_prime[l-1](z[l-1])
update:
W[l] -= learning_rate * dW[l]
b[l] -= learning_rate * db[l]
For a batch, sum or average the per-example deltas according to the loss reduction. If Lbatch=(1/m)∑rLr, its gradient is the average of individual gradients; a summed loss is larger by a factor of m.
Activation derivatives and gradient flow
| Activation | Derivative | Typical issue |
|---|---|---|
| Sigmoid | σ(z)(1-σ(z)) |
Saturates; gradients can vanish |
| Tanh | 1-tanh2(z) |
Saturates at large magnitude |
| ReLU | 1 for z>0, 0 for z<0 |
Dead units; undefined classical derivative at zero |
| Leaky ReLU | 1 or α |
Reduces, but does not eliminate, dead-unit risk |
| Softmax | Coupled Jacobian | Not an elementwise scalar derivative |
Across many layers, derivatives multiply. Repeated factors below one produce vanishing gradients; factors above one can produce exploding gradients. Initialization, normalization, architecture, activation choice, and sometimes gradient clipping affect these dynamics. Clipping changes the effective gradient, so it is not a universal cure.
Backpropagation versus automatic differentiation
Automatic differentiation (AD) applies the chain rule to programs built from differentiable operations. Forward-mode AD propagates tangents from inputs to outputs; reverse-mode AD propagates sensitivities from outputs back to inputs. Backpropagation is reverse-mode AD specialized to neural-network loss functions.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Best Value
- Versatile Chisel Tip: For broad, medium, or fine lines
- Low-Odor Ink: Ideal for classrooms, offices, and home use
- Multipurpose: Suitable for use on whiteboards and most non-porous surfaces
- Vivid & Quick Drying: Bold color that is easy to erase and see from a distance
- Pack Includes: 36 assorted color dry erase markers
| Method | Best use | Limitation |
|---|---|---|
| Forward-mode AD | Few inputs, many outputs | Expensive when parameters greatly outnumber outputs |
| Reverse-mode AD | Scalar loss, many parameters | Stores or recomputes intermediates |
| Backpropagation | Training differentiable networks | Memory and differentiability constraints |
| Finite differences | Gradient checking | Slow and step-size sensitive |
| Symbolic differentiation | Algebraic analysis | Expression growth and poor scalability |
AD avoids finite-difference approximation for supported elementary operations, but floating-point arithmetic still introduces numerical error. PyTorch documents reverse-mode autograd and a forward-mode API (autograd reference).
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Why the forward pass uses memory
Backward computation commonly needs inputs, preactivations, activations, ReLU masks, branch choices, normalization statistics, or other operation-specific state saved during the forward pass. Saving more values makes backward faster but increases memory. Recomputing or checkpointing selected activations saves memory at extra compute cost. TensorFlow documents tape-held intermediates and persistent-tape memory behavior (autodiff guide); PyTorch documents saved tensors and storage hooks (autograd reference).
Framework implementations
PyTorch
import torch
x = torch.tensor([[1.0, 2.0]])
y = torch.tensor([[1.0]])
W1 = torch.tensor([[0.1, 0.2], [0.3, 0.4]], requires_grad=True)
b1 = torch.tensor([0.1, 0.1], requires_grad=True)
W2 = torch.tensor([[0.5, -0.4]], requires_grad=True)
b2 = torch.tensor([0.2], requires_grad=True)
z1 = x @ W1.T + b1
a1 = torch.relu(z1)
z2 = a1 @ W2.T + b2
loss = 0.5 * (z2 - y).pow(2).mean()
loss.backward()
print(loss.item(), W1.grad, b1.grad, W2.grad, b2.grad)
Floating-point or complex tensors are required for ordinary gradients. PyTorch accumulates gradients when backward() runs, so clear them (for example with an optimizer’s zero_grad()) before the next step. Avoid in-place operations that overwrite values needed by backward; use torch.autograd.grad() for explicit gradients and torch.autograd.gradcheck() for numerical checks.
TensorFlow
import tensorflow as tf
x = tf.constant([[1.0, 2.0]])
y = tf.constant([[1.0]])
W1 = tf.Variable([[0.1, 0.2], [0.3, 0.4]])
b1 = tf.Variable([0.1, 0.1])
W2 = tf.Variable([[0.5, -0.4]])
b2 = tf.Variable([0.2])
with tf.GradientTape() as tape:
z1 = tf.matmul(x, W1, transpose_b=True) + b1
a1 = tf.nn.relu(z1)
z2 = tf.matmul(a1, W2, transpose_b=True) + b2
loss = 0.5 * tf.reduce_mean(tf.square(z2 - y))
grads = tape.gradient(loss, [W1, b1, W2, b2])
GradientTape watches trainable variables used inside its context. Use tape.watch() for ordinary tensors, and persistent=True only when multiple gradient calls are needed. NumPy conversions, integer or string tensors, stateful assignments, and nonscalar targets can interrupt or change differentiation; use jacobian() or an explicit upstream gradient for vector targets (advanced autodiff).
Gradient checking and troubleshooting
A central finite-difference check estimates one derivative as [L(θ+h)-L(θ-h)]/(2h). Compare it with backpropagation in double precision, choose a small but not tiny h, and test away from ReLU’s zero kink. Matching values within a sensible relative tolerance validates implementation details; finite differences are not a practical training method.
| Symptom | Likely cause | Check or fix |
|---|---|---|
| Shape error or transposed gradient | Mixed row/column conventions | Write every tensor dimension before each multiplication. |
| Missing bias gradient | Incorrectly multiplying bias by input | Use db=δ per example; reduce across the batch. |
| Unexpected scale difference | One loss is summed and the other averaged | Check mean, sum, or unreduced loss. |
| All-zero ReLU gradient | Unit is negative or at the kink | Inspect preactivations and activation convention. |
| Exploding or vanishing values | Repeated Jacobian factors, saturation, or poor initialization | Inspect gradient norms; review initialization, normalization, activations, and learning rate. |
| Backward error after mutation | In-place operation overwrote a saved value | Use out-of-place operations and inspect the framework’s anomaly diagnostics. |
| No gradient | Integer/discrete operation, detached tensor, or NumPy conversion | Keep the path differentiable and verify tracking. |
| Unstable softmax loss | Naive exponentials or logarithms | Use stabilized logits-based cross-entropy. |
Beyond dense feed-forward networks
The same local rule applies to convolutional layers, normalization, residual additions, and shared parameters; only the operation-specific derivative changes. In branches, gradients add at merges. Recurrent networks apply the same backward pass through an unrolled time graph, often with truncated backpropagation through time. Higher-order derivatives differentiate the backward computation itself, while custom operations require a correctly defined backward rule.
The complete picture
The essential pipeline is
forward pass → loss → backward pass → gradients → optimizer update.
Forward equations tell you what the network computes. The chain rule explains how each parameter influences the loss. Transposes and outer products follow from matrix dimensions, not memorization. Once loss reduction, activation derivatives, saved intermediates, and tensor shapes are explicit, manual derivations and framework autograd results can be checked against one another.
Free tools Windows power users keep installed
One-click scans. No signup required.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




