PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteYou can build and train a small neural network with Python’s built-in features alone: no PyTorch, NumPy, or other machine-learning library. This walkthrough implements a two-layer network for XOR, showing the forward pass, loss, backpropagation, and gradient-descent updates. It assumes you can already write basic Python; the Python Tutorial describes its intended readers as “programmers that are new to Python, not beginners who are new to programming.”
What this network will do
The network learns XOR: it outputs 1 when its two inputs differ and 0 when they match. The complete dataset is:
| Input A | Input B | Target |
|---|---|---|
| 0 | 0 | 0 |
| 0 | 1 | 1 |
| 1 | 0 | 1 |
| 1 | 1 | 0 |
A neuron multiplies each input by a weight, adds a bias, then applies an activation function. A network composes neurons in layers. This example has two input values, a hidden layer of two neurons, and one output neuron. The hidden layer’s nonlinearity lets the network represent XOR; a single linear decision boundary cannot separate its two classes.
The code uses lists, loops, and the math module only. Python’s documentation shows how nested lists can represent matrices, but this implementation keeps each neuron’s values explicit to make the calculations easier to follow. See Python’s data-structures documentation for nested-list examples.
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
Define the parameters and activation
Each hidden neuron has two input weights and a bias. The output neuron has two weights—one per hidden activation—and a bias. We’ll use the sigmoid function, which maps any real number to a value between 0 and 1, along with its derivative:
sigmoid(x) = 1 / (1 + exp(-x))
sigmoid'(x) = sigmoid(x) * (1 - sigmoid(x))
Here is a complete training script. The initial parameters are fixed rather than randomly generated so that the example is repeatable. Change them and the training behavior may change, too.
import math
# Hidden neuron weights: one list per neuron, one weight per input.
w1 = [[-0.4, 0.2], [0.3, -0.5]]
b1 = [0.1, -0.2]
# Output neuron weights: one per hidden neuron.
w2 = [0.2, -0.3]
b2 = [0.1]
xs = [[0.0, 0.0], [0.0, 1.0], [1.0, 0.0], [1.0, 1.0]]
ys = [0.0, 1.0, 1.0, 0.0]
def sigmoid(x):
return 1.0 / (1.0 + math.exp(-x))
def forward(x):
# Hidden layer: z = weighted sum + bias; h = sigmoid(z).
z1 = [sum(w * value for w, value in zip(weights, x)) + bias
for weights, bias in zip(w1, b1)]
h = [sigmoid(z) for z in z1]
# One output neuron.
z2 = sum(w * value for w, value in zip(w2, h)) + b2[0]
yhat = sigmoid(z2)
return z1, h, z2, yhat
def mean_squared_error():
return sum((forward(x)[3] - y) ** 2 for x, y in zip(xs, ys)) / len(xs)
learning_rate = 1.0
steps = 20_000
for step in range(steps):
# Full-batch gradients: accumulate each example's contribution.
gw1 = [[0.0, 0.0] for _ in range(2)]
gb1 = [0.0, 0.0]
gw2 = [0.0, 0.0]
gb2 = 0.0
for x, target in zip(xs, ys):
z1, h, z2, yhat = forward(x)
# For squared error (yhat - target)^2, including its factor of 2.
delta2 = 2.0 * (yhat - target) * yhat * (1.0 - yhat)
gb2 += delta2
for j in range(2):
gw2[j] += delta2 * h[j]
# Chain the output error back through each hidden neuron.
for j in range(2):
delta1 = delta2 * w2[j] * h[j] * (1.0 - h[j])
gb1[j] += delta1
for k in range(2):
gw1[j][k] += delta1 * x[k]
# Average over the four examples, then take a gradient-descent step.
n = len(xs)
for j in range(2):
for k in range(2):
w1[j][k] -= learning_rate * gw1[j][k] / n
b1[j] -= learning_rate * gb1[j] / n
w2[j] -= learning_rate * gw2[j] / n
b2[0] -= learning_rate * gb2 / n
print("MSE:", mean_squared_error())
for x in xs:
print(x, "->", round(forward(x)[3], 3))
This is a teaching script rather than a claim of a measured result: it has not been executed here, so no particular final loss or prediction is reported. Run it in a Python interpreter to inspect the output on your own machine.
Rank #2
Follow one forward pass
The forward pass turns an input into a prediction by applying the same sequence of operations through both layers. With the initial values above and input [1, 0], the hidden layer computes:
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errors- Hidden neuron 1:
z = (-0.4 × 1) + (0.2 × 0) + 0.1 = -0.3; its activation is approximately0.426. - Hidden neuron 2:
z = (0.3 × 1) + (-0.5 × 0) - 0.2 = 0.1; its activation is approximately0.525.
The output neuron then computes z = (0.2 × 0.426) + (-0.3 × 0.525) + 0.1, or approximately 0.027. Applying sigmoid gives a prediction of about 0.507. That is the network’s initial estimate for a target of 1; training adjusts the parameters to reduce the error.
Calculate loss and backpropagate
The script uses mean squared error (MSE): it squares each prediction’s difference from its target and averages the four results. For a single example, the loss is (prediction - target)^2. Loss gives training a numerical measure of error; backpropagation calculates how each weight and bias contributes to that error.
For the output neuron, the chain rule gives the derivative of the example’s squared error with respect to its pre-activation:
delta_output = 2 * (prediction - target) * prediction * (1 - prediction)
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →The first factor is the derivative of squared error with respect to the prediction. The remaining factors are the sigmoid derivative. An output weight’s gradient is this delta multiplied by the corresponding hidden activation; the output bias’s gradient is the delta itself.
For a hidden neuron, the error is passed back through the connected output weight and the hidden neuron’s sigmoid derivative:
delta_hidden = delta_output * output_weight * hidden_activation * (1 - hidden_activation)
Its input-weight gradient is delta_hidden * input_value, and its bias gradient is delta_hidden. The code accumulates these gradients across all four examples, averages them, then subtracts the result multiplied by the learning rate from each parameter. This is full-batch gradient descent.
Best Value
Train it and inspect the result
The loop repeats the same cycle: compute predictions, calculate gradients, and update parameters. After the loop, the script prints the final average loss and one prediction per training input. If training is working, predictions for the two XOR-positive rows should move toward 1 and those for the other rows toward 0; exact values depend on the initial parameters and learning rate.
If the output barely changes, oscillates, or becomes unhelpful, adjust the learning rate or initialization and run again. A learning rate that is too large can overshoot useful parameter values; a very small one can make progress slow. This compact example does not guarantee convergence for every possible initialization or learning rate.
What “from scratch” does—and does not—mean
This version uses neither PyTorch nor NumPy. “Without PyTorch” alone would not rule out NumPy: those are different constraints. Using NumPy would make array arithmetic shorter, while explicit Python lists make each multiply, sum, and update visible. The trade-off is that manual list-based operations are verbose and provide no automatic shape checks.
A machine-learning framework automates much of this bookkeeping, including differentiation and optimization workflows, and provides tools for batching and hardware acceleration. Writing the operations yourself is useful for understanding them, but this tiny network is not evidence that hand-written Python is appropriate for large models or production workloads.
For readers who want a Python refresher, the Python Tutorial is aimed at people with programming experience; the Python Software Foundation says it helps to have an interpreter available for hands-on work.
Why XOR is a useful test
XOR is small enough to inspect completely, yet it exposes a key limitation of a single linear classifier: the positive examples occupy opposite corners of the input square, so one straight decision boundary cannot separate them from the negative examples. A hidden layer with nonlinear activations can transform the inputs into a representation that the output neuron can separate. University teaching materials use XOR to teach multilayer networks, backpropagation, and gradient descent, including the University of Göttingen’s deep-neural-networks course and the University of Tübingen’s Deep Learning curriculum.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




