Hardware FixRecommendedDevice not working? Your driver may be the problemCheck updates for common hardware issues.Fix DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PC×
Skip to content
World desk5 min

Build a Tiny Neural Network From Scratch in Python—No PyTorch

Build a small XOR neural network with Python built-ins and follow each step from weighted inputs to backpropagation and parameter updates.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

You can build and train a small neural network with Python’s built-in features alone: no PyTorch, NumPy, or other machine-learning library. This walkthrough implements a two-layer network for XOR, showing the forward pass, loss, backpropagation, and gradient-descent updates. It assumes you can already write basic Python; the Python Tutorial describes its intended readers as “programmers that are new to Python, not beginners who are new to programming.”

What this network will do

The network learns XOR: it outputs 1 when its two inputs differ and 0 when they match. The complete dataset is:

Input A Input B Target
0 0 0
0 1 1
1 0 1
1 1 0

A neuron multiplies each input by a weight, adds a bias, then applies an activation function. A network composes neurons in layers. This example has two input values, a hidden layer of two neurons, and one output neuron. The hidden layer’s nonlinearity lets the network represent XOR; a single linear decision boundary cannot separate its two classes.

The code uses lists, loops, and the math module only. Python’s documentation shows how nested lists can represent matrices, but this implementation keeps each neuron’s values explicit to make the calculations easier to follow. See Python’s data-structures documentation for nested-list examples.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Define the parameters and activation

Each hidden neuron has two input weights and a bias. The output neuron has two weights—one per hidden activation—and a bias. We’ll use the sigmoid function, which maps any real number to a value between 0 and 1, along with its derivative:

sigmoid(x) = 1 / (1 + exp(-x))

sigmoid'(x) = sigmoid(x) * (1 - sigmoid(x))

Here is a complete training script. The initial parameters are fixed rather than randomly generated so that the example is repeatable. Change them and the training behavior may change, too.

import math

# Hidden neuron weights: one list per neuron, one weight per input.
w1 = [[-0.4, 0.2], [0.3, -0.5]]
b1 = [0.1, -0.2]

# Output neuron weights: one per hidden neuron.
w2 = [0.2, -0.3]
b2 = [0.1]

xs = [[0.0, 0.0], [0.0, 1.0], [1.0, 0.0], [1.0, 1.0]]
ys = [0.0, 1.0, 1.0, 0.0]

def sigmoid(x):
    return 1.0 / (1.0 + math.exp(-x))

def forward(x):
    # Hidden layer: z = weighted sum + bias; h = sigmoid(z).
    z1 = [sum(w * value for w, value in zip(weights, x)) + bias
          for weights, bias in zip(w1, b1)]
    h = [sigmoid(z) for z in z1]

    # One output neuron.
    z2 = sum(w * value for w, value in zip(w2, h)) + b2[0]
    yhat = sigmoid(z2)
    return z1, h, z2, yhat

def mean_squared_error():
    return sum((forward(x)[3] - y) ** 2 for x, y in zip(xs, ys)) / len(xs)

learning_rate = 1.0
steps = 20_000

for step in range(steps):
    # Full-batch gradients: accumulate each example's contribution.
    gw1 = [[0.0, 0.0] for _ in range(2)]
    gb1 = [0.0, 0.0]
    gw2 = [0.0, 0.0]
    gb2 = 0.0

    for x, target in zip(xs, ys):
        z1, h, z2, yhat = forward(x)

        # For squared error (yhat - target)^2, including its factor of 2.
        delta2 = 2.0 * (yhat - target) * yhat * (1.0 - yhat)
        gb2 += delta2
        for j in range(2):
            gw2[j] += delta2 * h[j]

        # Chain the output error back through each hidden neuron.
        for j in range(2):
            delta1 = delta2 * w2[j] * h[j] * (1.0 - h[j])
            gb1[j] += delta1
            for k in range(2):
                gw1[j][k] += delta1 * x[k]

    # Average over the four examples, then take a gradient-descent step.
    n = len(xs)
    for j in range(2):
        for k in range(2):
            w1[j][k] -= learning_rate * gw1[j][k] / n
        b1[j] -= learning_rate * gb1[j] / n
        w2[j] -= learning_rate * gw2[j] / n
    b2[0] -= learning_rate * gb2 / n

print("MSE:", mean_squared_error())
for x in xs:
    print(x, "->", round(forward(x)[3], 3))

This is a teaching script rather than a claim of a measured result: it has not been executed here, so no particular final loss or prediction is reported. Run it in a Python interpreter to inspect the output on your own machine.

Follow one forward pass

The forward pass turns an input into a prediction by applying the same sequence of operations through both layers. With the initial values above and input [1, 0], the hidden layer computes:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Hidden neuron 1: z = (-0.4 × 1) + (0.2 × 0) + 0.1 = -0.3; its activation is approximately 0.426.
  • Hidden neuron 2: z = (0.3 × 1) + (-0.5 × 0) - 0.2 = 0.1; its activation is approximately 0.525.

The output neuron then computes z = (0.2 × 0.426) + (-0.3 × 0.525) + 0.1, or approximately 0.027. Applying sigmoid gives a prediction of about 0.507. That is the network’s initial estimate for a target of 1; training adjusts the parameters to reduce the error.

Calculate loss and backpropagate

The script uses mean squared error (MSE): it squares each prediction’s difference from its target and averages the four results. For a single example, the loss is (prediction - target)^2. Loss gives training a numerical measure of error; backpropagation calculates how each weight and bias contributes to that error.

For the output neuron, the chain rule gives the derivative of the example’s squared error with respect to its pre-activation:

delta_output = 2 * (prediction - target) * prediction * (1 - prediction)

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The first factor is the derivative of squared error with respect to the prediction. The remaining factors are the sigmoid derivative. An output weight’s gradient is this delta multiplied by the corresponding hidden activation; the output bias’s gradient is the delta itself.

For a hidden neuron, the error is passed back through the connected output weight and the hidden neuron’s sigmoid derivative:

delta_hidden = delta_output * output_weight * hidden_activation * (1 - hidden_activation)

Its input-weight gradient is delta_hidden * input_value, and its bias gradient is delta_hidden. The code accumulates these gradients across all four examples, averages them, then subtracts the result multiplied by the learning rate from each parameter. This is full-batch gradient descent.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Train it and inspect the result

The loop repeats the same cycle: compute predictions, calculate gradients, and update parameters. After the loop, the script prints the final average loss and one prediction per training input. If training is working, predictions for the two XOR-positive rows should move toward 1 and those for the other rows toward 0; exact values depend on the initial parameters and learning rate.

If the output barely changes, oscillates, or becomes unhelpful, adjust the learning rate or initialization and run again. A learning rate that is too large can overshoot useful parameter values; a very small one can make progress slow. This compact example does not guarantee convergence for every possible initialization or learning rate.

What “from scratch” does—and does not—mean

This version uses neither PyTorch nor NumPy. “Without PyTorch” alone would not rule out NumPy: those are different constraints. Using NumPy would make array arithmetic shorter, while explicit Python lists make each multiply, sum, and update visible. The trade-off is that manual list-based operations are verbose and provide no automatic shape checks.

A machine-learning framework automates much of this bookkeeping, including differentiation and optimization workflows, and provides tools for batching and hardware acceleration. Writing the operations yourself is useful for understanding them, but this tiny network is not evidence that hand-written Python is appropriate for large models or production workloads.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For readers who want a Python refresher, the Python Tutorial is aimed at people with programming experience; the Python Software Foundation says it helps to have an interpreter available for hands-on work.

Why XOR is a useful test

XOR is small enough to inspect completely, yet it exposes a key limitation of a single linear classifier: the positive examples occupy opposite corners of the input square, so one straight decision boundary cannot separate them from the negative examples. A hidden layer with nonlinear activations can transform the inputs into a representation that the output neuron can separate. University teaching materials use XOR to teach multilayer networks, backpropagation, and gradient descent, including the University of Göttingen’s deep-neural-networks course and the University of Tübingen’s Deep Learning curriculum.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.