Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

An autoencoder is a neural network that learns to reconstruct its input. An encoder maps an input x to a latent representation z, and a decoder maps z back to a reconstruction x̂. Training minimizes reconstruction error. In this guide, you will build a dense autoencoder for Fashion-MNIST with Keras, inspect its latent vectors and errors, then adapt the workflow to convolutional denoising and anomaly detection.

What an autoencoder actually learns

The basic computation is:

z = fθ(x)
x̂ = gφ(z)

The encoder, f, produces the latent code; the decoder, g, reconstructs the original feature space. In a standard autoencoder the input is also the target:

model.fit(x_train, x_train)

This is more precisely self-supervised reconstruction than “learning without targets.” The bottleneck, architecture, noise, regularization and loss determine whether the latent code contains useful structure. A large, unconstrained network can learn an almost-identity mapping instead of a meaningful compact representation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose the right autoencoder variant

Variant Objective Typical use
Dense autoencoder Reconstruct vectors or small flattened inputs Learning the basic workflow and embeddings
Convolutional autoencoder Reconstruct images while preserving spatial locality Image reconstruction and denoising
Denoising autoencoder Map corrupted inputs to clean targets Noise removal and robust features
Sparse autoencoder Reconstruct while encouraging sparse activations Feature discovery
Variational autoencoder (VAE) Reconstruct while regularizing a probability distribution in latent space Structured latent modeling and generation
Anomaly-detection autoencoder Reconstruct mostly normal data and score the error Novelty or fault screening

Autoencoders are a poor substitute for every problem. PCA is often a better first baseline for linear dimensionality reduction, and a supervised classifier is usually preferable when labeled examples and a classification objective are available. Good reconstructions also do not guarantee useful embeddings or a reliable anomaly detector.

Prerequisites and environment

  • Python, NumPy and basic plotting
  • Train, validation and test-set concepts
  • Neural-network layers, activations, losses, gradients, epochs and batches
  • A local virtual environment or hosted notebook

Fashion-MNIST is small enough for a CPU, although a GPU helps with larger convolutional or high-resolution datasets. Create an isolated environment with python -m venv .venv, activate it with source .venv/bin/activate on macOS/Linux or .venvScriptsActivate.ps1 in Windows PowerShell, then follow the official installation instructions for your operating system and accelerator. Pin and record the Python and framework versions used for a reproducible project. PyTorch users can follow the current workflow covering tensors, data loaders, models, autograd and optimization at the official beginner guide.

Load and prepare Fashion-MNIST

Fashion-MNIST contains 60,000 training and 10,000 test grayscale images, each 28×28 pixels, in the TensorFlow tutorial’s workflow (TensorFlow autoencoder tutorial). Labels are not needed to train a basic reconstruction model, but retaining them lets you compare error by clothing class.

import numpy as np
import keras
from keras import layers

(x_train, y_train), (x_test, y_test) = keras.datasets.fashion_mnist.load_data()

x_train = x_train.astype("float32") / 255.0
x_test = x_test.astype("float32") / 255.0

# Dense layers expect one vector per image.
x_train = x_train.reshape((len(x_train), -1))
x_test = x_test.reshape((len(x_test), -1))

Keep preprocessing identical at training and inference time. For a convolutional model, retain spatial dimensions and add a channel axis instead:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
x_train = x_train[..., None]
x_test = x_test[..., None]

Because these targets are scaled to [0, 1], a sigmoid decoder output is a reasonable default. Continuous, unbounded targets generally call for a linear output and a loss matched to their scale.

Rank #2
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning

Build the smallest working dense model

This baseline compresses each 784-pixel image to 64 values and expands it again:

input_dim = x_train.shape[1]
latent_dim = 64

inputs = keras.Input(shape=(input_dim,))
encoded = layers.Dense(latent_dim, activation="relu")(inputs)
decoded = layers.Dense(input_dim, activation="sigmoid")(encoded)

autoencoder = keras.Model(inputs, decoded)
encoder = keras.Model(inputs, encoded)

autoencoder.compile(optimizer="adam", loss="binary_crossentropy")

Binary cross-entropy is common when normalized pixels are treated as Bernoulli-like values and follows the introductory TensorFlow example. Mean squared error (MSE) penalizes large pixel deviations more strongly; mean absolute error (MAE) is less sensitive to outliers. Choose among them according to target scaling and the behavior you need, rather than assuming one is universally correct:

autoencoder.compile(optimizer="adam", loss="mse")

Train with a validation protocol

Use a validation split for tuning and keep the test set for final evaluation. The values below are illustrative, not universal hyperparameters.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
history = autoencoder.fit(
    x_train,
    x_train,
    epochs=50,
    batch_size=256,
    shuffle=True,
    validation_split=0.1,
    callbacks=[
        keras.callbacks.EarlyStopping(
            monitor="val_loss",
            patience=5,
            restore_best_weights=True,
        )
    ],
)

Plot both training and validation loss. A widening gap suggests overfitting; two high curves suggest insufficient capacity, an overly severe bottleneck or a preprocessing problem. Fix a random seed when comparing experiments, and record the latent dimension, loss, preprocessing and framework version with each run.

Inspect reconstructions and error

Always combine visual and numerical checks. Reconstruct held-out images and reshape them for display:

reconstructed = autoencoder.predict(x_test[:10], verbose=0)
original_images = x_test[:10].reshape(-1, 28, 28)
reconstructed_images = reconstructed.reshape(-1, 28, 28)
difference_images = np.abs(original_images - reconstructed_images)

Display each original, reconstruction and absolute-difference image in a three-row grid. Then calculate one error value per image:

errors = np.mean(np.square(x_test - reconstructed), axis=1)

For a convolutional tensor, reduce across every non-batch axis:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
errors = np.mean(
    np.square(x_test - reconstructed),
    axis=tuple(range(1, x_test.ndim)),
)

Overall validation loss averages all pixels and examples. Per-image error exposes difficult cases; class-specific error can reveal that a model performs poorly on one clothing category. Examine typical and worst examples rather than reporting only a single mean.

Explore the latent representation

latent_vectors = encoder.predict(x_test, verbose=0)
print(latent_vectors.shape)  # (number_of_images, 64)

A two-dimensional bottleneck can be plotted directly and colored by Fashion-MNIST label. A 64-dimensional code requires another dimensionality-reduction method for visualization, which adds a second modeling step. Standard autoencoder coordinates are not guaranteed to be semantic: they may rotate, rescale or reorganize between runs. A smaller bottleneck increases compression and information loss; a larger one usually improves reconstruction while weakening the compression constraint.

Use convolutional layers for images

Flattening discards explicit spatial locality. A convolutional encoder and decoder usually provide a better image inductive bias:

inputs = keras.Input(shape=(28, 28, 1))

x = layers.Conv2D(16, 3, activation="relu", padding="same", strides=2)(inputs)
x = layers.Conv2D(8, 3, activation="relu", padding="same", strides=2)(x)
x = layers.Conv2DTranspose(8, 3, activation="relu", padding="same", strides=2)(x)
x = layers.Conv2DTranspose(16, 3, activation="relu", padding="same", strides=2)(x)
outputs = layers.Conv2D(1, 3, activation="sigmoid", padding="same")(x)

conv_autoencoder = keras.Model(inputs, outputs)
conv_autoencoder.compile(optimizer="adam", loss="mse")

This pattern follows the image-denoising approach in Keras’s convolutional autoencoder example. Print or inspect every intermediate shape before a long run. Downsampling odd dimensions, mismatched channel counts or padding choices can leave the decoder one pixel larger or smaller than its target. Transposed convolutions can also produce checkerboard artifacts; alternative upsampling-plus-convolution designs are worth testing when that occurs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Train a denoising autoencoder

A denoising model receives a corrupted image but is trained against the clean image:

noise_factor = 0.2

x_train_noisy = x_train + noise_factor * np.random.normal(
    0.0, 1.0, size=x_train.shape
)
x_test_noisy = x_test + noise_factor * np.random.normal(
    0.0, 1.0, size=x_test.shape
)

x_train_noisy = np.clip(x_train_noisy, 0.0, 1.0)
x_test_noisy = np.clip(x_test_noisy, 0.0, 1.0)

autoencoder.fit(
    x_train_noisy,
    x_train,
    epochs=20,
    batch_size=256,
    validation_data=(x_test_noisy, x_test),
)

The same input/target distinction applies to convolutional tensors. Gaussian noise is only one corruption model; deployment may involve blur, missing pixels, compression artifacts, salt-and-pepper noise or sensor-specific noise. The network learns the conditional reconstruction favored by its training distribution and loss, not a guaranteed historically “true” image.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Use reconstruction error for anomaly detection

  1. Train using normal examples, excluding known anomalies.
  2. Measure reconstruction errors on a representative normal validation set.
  3. Select a threshold without using the final test set for tuning.
  4. Apply the threshold to future data and report precision, recall and false-positive rate.

An instructional ECG example in TensorFlow’s tutorial uses normal rhythms and a mean-plus-one-standard-deviation threshold. That rule is not universal:

normal_reconstructions = autoencoder.predict(normal_train_data, verbose=0)
normal_errors = np.mean(
    np.abs(normal_reconstructions - normal_train_data),
    axis=1,
)
threshold = normal_errors.mean() + normal_errors.std()

Thresholds must reflect the cost of false alarms and missed detections. Contaminated training data, distribution drift, subgroup-specific normal behavior, temporal dependence and a powerful decoder that reconstructs anomalies well can all invalidate a fixed cutoff. Recalibrate on a representative validation period or compare against supervised and classical anomaly-detection baselines.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How a VAE differs

A variational autoencoder does not assign each input one deterministic code. Its encoder estimates a mean and log variance, samples a latent vector, and trains with reconstruction loss plus a KL-divergence penalty:

L = Lreconstruction + βDKL(qφ(z|x) || p(z))

The probabilistic constraint makes sampling and structured latent spaces possible, but may trade reconstruction sharpness for regularity. Keras’s VAE example demonstrates the sampling layer and custom training objective. Monitor reconstruction and KL terms separately; posterior collapse occurs when a powerful decoder ignores the latent variable. KL-weight schedules, a smaller decoder or a different latent size may help.

Compact PyTorch translation

The same design transfers directly to PyTorch:

import torch
from torch import nn

class Autoencoder(nn.Module):
    def __init__(self, input_dim, latent_dim=64):
        super().__init__()
        self.encoder = nn.Sequential(
            nn.Linear(input_dim, latent_dim), nn.ReLU()
        )
        self.decoder = nn.Sequential(
            nn.Linear(latent_dim, input_dim), nn.Sigmoid()
        )

    def forward(self, x):
        z = self.encoder(x)
        return self.decoder(z)

model = Autoencoder(input_dim=x_train.shape[1])
optimizer = torch.optim.Adam(model.parameters())
criterion = nn.MSELoss()

for epoch in range(epochs):
    model.train()
    for batch_x, _ in train_loader:
        optimizer.zero_grad()
        reconstruction = model(batch_x)
        loss = criterion(reconstruction, batch_x)
        loss.backward()
        optimizer.step()

Use PyTorch’s official beginner workflow and optimization tutorial for data loading, device management and evaluation. This is a framework translation, not a second fully tested dataset pipeline.

Troubleshooting checklist

  • Shape mismatch: Print every tensor shape, test one batch, and keep image height, width and channels in one source of truth.
  • Wrong output range: Match sigmoid outputs to [0, 1] targets, or use a linear output for unconstrained values.
  • Identity mapping: Reduce latent size or decoder capacity; add masking, noise, sparsity or weight regularization.
  • Blurry results: MSE often averages plausible details. Try MAE, a convolutional architecture or a task-specific perceptual objective, and judge accuracy rather than sharpness alone.
  • Overfitting: Use a validation split, early stopping, regularization and more representative data.
  • Unstable anomaly scores: Check drift, subgroup distributions, temporal structure and threshold selection; never fit production preprocessing separately from training preprocessing.

Practical decision checklist

  • Define the reconstruction target and its numeric range.
  • Compare against PCA or another simple baseline.
  • Choose dense layers for small vectors and convolutions for images.
  • Select latent size using validation results and the downstream task, not reconstruction loss alone.
  • Keep test data separate from hyperparameter and threshold tuning.
  • Inspect curves, error distributions, class-specific results and failure examples.
  • Save preprocessing parameters, model weights, framework versions and random seeds together.

An autoencoder is valuable when its constraint and evaluation match the problem: a bottleneck for representation learning, corruption for denoising, spatial inductive bias for images or a carefully validated error threshold for novelty screening. It is not automatically a compressor, generator or anomaly detector simply because it reconstructs its inputs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.