Hardware FixRecommendedDevice not working? Your driver may be the problemCheck updates for common hardware issues.Fix DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix Now×
Skip to content
World desk6 min

Image Classification with Vision Transformer in Keras

A step-by-step walkthrough of Keras's Vision Transformer image classifier: how patches become tokens, what the CIFAR-100 example's settings and results actually mean, and how to load your own labeled images.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Yes. Keras’s official Vision Transformer example implements the full pipeline in ordinary Keras layers: patch extraction, patch encoding, Transformer blocks, and a classifier head. It trains from scratch on CIFAR-100, and its reported run reaches about 55% test accuracy after 100 epochs, which the example itself describes as not competitive on that dataset. Treat the from-scratch model as a way to learn the architecture and prototype on your own images. The much stronger results associated with the original Vision Transformer paper depended on large-scale pretraining.

How a Vision Transformer turns pixels into class scores

A Vision Transformer (ViT) does not scan an image with convolutional filters. It cuts the image into a sequence of patches, treats each patch like a token in a sentence, and lets self-attention relate the patches to one another before a final layer predicts a class. The Keras example, written by Khalid Salama, follows the design of the ViT model introduced by Alexey Dosovitskiy and coauthors.

Step 1: Split the image into patches

The example resizes each input to 72 by 72 pixels and cuts it into 6 by 6 patches. That produces 12 × 12 = 144 patches per image. Each RGB patch contains 6 × 6 × 3 = 108 values once flattened. The model therefore receives a sequence of 144 vectors rather than a grid.

This is also why the input size matters. The side length has to divide evenly by the patch size. 72 divides by 6 cleanly; a size such as 70 would leave a partial row and column of pixels that the patch extraction step typically discards.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning

Step 2: Project each patch and add position information

Each flattened patch passes through a dense projection into the embedding dimension, which is 64 in the example. Self-attention by itself does not know where a vector sits in the sequence, so the model adds a learned position embedding to each patch vector. Without it, a patch from the top-left corner would look identical to the same content in the bottom-right corner.

Step 3: Process the sequence with Transformer blocks

Each block applies layer normalization, multi-head self-attention, a residual connection, a second layer normalization, an MLP, and another residual connection. The example uses four attention heads and stacks eight of these blocks. In each block, every patch can attend to every other patch, which is how the model learns relationships across the whole image instead of only within a local window.

Step 4: Build the representation and classify

After the final block, the example normalizes the outputs, flattens them into one vector, and passes that vector through a dense classifier. For CIFAR-100 the output layer has 100 units, one per class.

This is a deliberate departure from the original paper. The paper prepends a learnable class embedding and classifies from it. The Keras example instead flattens all final patch outputs, and the page notes global average pooling as another way to aggregate them. Both choices are valid design options, but a model built this way is not a literal reproduction of the paper.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the example’s settings mean

The example’s configuration is a tutorial setup, not a set of defaults that will suit every dataset or hardware budget. The values below come from the example page.

Setting Value in the example What it controls
Dataset CIFAR-100: 50,000 training and 10,000 test images Data the model learns from and is scored on
Input size 72 × 72 pixels Resolution fed to the model
Patch size 6 × 6 pixels Produces 144 patches per image
Embedding dimension 64 Width of each patch vector through the network
Attention heads 4 Parallel attention patterns per block
Transformer layers 8 Depth of the encoder stack
Epochs 10 as a test value; 100 for real training, per the example How long the model trains

If you change the dataset, resolution, or patch size, recalculate the patch count and check that memory use stays within your hardware limits before committing to a long run.

Reading the reported results

The example reports two numbers for its from-scratch ViT, and one reference point for a convolutional baseline. Keep their conditions attached whenever you quote them.

Model or setup Reported figure Conditions and source
ViT trained from scratch About 55% test accuracy CIFAR-100, 100 epochs, Keras example page (last updated 2021)
ViT trained from scratch About 82% test top-5 accuracy Same run and source as above
ResNet50V2 trained from scratch About 67% accuracy Comparison figure stated in the same Keras example; not rerun here
ViT pretrained on JFT-300M, then fine-tuned Not stated in the Keras example The example names JFT-300M as the pretraining dataset behind the paper’s stronger results; the figures appear in the original paper by Dosovitskiy et al.

These numbers come from one tutorial configuration. They are not a general benchmark, and they should not be compared with results obtained under different splits, epoch counts, or augmentation settings.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Using your own labeled images

To train on your own data, Keras’s image_dataset_from_directory utility builds a labeled dataset from folders, with one subfolder per class. Its from-scratch image-classification example demonstrates loading JPEG files from disk with preprocessing and augmentation layers.

  1. Create one folder per class under a single root, for example data/train/cats/ and data/train/dogs/. Folder names become the labels.
  2. Load the data with the directory utility, setting image_size to the size your model expects and splitting off a validation set with a fixed seed.
  3. Print class_names and confirm the order matches your folders. Keras sorts class names alphabetically, so the index assigned to each class follows that order.
  4. Set the classifier’s output units to the number of folders.
  5. Add augmentation, such as random flips, crops, or small rotations, as preprocessing layers so they run only during training.
import keras

train_ds, val_ds = keras.utils.image_dataset_from_directory(
    "data/train",
    validation_split=0.2,
    subset="both",
    seed=123,
    image_size=(72, 72),
    batch_size=32,
)
print(train_ds.class_names)

Run the loader once before training and check a batch of images visually. Mislabeled or mis-sorted folders are a common source of results that look plausible but are wrong.

Choosing an approach

The sources present these as options to weigh, not as a ranking. Match the choice to your data size, whether pretrained weights are available for your task, your resolution, and your compute budget.

Option Where it fits Caveat
Train the Keras example model from scratch Learning the architecture, small experiments, prototyping on your own labels The reported accuracy is modest; pretraining is what drove the paper’s stronger results
Fine-tune a pretrained ViT Tasks where pretrained weights exist and you need higher accuracy Pretrained-weight fine-tuning is outside the example’s walkthrough, so you will need to adapt the code
Shifted patch tokenization and locality self-attention Small datasets, according to a separate Keras example It is a distinct variant, not the same architecture as the basic example
Flatten final outputs (example default) Simple, close to the tutorial code Keeps every patch output in the classifier input, which grows with patch count
Global average pooling A lighter aggregation step the example suggests Changes what the classifier sees; compare results on your own validation split
Learnable class embedding (original paper) Closest to the published design Requires changing the example’s representation step

Limits and troubleshooting

  • Low accuracy on small datasets. A ViT trained from scratch has little built-in image prior. Use a validation split, augmentation, and the locality-focused variant if your data is small.
  • Code that fails on a newer release. The example page was last updated in 2021. Keras and TensorFlow change their APIs, so confirm that imports and layer calls run on your installed version before debugging the model itself.
  • Input size error. If your image side length is not a multiple of the patch size, change the resolution or patch size so the division is exact.
  • Wrong predicted labels. Check class_names against your folder names and make sure the output layer size matches.
  • Out-of-memory errors. Reduce the batch size or the input resolution first. The sources do not provide hardware sizing guidance, so measure memory on a short run before starting a full schedule.

The Keras example is a sound foundation for understanding ViT in code. Its figures describe one configuration on CIFAR-100, and the path to strong accuracy on your own images runs through pretrained weights, careful data preparation, and version-checked code.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The Bottom Line

Use the Keras Vision Transformer example to learn how patches, position embeddings, Transformer blocks, and the classifier fit together, and to prototype on folder-based data. Do not expect its from-scratch CIFAR-100 result to represent what a Vision Transformer can achieve with large-scale pretraining.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.