Yes. Keras’s official Vision Transformer example implements the full pipeline in ordinary Keras layers: patch extraction, patch encoding, Transformer blocks, and a classifier head. It trains from scratch on CIFAR-100, and its reported run reaches about 55% test accuracy after 100 epochs, which the example itself describes as not competitive on that dataset. Treat the from-scratch model as a way to learn the architecture and prototype on your own images. The much stronger results associated with the original Vision Transformer paper depended on large-scale pretraining.
How a Vision Transformer turns pixels into class scores
A Vision Transformer (ViT) does not scan an image with convolutional filters. It cuts the image into a sequence of patches, treats each patch like a token in a sentence, and lets self-attention relate the patches to one another before a final layer predicts a class. The Keras example, written by Khalid Salama, follows the design of the ViT model introduced by Alexey Dosovitskiy and coauthors.
Step 1: Split the image into patches
The example resizes each input to 72 by 72 pixels and cuts it into 6 by 6 patches. That produces 12 × 12 = 144 patches per image. Each RGB patch contains 6 × 6 × 3 = 108 values once flattened. The model therefore receives a sequence of 144 vectors rather than a grid.
This is also why the input size matters. The side length has to divide evenly by the patch size. 72 divides by 6 cleanly; a size such as 70 would leave a partial row and column of pixels that the patch extraction step typically discards.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →#1 Best Overall
- Use scikit-learn to track an example ML project end to end
- Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
- Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
- Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
- Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
Step 2: Project each patch and add position information
Each flattened patch passes through a dense projection into the embedding dimension, which is 64 in the example. Self-attention by itself does not know where a vector sits in the sequence, so the model adds a learned position embedding to each patch vector. Without it, a patch from the top-left corner would look identical to the same content in the bottom-right corner.
Step 3: Process the sequence with Transformer blocks
Each block applies layer normalization, multi-head self-attention, a residual connection, a second layer normalization, an MLP, and another residual connection. The example uses four attention heads and stacks eight of these blocks. In each block, every patch can attend to every other patch, which is how the model learns relationships across the whole image instead of only within a local window.
Rank #2
Step 4: Build the representation and classify
After the final block, the example normalizes the outputs, flattens them into one vector, and passes that vector through a dense classifier. For CIFAR-100 the output layer has 100 units, one per class.
This is a deliberate departure from the original paper. The paper prepends a learnable class embedding and classifies from it. The Keras example instead flattens all final patch outputs, and the page notes global average pooling as another way to aggregate them. Both choices are valid design options, but a model built this way is not a literal reproduction of the paper.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11What the example’s settings mean
The example’s configuration is a tutorial setup, not a set of defaults that will suit every dataset or hardware budget. The values below come from the example page.
| Setting | Value in the example | What it controls |
|---|---|---|
| Dataset | CIFAR-100: 50,000 training and 10,000 test images | Data the model learns from and is scored on |
| Input size | 72 × 72 pixels | Resolution fed to the model |
| Patch size | 6 × 6 pixels | Produces 144 patches per image |
| Embedding dimension | 64 | Width of each patch vector through the network |
| Attention heads | 4 | Parallel attention patterns per block |
| Transformer layers | 8 | Depth of the encoder stack |
| Epochs | 10 as a test value; 100 for real training, per the example | How long the model trains |
If you change the dataset, resolution, or patch size, recalculate the patch count and check that memory use stays within your hardware limits before committing to a long run.
Rank #4
Reading the reported results
The example reports two numbers for its from-scratch ViT, and one reference point for a convolutional baseline. Keep their conditions attached whenever you quote them.
| Model or setup | Reported figure | Conditions and source |
|---|---|---|
| ViT trained from scratch | About 55% test accuracy | CIFAR-100, 100 epochs, Keras example page (last updated 2021) |
| ViT trained from scratch | About 82% test top-5 accuracy | Same run and source as above |
| ResNet50V2 trained from scratch | About 67% accuracy | Comparison figure stated in the same Keras example; not rerun here |
| ViT pretrained on JFT-300M, then fine-tuned | Not stated in the Keras example | The example names JFT-300M as the pretraining dataset behind the paper’s stronger results; the figures appear in the original paper by Dosovitskiy et al. |
These numbers come from one tutorial configuration. They are not a general benchmark, and they should not be compared with results obtained under different splits, epoch counts, or augmentation settings.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Best Value
Using your own labeled images
To train on your own data, Keras’s image_dataset_from_directory utility builds a labeled dataset from folders, with one subfolder per class. Its from-scratch image-classification example demonstrates loading JPEG files from disk with preprocessing and augmentation layers.
- Create one folder per class under a single root, for example
data/train/cats/anddata/train/dogs/. Folder names become the labels. - Load the data with the directory utility, setting
image_sizeto the size your model expects and splitting off a validation set with a fixed seed. - Print
class_namesand confirm the order matches your folders. Keras sorts class names alphabetically, so the index assigned to each class follows that order. - Set the classifier’s output units to the number of folders.
- Add augmentation, such as random flips, crops, or small rotations, as preprocessing layers so they run only during training.
import keras
train_ds, val_ds = keras.utils.image_dataset_from_directory(
"data/train",
validation_split=0.2,
subset="both",
seed=123,
image_size=(72, 72),
batch_size=32,
)
print(train_ds.class_names)
Run the loader once before training and check a batch of images visually. Mislabeled or mis-sorted folders are a common source of results that look plausible but are wrong.
Choosing an approach
The sources present these as options to weigh, not as a ranking. Match the choice to your data size, whether pretrained weights are available for your task, your resolution, and your compute budget.
| Option | Where it fits | Caveat |
|---|---|---|
| Train the Keras example model from scratch | Learning the architecture, small experiments, prototyping on your own labels | The reported accuracy is modest; pretraining is what drove the paper’s stronger results |
| Fine-tune a pretrained ViT | Tasks where pretrained weights exist and you need higher accuracy | Pretrained-weight fine-tuning is outside the example’s walkthrough, so you will need to adapt the code |
| Shifted patch tokenization and locality self-attention | Small datasets, according to a separate Keras example | It is a distinct variant, not the same architecture as the basic example |
| Flatten final outputs (example default) | Simple, close to the tutorial code | Keeps every patch output in the classifier input, which grows with patch count |
| Global average pooling | A lighter aggregation step the example suggests | Changes what the classifier sees; compare results on your own validation split |
| Learnable class embedding (original paper) | Closest to the published design | Requires changing the example’s representation step |
Limits and troubleshooting
- Low accuracy on small datasets. A ViT trained from scratch has little built-in image prior. Use a validation split, augmentation, and the locality-focused variant if your data is small.
- Code that fails on a newer release. The example page was last updated in 2021. Keras and TensorFlow change their APIs, so confirm that imports and layer calls run on your installed version before debugging the model itself.
- Input size error. If your image side length is not a multiple of the patch size, change the resolution or patch size so the division is exact.
- Wrong predicted labels. Check
class_namesagainst your folder names and make sure the output layer size matches. - Out-of-memory errors. Reduce the batch size or the input resolution first. The sources do not provide hardware sizing guidance, so measure memory on a short run before starting a full schedule.
The Keras example is a sound foundation for understanding ViT in code. Its figures describe one configuration on CIFAR-100, and the path to strong accuracy on your own images runs through pretrained weights, careful data preparation, and version-checked code.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
The Bottom Line
Use the Keras Vision Transformer example to learn how patches, position embeddings, Transformer blocks, and the classifier fit together, and to prototype on folder-based data. Do not expect its from-scratch CIFAR-100 result to represent what a Vision Transformer can achieve with large-scale pretraining.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




