You can use unlabeled images to improve image classification by first training an image encoder to recognize two augmented views of the same image, then training a classifier with the labeled subset. Keras demonstrates this SimCLR workflow on STL-10, including contrastive pretraining, a frozen-feature linear probe, and final fine-tuning. Its settings are a teaching example—not universal requirements or a promise of better accuracy on another dataset.
What SimCLR learns from unlabeled images
In semi-supervised learning, some training images have labels and others do not. During SimCLR pretraining, the model uses images without their labels: it creates two different augmented views of each image and treats those views as a positive pair. The training objective pulls their representations together and distinguishes them from representations of other images in the batch.
The Keras example uses an encoder to turn each view into a feature representation, then a nonlinear projection head to map that representation into the space used by the contrastive loss. It normalizes the projections, calculates temperature-scaled pairwise similarities, and applies a symmetrized cross-entropy loss with the matching view as the target. The projection head matters because the contrastive objective is applied to its output, while the encoder’s representation is what the downstream classifier uses. The original SimCLR paper also reports that augmentation composition, a learnable nonlinear transformation, and larger batches and more training steps were important in its experiments: A Simple Framework for Contrastive Learning of Visual Representations.
How to use unlabeled images in the Keras workflow
- Prepare the image pools. The Keras STL-10 example configures 100,000 unlabeled training examples and 5,000 labeled training examples. The labels are not part of the contrastive objective.
- Create paired views. Apply the contrastive augmentation pipeline twice to each image so the model sees two distinct versions of the same underlying image.
- Pretrain the encoder. Train the encoder and projection head with the contrastive loss on the unlabeled images. In the example’s combined stream, a batch contains 500 unlabeled and 25 labeled images; labels do not contribute to the contrastive loss.
- Monitor a linear probe. Freeze the encoder and train a linear classifier on its features using labeled examples. This provides a way to monitor whether the learned representation supports classification without adapting the encoder.
- Fine-tune for classification. Attach a classifier to the pretrained encoder and train the resulting model on labeled images. The tutorial also trains a randomly initialized supervised baseline using labeled data, then compares validation curves; it uses the test split for validation.
This sequence is described in the Keras SimCLR example, created on 2021-04-24 and last modified on 2024-03-04. It is one reproducible teaching configuration, not a prescription to use the same data ratio, encoder width, batch composition, or number of epochs on a different task.
#1 Best Overall
How many labeled images do you need?
There is no universal labeled-image threshold established by these sources. The Keras example’s 5,000 labeled examples are a configuration for its STL-10 experiment, not a minimum requirement. The useful amount depends on the target task, the quality and relevance of the unlabeled images, and how well the learned features transfer; evaluate with a held-out validation set rather than assuming that pretraining will beat a supervised model.
For perspective, the original SimCLR paper and its follow-up report ImageNet results using different protocols. Those figures illustrate what the papers achieved in their own experiments, not what to expect from the Keras STL-10 notebook:
Rank #2
- Use scikit-learn to track an example ML project end to end
- Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
- Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
- Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
- Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
| Work and method | Reported result | How to interpret it |
|---|---|---|
| Keras STL-10 example | No numerical accuracy figure is specified here. The tutorial reports that its pretraining-and-fine-tuning path had higher validation accuracy and lower validation loss than its randomly initialized supervised baseline. | The reported comparison is for that tutorial experiment; it is not an independently reproduced result or a guarantee on another dataset. See the Keras example. |
| Original SimCLR paper, 2020 | 76.5% ImageNet top-1 accuracy with linear evaluation; 85.8% ImageNet top-5 accuracy after fine-tuning with 1% of labels. | These are distinct evaluation measures and protocols reported by Chen, Kornblith, Norouzi, and Hinton—not results from the Keras STL-10 example. See the paper. |
| SimCLRv2 paper, 2020 | 73.9% ImageNet top-1 accuracy with ResNet-50 and 1% of labels after distillation; 77.5% with 10% of labels. | This is a larger three-stage pipeline: self-supervised pretraining, supervised fine-tuning, then distillation using unlabeled examples. It is not the original SimCLR setup or the Keras example. See Big Self-Supervised Models are Strong Semi-Supervised Learners. |
Which augmentations should you use?
For contrastive pretraining, the Keras tutorial emphasizes random crops, color jitter, and horizontal flips. It uses stronger transformations for contrastive training than for supervised classification, aiming to teach the encoder useful invariances while limiting overfitting on the smaller labeled subset. The example’s custom preprocessing layers keep augmentation in the model pipeline; its author notes that batched augmentation can run on a GPU, which may help when CPU processing is constrained.
Do not copy augmentation strengths as universal defaults. A transformation is useful only if the resulting view still preserves information needed for the target label. Excessively strong augmentation can hurt downstream performance, and the appropriate strength depends on the task and architecture. The original SimCLR paper’s finding that augmentation composition is central is a reason to validate the pipeline, not to assume one recipe transfers unchanged.
Rank #3
What batch size, architecture, and training settings should you use?
The Keras demonstration uses a compact convolutional encoder and a two-layer projection head. Its configured batch size is 525 images in total—500 unlabeled plus 25 labeled—its training duration is 20 epochs, and its contrastive temperature is 0.1. These settings describe the example, not recommended minimums or defaults for other hardware and datasets.
Batch size, temperature, augmentation strength, learning-rate schedule, and optimizer all affect training. The tutorial uses Adam with a constant schedule for its demonstration and discusses cosine decay and SGD with momentum as alternatives that may require tuning. Its author notes that larger or deeper encoders, including ResNet-50 as a common choice in the literature, can improve results while increasing training time and memory use; memory limits can also constrain batch size. The original SimCLR paper found benefits from larger batches and more training steps in its experiments, but those gains must be weighed against compute costs.
Rank #4
A GPU is an optional performance resource, not a stated prerequisite. The Keras example discusses GPU execution and Colab or a personal machine; actual hardware needs depend on image size, batch size, architecture, and training duration. Start with a model and batch your available memory can support, then tune against validation performance.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How to judge whether pretraining helped
Compare models under the same dataset split and validation procedure. The Keras example’s comparison is between a randomly initialized supervised baseline and a model that is contrastively pretrained, then fine-tuned on labeled examples; the linear probe separately tracks performance with a frozen encoder. Its prose reports higher validation accuracy and lower validation loss for the pretraining-and-fine-tuning path in that experiment.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsBest Value
When comparing with other published methods, check the dataset, labeled fraction, evaluation method, and metric. Linear evaluation of a frozen representation is not the same as fine-tuning the encoder, and top-1 accuracy is not interchangeable with top-5 accuracy. SimCLR uses negative examples in its contrastive objective; the Keras page contrasts it with SimSiam, which avoids negatives, and mentions related approaches using clustering or cross-correlation. These methods should be compared using matched data and evaluation protocols, not a single headline score.
Reproducing the tutorial
The Keras page documents the workflow and its example configuration, but does not establish compatibility across current Keras and TensorFlow releases or provide a package-version matrix. Before reproducing it, check the live notebook’s dependency versions and execution environment. Treat the reported validation comparison as the tutorial author’s result, not as a benchmark you have reproduced on your own data.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




