Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

The most reliable starting point for a custom image-classification project is transfer learning: use a model pretrained on a large image collection, replace its original classification head, train the new head on your classes, then fine-tune part of the backbone with a much smaller learning rate if validation results justify it. This approach usually reaches a useful baseline faster and with less data than training every layer from scratch, but it still depends on consistent labels, leakage-free splits, production-like images and task-appropriate evaluation.

First decide whether classification is the right problem

Image classification assigns labels to an entire image. It is appropriate when the required answer concerns the image as a whole, such as “healthy leaf” versus “diseased leaf.” It is not the right output when users need object locations or pixel boundaries.

Task Output Example
Classification One or more labels for the whole image Cat, dog or healthy leaf
Object detection Bounding boxes and labels Three cars and their locations
Instance segmentation A pixel mask for each object Exact pixels belonging to each person
Semantic segmentation A class for every pixel Road, sky and building pixels
Multilabel classification Several independent labels Dog, grass and vehicle in one image

Binary, multiclass and multilabel outputs

  • Binary: two mutually exclusive classes.
  • Multiclass: exactly one class from three or more choices.
  • Multilabel: any number of labels can be true at once.

This choice determines the label format, output activation, loss function, thresholds and metrics. Forcing a single whole-image label onto a scene containing several relevant objects is usually a task-definition error, not a model-architecture problem.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Write the label policy before writing code

Document what qualifies for every class, including positive and negative examples. Specify how to label borderline images, images containing multiple categories, unusable images and genuinely unknown cases. Decide whether an “other” or reject path is needed, and record whether false positives or false negatives are more costly.

Include escalation rules for annotators and keep a versioned annotation guide. A model learns statistical associations from the labels it receives; it cannot repair a policy in which two equivalent images are assigned different classes.

Build a trustworthy dataset

Use a clear directory layout

dataset/
  train/
    class_a/
    class_b/
    class_c/
  validation/
    class_a/
    class_b/
    class_c/
  test/
    class_a/
    class_b/
    class_c/

Keras can read class-specific directories directly. TensorFlow’s transfer-learning guide and image tutorial demonstrate resizing, batching, caching and prefetching: transfer-learning workflow and end-to-end tutorial.

Split by the source of correlation

Do not randomly distribute frames from one video, photos of one product, images from one patient or repeated views of one object across all splits. Deduplicate first, then split by person, patient, product, location, acquisition session or time period as appropriate. Augmented copies belong only in the training path. Keep the test set untouched until the final evaluation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Run a data-quality checklist

  • Decode every file and remove corrupt or empty images.
  • Record dimensions, aspect ratios, color channels and formats.
  • Find exact and near duplicates before splitting.
  • Count examples per class and inspect rare classes.
  • Review suspected mislabels and ambiguous cases.
  • Look for watermarks, backgrounds, camera types or locations that reveal the label accidentally.
  • Compare training images with the camera, lighting, geography and workflow expected in production.
  • Record dataset provenance, licensing, consent and privacy constraints.

AWS’s managed TensorFlow image-classification algorithm accepts JPEG and PNG training images, but any local pipeline should still validate decoding and channel order itself: AWS TensorFlow image classification.

Prepare images without changing their meaning

Choose a resize and crop policy that preserves the information needed for the label. Use the preprocessing function required by the selected backbone, and keep RGB versus grayscale assumptions consistent. Training augmentation should represent plausible production variation; validation, testing and inference should use deterministic preprocessing.

Rank #2
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
  • Useful candidates include modest flips, rotations, crops, translations, brightness or contrast changes, zoom, blur and compression simulation.
  • Do not flip text, road signs, medical laterality or directional symbols when orientation matters.
  • Do not crop away the object, create impossible rotations or alter scientifically meaningful colors.

The official TensorFlow tutorial demonstrates random horizontal flips and rotations as examples of label-preserving augmentation, not universal settings: TensorFlow transfer-learning tutorial.

Recommended baseline: TensorFlow and Keras transfer learning

1. Create an isolated environment

python -m venv .venv
source .venv/bin/activate        # macOS/Linux
# .venvScriptsactivate         # Windows PowerShell
python -m pip install --upgrade pip
pip install tensorflow scikit-learn matplotlib

Pin the versions used by your project in its lockfile. GPU compatibility depends on the operating system, Python version, TensorFlow release, drivers and hardware; verify those combinations in the current official documentation rather than assuming a command will work everywhere.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

2. Load and pipeline the datasets

import tensorflow as tf

IMG_SIZE = (224, 224)
BATCH_SIZE = 32
SEED = 42

train_ds = tf.keras.utils.image_dataset_from_directory(
    "dataset/train", image_size=IMG_SIZE, batch_size=BATCH_SIZE,
    seed=SEED, shuffle=True)
val_ds = tf.keras.utils.image_dataset_from_directory(
    "dataset/validation", image_size=IMG_SIZE, batch_size=BATCH_SIZE,
    seed=SEED, shuffle=False)
test_ds = tf.keras.utils.image_dataset_from_directory(
    "dataset/test", image_size=IMG_SIZE, batch_size=BATCH_SIZE,
    seed=SEED, shuffle=False)

class_names = train_ds.class_names
num_classes = len(class_names)
AUTOTUNE = tf.data.AUTOTUNE
train_ds = train_ds.prefetch(AUTOTUNE)
val_ds = val_ds.prefetch(AUTOTUNE)
test_ds = test_ds.prefetch(AUTOTUNE)

If you have not created the splits yet, perform the group-level split before calling this loader. The image size and batch size above are starting points, not guarantees.

3. Freeze a pretrained backbone and add a head

from tensorflow import keras
from tensorflow.keras import layers

augmentation = keras.Sequential([
    layers.RandomFlip("horizontal"),
    layers.RandomRotation(0.1),
    layers.RandomZoom(0.1),
], name="data_augmentation")

base_model = keras.applications.MobileNetV2(
    input_shape=IMG_SIZE + (3,), include_top=False, weights="imagenet")
base_model.trainable = False

inputs = keras.Input(shape=IMG_SIZE + (3,))
x = augmentation(inputs)
x = keras.applications.mobilenet_v2.preprocess_input(x)
x = base_model(x, training=False)
x = layers.GlobalAveragePooling2D()(x)
x = layers.Dropout(0.2)(x)
outputs = layers.Dense(num_classes, activation="softmax")(x)
model = keras.Model(inputs, outputs)
model.compile(optimizer=keras.optimizers.Adam(learning_rate=1e-3),
              loss="sparse_categorical_crossentropy", metrics=["accuracy"])

TensorFlow’s documented workflow freezes the base, trains new layers and calls the base with training=False, which matters for batch-normalization layers: Keras transfer learning. The example uses a 224×224 input, 0.2 dropout and a 1e-3 head learning rate only as illustrative starting values. The correct preprocessing depends on the backbone.

Use the matching output and loss

# Binary
outputs = layers.Dense(1, activation="sigmoid")(x)
loss = "binary_crossentropy"

# Single-label multiclass with integer IDs
outputs = layers.Dense(num_classes, activation="softmax")(x)
loss = "sparse_categorical_crossentropy"

# Single-label multiclass with one-hot labels
loss = "categorical_crossentropy"

# Multilabel
outputs = layers.Dense(num_classes, activation="sigmoid")(x)
loss = "binary_crossentropy"

Softmax makes classes compete and sum to one. Sigmoid treats labels independently; substituting one for the other changes the problem being modeled.

4. Train with checkpoints and early stopping

callbacks = [
    keras.callbacks.ModelCheckpoint("best_model.keras",
                                   monitor="val_loss", save_best_only=True),
    keras.callbacks.EarlyStopping(monitor="val_loss", patience=5,
                                  restore_best_weights=True),
    keras.callbacks.ReduceLROnPlateau(monitor="val_loss", factor=0.2,
                                      patience=2, min_lr=1e-7),
]

history = model.fit(train_ds, validation_data=val_ds,
                    epochs=20, callbacks=callbacks)

Training accuracy alone is insufficient. The best checkpoint may occur before the final epoch, and validation loss can reveal deteriorating confidence even when accuracy is unchanged. Early stopping saves compute; it does not replace a final test evaluation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

5. Fine-tune cautiously

base_model.trainable = True
for layer in base_model.layers[:-30]:
    layer.trainable = False

model.compile(optimizer=keras.optimizers.Adam(learning_rate=1e-5),
              loss="sparse_categorical_crossentropy", metrics=["accuracy"])
model.fit(train_ds, validation_data=val_ds, epochs=10,
          callbacks=callbacks)

Recompile after changing trainability and use a much lower learning rate than for the new head. Fine-tuning too aggressively can overwrite useful pretrained representations. If validation performance collapses, restore the best checkpoint, lower the rate, unfreeze fewer layers, verify preprocessing and recheck labels and split integrity.

Evaluate what matters in production

Report accuracy alongside balanced accuracy for uneven classes, per-class precision, recall, F1 and support, a confusion matrix, and ROC-AUC or PR-AUC where appropriate. Add latency and throughput when deployment constraints matter. Evaluate on a production-like holdout whose images were not used to select thresholds or tune the model.

For binary and multilabel outputs, choose thresholds on validation data according to false-positive and false-negative costs; 0.5 is not automatically optimal. Keep the test set untouched while doing so. A softmax score is a score distribution, not automatically a calibrated probability. Consider rejecting low-confidence cases for human review and monitor the rejection rate.

Choose a model and compute strategy

Option Strength Trade-off
MobileNet family Small and fast for edge or low-latency use May lose accuracy on difficult classes
EfficientNet family Strong accuracy-efficiency balance More preprocessing and deployment considerations
ResNet family Well-understood, dependable baseline Often heavier than mobile-oriented models
Vision Transformer Competitive with suitable data and hardware Can require more data, tuning or compute
Custom CNN Maximum simplicity and control Usually weaker than a good pretrained model unless the domain is specialized

Transfer learning is generally preferable for small or moderate datasets and common RGB imagery. Training from scratch becomes more defensible with a large representative dataset, unusual sensors or channels, a substantially different domain, privacy or licensing restrictions on pretrained weights, and enough compute and engineering capacity. AWS documents MobileNet, ResNet, Inception and EfficientNet as common choices and describes fine-tuning a head attached to a pretrained model: how the SageMaker algorithm works.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A CPU can handle small models and datasets. A GPU mainly reduces training time for larger experiments; PyTorch’s cloud guidance lists AWS, Google Cloud, Azure and Lightning paths but does not make a GPU universal: PyTorch cloud partners.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Diagnose the common failures

Overfitting

Rising training accuracy with stagnant or worsening validation results points to overfitting. Add representative data, use label-preserving augmentation, weight decay or dropout, simplify the head, stop earlier or fine-tune fewer layers.

Leakage

Suspiciously strong validation results followed by poor production performance often indicate duplicates or correlated entities across splits. Deduplicate, split by entity or acquisition session, and keep augmentation inside the training path.

Class imbalance

High overall accuracy with poor minority recall means the majority class dominates. Use class-weighted loss or balanced sampling, collect minority examples, inspect per-class metrics and tune thresholds. Use focal loss only when its behavior is understood.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Background shortcuts

If performance collapses when backgrounds, cameras or locations change, collect diverse scenes, crop or segment where appropriate, use background-aware augmentation and test deliberately altered backgrounds.

Domain shift

Differences by device, season, geography or lighting require a production-like holdout, provenance tracking, image-quality and class-distribution monitoring, periodic relabeling and retraining with representative new data.

Preprocessing mismatch

Poor real-world predictions despite normal training metrics commonly result from different resizing, cropping, color order or normalization. Put preprocessing in the saved model when practical, reuse the exact inference code, test known examples and store the class-index mapping.

Export and deploy responsibly

Save the complete model contract

  • Architecture, weights and best checkpoint.
  • Class names and class-index mapping.
  • Input dimensions, color-channel order and preprocessing.
  • Decision thresholds and reject policy.
  • Dataset version, split method and evaluation results.
  • Framework, dependency and hardware details.
  • Model license and pretrained-weight provenance.
  • Representative inference inputs and outputs.

Select a serving target

Target Suitable use
Local Python service Internal tools and prototypes
REST API Web and mobile clients
Batch inference Large offline image collections
Mobile or edge Offline or low-latency applications
Managed cloud endpoint Scalable serving and managed infrastructure
Browser inference Small models and client-side privacy

AWS SageMaker documents deployment paths for TensorFlow, PyTorch, ONNX and other common frameworks: SageMaker AI deployment. Managed services trade infrastructure work for cloud administration and usage-based cost; exact rates vary by region, instance, storage, traffic and endpoint uptime.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Monitor after release

  • Format, dimension and decoding failures.
  • Latency, throughput and service errors.
  • Prediction and confidence distributions.
  • Reject rate and class-frequency drift.
  • Subgroup performance and image-quality changes.
  • Accuracy on a continuously labeled sample when ground truth arrives.

Without delayed labels, accuracy cannot be observed directly; proxy signals help identify when a fresh evaluation is needed.

When local tools, cloud GPUs or managed platforms make sense

  • Learning or a small prototype: local Keras or PyTorch on a CPU, with an occasional rented GPU if training time becomes inconvenient.
  • Repeated experiments: add dataset versioning and experiment tracking.
  • Team development: use shared storage, labeling review, a model registry and reproducible training.
  • Production API: use a managed endpoint or containerized service with monitoring.
  • Regulated or large-scale deployment: choose a cloud platform based on existing identity, security, governance, data residency and operational expertise.

SageMaker, Vertex AI and Azure Machine Learning can reduce infrastructure work, while direct GPU instances provide more control. Annotation services are useful when review and consensus are the bottleneck, but no platform replaces a clear labeling guide and quality checks. Product pages include SageMaker AI, Vertex AI and Azure Machine Learning; verify current availability and pricing before committing.

The Bottom Line

Start with a leakage-free dataset, a written label policy and a frozen pretrained backbone. Train the correct head for binary, multiclass or multilabel outputs, fine-tune only after the baseline is stable, and judge success with class-level metrics and production-like data rather than accuracy alone.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.