Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

There is no single image-similarity algorithm that works for every task. Use a cryptographic hash to find byte-for-byte duplicate files, a perceptual hash for near-duplicates, SSIM for aligned images, local features for matching shared regions, and neural embeddings for images with related subjects or meanings. The right choice depends on which differences should count—and which should not.

Choose a method based on what “similar” means

Goal Use What the result tells you
Are these files exactly the same? Byte comparison or a cryptographic hash Whether the file contents match exactly
Is one a resized or recompressed copy? Perceptual hashing Distance between compact visual fingerprints
Are aligned images visually alike at the pixel or structure level? MSE or SSIM Pixel error or structural similarity after preprocessing
Do the images share a logo, object, or other region despite a viewpoint change? Local-feature matching, such as ORB How many distinctive regions match geometrically
Do they show related objects, scenes, or concepts? Neural embeddings and cosine similarity Closeness in a model’s learned representation
Do you need to search a large collection? Embeddings plus an index such as Faiss Nearest image vectors for a query

These methods are not interchangeable. A shifted copy may score poorly under pixel comparison but closely under pHash; two different photos of dogs may be close in an embedding space but not duplicates. For a workflow that needs several kinds of matching, combine inexpensive filters with a method suited to the final decision.

Set up a Python environment

Create an isolated environment and install the packages for local comparison. PyTorch installation varies by operating system and CPU or accelerator, so follow the official PyTorch installation selector for that part.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
python -m venv .venv

Activate it on macOS or Linux:

source .venv/bin/activate

In Windows PowerShell:

.venvScriptsActivate.ps1

Install the basic comparison stack:

python -m pip install --upgrade pip
python -m pip install pillow numpy scikit-image imagehash

For the local-feature example, also install OpenCV:

python -m pip install opencv-python

For embeddings, install PyTorch and torchvision using the platform-specific PyTorch instructions. Torchvision’s current model and weight APIs are documented in its models guide. Faiss installation depends on platform and CPU/GPU needs; consult the Faiss installation instructions.

Find exact duplicate files with SHA-256

A cryptographic hash answers whether two files have exactly the same bytes. It does not tell you whether two differently encoded files look the same: metadata, compression, or file format changes will produce a different hash.

from pathlib import Path
import hashlib


def sha256_file(path: str, chunk_size: int = 1024 * 1024) -> str:
    digest = hashlib.sha256()
    with Path(path).open("rb") as file:
        while chunk := file.read(chunk_size):
            digest.update(chunk)
    return digest.hexdigest()

print(sha256_file("image_a.jpg") == sha256_file("image_b.jpg"))

Use this as the first, fast check in a duplicate-file pipeline. If it returns False, the images may still be visually identical.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Detect near-duplicates with perceptual hashes

A perceptual hash reduces an image to a compact fingerprint. With ImageHash, subtracting two hash objects gives their Hamming distance: the number of differing bits. A smaller distance means more similar hash patterns, not a guaranteed probability of duplication. The ImageHash library provides average, perceptual, difference, wavelet, color, and crop-resistant hashes.

from PIL import Image
import imagehash


def phash_distance(path_a: str, path_b: str) -> int:
    image_a = Image.open(path_a)
    image_b = Image.open(path_b)
    return imagehash.phash(image_a) - imagehash.phash(image_b)

print("Hamming distance:", phash_distance("image_a.jpg", "image_b.jpg"))

You can compare several hash types on the same pair:

from PIL import Image
import imagehash

image_a = Image.open("image_a.jpg")
image_b = Image.open("image_b.jpg")

hash_functions = {
    "average_hash": imagehash.average_hash,
    "phash": imagehash.phash,
    "dhash": imagehash.dhash,
    "whash": imagehash.whash,
    "colorhash": imagehash.colorhash,
}

for name, function in hash_functions.items():
    print(name, function(image_a) - function(image_b))
  • Average hash is simple and fast, but its global brightness pattern can be less discriminating.
  • Difference hash captures neighboring-pixel differences.
  • Perceptual hash is a common starting point for resized or recompressed near-duplicates.
  • Wavelet hash uses a wavelet-based representation.
  • Color hash describes color distribution, not the exact spatial arrangement of colors.
  • Crop-resistant hashing can help with some crops, but should still be tested on the target images.

Perceptual hashing is a fingerprinting technique, not semantic image understanding. Repeated textures, skies, screenshots, or product grids can create false positives; a crop or major composition change can cause a missed match.

Choose a threshold from labeled examples

There is no universal pHash cutoff. Build a validation set with known near-duplicate pairs and known non-matches, then inspect the distances:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from PIL import Image
import imagehash


def phash_distance(path_a: str, path_b: str) -> int:
    return imagehash.phash(Image.open(path_a)) - imagehash.phash(Image.open(path_b))

positive = [
    phash_distance("original.jpg", "resized.jpg"),
    phash_distance("original.jpg", "compressed.jpg"),
]
negative = [phash_distance("original.jpg", "unrelated.jpg")]

print("Known near-duplicates:", positive)
print("Known non-match:", negative)

Choose a threshold that reflects the cost of each error. A false positive treats unrelated images as duplicates; a false negative misses a true duplicate. Evaluate both on examples representative of the collection rather than copying a cutoff from another dataset.

Compare aligned images with SSIM or MSE

Structural Similarity Index (SSIM) is useful for comparing images of the same scene when their dimensions and alignment are compatible. It is often used for image-processing checks, not open-ended search. The score depends on preprocessing and does not measure whether two pictures depict the same concept. See the scikit-image metrics API.

import numpy as np
from PIL import Image
from skimage.metrics import structural_similarity


def load_rgb(path: str, size=(512, 512)) -> np.ndarray:
    image = Image.open(path).convert("RGB").resize(size)
    return np.asarray(image)

image_a = load_rgb("image_a.jpg")
image_b = load_rgb("image_b.jpg")

score, difference = structural_similarity(
    image_a,
    image_b,
    channel_axis=-1,
    data_range=255,
    full=True,
)
print(f"SSIM: {score:.4f}")

Here, both arrays are converted to RGB and resized to 512 by 512, so that resizing is part of the measurement. The arrays must have the same shape; channel_axis=-1 identifies the RGB axis, and data_range=255 matches 8-bit pixel values. A high score means structural similarity under these settings, not a universal probability of depicting the same subject.

Mean squared error (MSE) is a simpler pixel metric. Lower values mean smaller average squared pixel differences:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import numpy as np


def mean_squared_error(image_a, image_b) -> float:
    a = image_a.astype(np.float32)
    b = image_b.astype(np.float32)
    return float(np.mean((a - b) ** 2))

MSE is highly sensitive to alignment: moving an image a few pixels can raise the error even when a person sees nearly the same picture. Use MSE or SSIM when that sensitivity is useful, such as validating an image transformation—not when searching for images of the same object under different framing.

Match shared regions with OpenCV ORB

Local-feature methods find distinctive keypoints and compare their descriptors. They can help match a logo, poster, book cover, or building despite scale or viewpoint changes, but struggle on tiny, blurry, textureless, or repetitive images. Raw match count is not a universal similarity score. The OpenCV documentation covers its feature-detection and descriptor-matching APIs.

import cv2


def orb_match_count(path_a: str, path_b: str) -> int:
    image_a = cv2.imread(path_a, cv2.IMREAD_GRAYSCALE)
    image_b = cv2.imread(path_b, cv2.IMREAD_GRAYSCALE)
    if image_a is None or image_b is None:
        raise FileNotFoundError("Could not read one of the images")

    orb = cv2.ORB_create(nfeatures=1500)
    keypoints_a, descriptors_a = orb.detectAndCompute(image_a, None)
    keypoints_b, descriptors_b = orb.detectAndCompute(image_b, None)
    if descriptors_a is None or descriptors_b is None:
        return 0

    matcher = cv2.BFMatcher(cv2.NORM_HAMMING, crossCheck=True)
    matches = matcher.match(descriptors_a, descriptors_b)
    matches.sort(key=lambda match: match.distance)
    return len(matches)

print(orb_match_count("image_a.jpg", "image_b.jpg"))

This compact example returns a raw count, suitable for illustrating the mechanics but not for a production decision. A stronger matcher filters by descriptor distance or uses k-nearest neighbors with Lowe’s ratio test. For planar images or object matching, use RANSAC with a homography and base the decision on geometrically consistent matches. Resolution, texture, blur, lighting, and detector settings all affect the result.

Compare semantic content with image embeddings

An embedding model maps an image to a vector. Comparing normalized vectors with a dot product gives cosine similarity: higher values indicate closer representations for that model. A score is neither a probability nor a universal percentage match. The model’s training and the target domain determine whether its ranking is useful.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

This torchvision example uses a ResNet-50 feature representation and the preprocessing associated with its default weights. A classification backbone’s features are a practical starting point, but a model trained specifically for image retrieval or image-text alignment may be better for semantic search.

import torch
import torch.nn.functional as F
from PIL import Image
from torchvision.models import resnet50, ResNet50_Weights

weights = ResNet50_Weights.DEFAULT
model = resnet50(weights=weights)
model.fc = torch.nn.Identity()
model.eval()
preprocess = weights.transforms()


def image_embedding(path: str) -> torch.Tensor:
    image = Image.open(path).convert("RGB")
    tensor = preprocess(image).unsqueeze(0)
    with torch.inference_mode():
        vector = model(tensor)
    return F.normalize(vector, p=2, dim=1)

embedding_a = image_embedding("image_a.jpg")
embedding_b = image_embedding("image_b.jpg")
score = float(embedding_a @ embedding_b.T)
print(f"Cosine similarity: {score:.4f}")

Preprocessing must match the model’s expectations. Convert consistently to RGB, use the model’s supplied transforms, and validate rankings on examples from the intended use case. Embeddings can group related subjects, but they do not guarantee that two images show the same object or establish identity.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Search an image collection with Faiss

For a collection, compute and store one embedding per image, then search vectors rather than repeatedly opening and comparing every pair. A naive all-pairs comparison across N images requires about N(N−1)/2 comparisons. Faiss supports dense-vector search, including exact and approximate indexes. Its getting-started guide describes the Python workflow and float32 matrices.

For cosine similarity, normalize both the stored vectors and query, then use inner-product search. Faiss documents this equivalence in its FAQ.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import faiss
import numpy as np

# embeddings: (number_of_images, embedding_dimension)
embeddings = np.asarray(embeddings, dtype="float32")
query = np.asarray(query_embedding, dtype="float32")

faiss.normalize_L2(embeddings)
faiss.normalize_L2(query)

index = faiss.IndexFlatIP(embeddings.shape[1])
index.add(embeddings)

scores, indices = index.search(query, 5)
for score, image_id in zip(scores[0], indices[0]):
    print(image_id, float(score))

Keep a stable mapping from each vector position to the image path or database record; the returned index IDs are not filenames. IndexFlatIP performs exact inner-product search and is a straightforward baseline. Approximate indexes, including IVF, HNSW, or product quantization, can improve speed or memory use at the cost of some recall or additional tuning. For small collections, brute-force search may be simpler; see Faiss’s guidance on search without an index.

Build a hybrid pipeline when one score is not enough

A practical system can use each method for the stage it handles well:

  1. Use SHA-256 to eliminate exact byte-for-byte duplicates.
  2. Use pHash to cheaply flag likely resized or recompressed copies.
  3. Use embedding search to retrieve images with related subjects or concepts.
  4. For candidates where shared visual regions matter, verify with ORB and geometric consistency, or with a domain-specific model.
  5. Apply a decision threshold calibrated on labeled examples and record which method produced the score.

Before choosing the final method, specify whether resizing, compression, rotation, cropping, color shifts, or a changed scene should preserve a match. Those invariances define what the system should treat as similar.

Preprocessing and common failure modes

  • Different dimensions: SSIM and MSE need compatible array shapes. Resize or align deliberately; resizing changes the measurement.
  • Orientation: Apply EXIF orientation consistently before pixel comparison, or displayed images may match while stored pixel matrices do not.
  • Color channels: Pillow examples convert to RGB; OpenCV’s imread color mode normally returns BGR, while this ORB example explicitly reads grayscale. Keep channel conventions clear.
  • Transparency: Decide whether to composite alpha over a specified background or discard it. For AWS Rekognition workflows, its image guidance recommends removing alpha for certain comparison inputs and notes that inference uses RGB channels.
  • Unreadable files: OpenCV returns None when it cannot decode a path; the ORB example checks this. Also handle corrupt or unsupported files in batch jobs.
  • No descriptors: Blank, tiny, or textureless images may yield no ORB descriptors; return no match or route them to another method.
  • Faiss dtype or normalization: Use two-dimensional NumPy arrays of float32 and normalize both the database vectors and query for cosine-by-inner-product search.
  • Model setup: Torchvision weights may need to be downloaded the first time. Inference can run on CPU or a supported accelerator; ensure the model and input tensor use the same device if you move either.
  • Threshold drift: Scores shift with image population, preprocessing, model, and collection domain. Revalidate thresholds when those change.

When to use a cloud image service

Hosted services can be useful when you need managed scaling or a defined capability and have approval to process the images in that environment. Google Cloud’s Image Warehouse overview describes image and text queries over indexed collections. For an AWS-native application needing supported image analysis or face-search operations, consult Amazon Rekognition documentation; it should not be treated as a general-purpose semantic similarity engine.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Check current regional availability, input limits, pricing, retention, and data-governance terms in the provider’s current documentation before deployment. For privacy-sensitive workloads, local embeddings and indexing avoid sending images to a hosted service, subject to the requirements of the application.

Keep face matching separate

Face similarity is a biometric-identification problem, not ordinary image similarity. AWS documents SearchFacesByImage as a face-collection search operation that detects the largest face and returns similarity scores for matches. Do not use general image-similarity code—or an unvalidated score—for identity, employment, housing, lending, surveillance, or other high-impact decisions.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.