Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix Now×
Skip to content
World desk5 min

Near-Duplicate Image Search in Python with Keras: A Practical Guide

Keras can turn images into embeddings and retrieve likely near-duplicates. Learn how the example works, when to use an approximate index, and how to check results safely.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To find near-duplicate images with Keras, turn each image into a feature embedding, then retrieve images with nearby embeddings. Keras’s official example demonstrates this with a pretrained BiT-ResNet, normalized features, and locality-sensitive hashing (LSH). Treat its results as candidate matches—not proof that two files are duplicates—and verify candidates before deleting or merging anything.

What the Keras workflow does

The official Keras near-duplicate image search tutorial, by Sayak Paul, builds an approximate retrieval system in four stages:

  1. Prepare each image. The example resizes demonstration images to 224 × 224 pixels.
  2. Extract features. A pretrained BiT-ResNet classifier produces a 2,048-dimensional representation for each image.
  3. Normalize and project. The example normalizes those representations and applies random projections to reduce them to bitwise hash values.
  4. Retrieve candidates. It queries multiple hash tables to find images that land in nearby buckets, then displays candidate matches for review.

Similar images can fall into different buckets because the projections are random, so querying multiple tables helps retrieval. The number of tables and the reduced dimensionality affect the trade-off between finding matches and keeping the index manageable. LSH is approximate: a true match can be missed, and a returned candidate can be wrong.

For an application, store each embedding with a stable image identifier and file path. Deduplicate hits returned from multiple tables and rank the remaining candidates with a similarity measure appropriate to your definition of “duplicate.” Don’t rely on a hash-bucket match alone.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose the method that fits your duplicate definition

“Near-duplicate” can mean the same photograph after recompression, a resized copy, or two images that merely depict similar subjects. Choose a representation and comparison method for the kind of match you actually want.

Approach Useful for Main limitation
Exact file or pixel hash Finding byte-identical files, or files with identical pixel data when decoded consistently. Recompression, resizing, cropping, or other edits can change the hash, so it does not solve visual near-duplicate search.
Perceptual or structural comparison Checking whether two images remain visually alike after relatively light changes. Keras image operations include SSIM for comparing an image pair: Keras image ops. Thresholds depend on image content and transformations. Pairwise comparison by itself is not an indexed retrieval system for a large collection.
Learned embeddings with exact ranking A straightforward baseline for a modest collection: compute embeddings, normalize them, and rank by cosine similarity (equivalently, dot product for normalized vectors). A model may rank semantically similar but distinct images highly. That is useful for visual search, but can be a false match for strict duplicate detection.
Learned embeddings with an approximate index or LSH Retrieving candidates quickly as a collection grows, with the model supplying visual features and an index reducing the search work. Approximate retrieval may miss matches; accuracy, latency, memory use, and operational effort depend on the model, index, and data.

The Keras tutorial’s own example shows incorrect retrievals. Its author notes that better embeddings may help and mentions ArcFace and supervised contrastive learning as possible representation approaches. A separate Keras metric-learning example explores image similarity with metric learning; a Keras dual-encoder example demonstrates another learned-representation approach for image search. Neither makes one model or threshold universally correct for duplicate detection.

Build a trustworthy candidate-review workflow

  1. Define what counts as a match. Decide whether you want exact copies, the same underlying photo after edits, or broader visual similarity. This determines which transformations and false matches matter.
  2. Create representative labeled pairs. Include same-image variants such as resized, recompressed, cropped, color-adjusted, rotated, or watermarked images when those changes occur in your collection. Include hard negatives: distinct images that look alike.
  3. Set a retrieval and decision rule. For exact ranking, choose a similarity threshold or review a top-k list; for LSH or another approximate index, retrieve candidates first and then rank them. A threshold is not portable across models or datasets.
  4. Measure outcomes on held-out examples. Track precision and recall at the chosen threshold or top-k. Inspect false positives and missed matches, especially for image types that matter to your use case.
  5. Keep human review before destructive actions. Use candidate results to guide review; don’t automatically delete or merge files until the workflow has been evaluated against representative labeled pairs and its errors are acceptable.

These evaluation steps are practical safeguards, not a protocol specified by the tutorial. They address the central limitation it demonstrates: approximate search can return incorrect matches.

When to use exact search, LSH, or an ANN library

Start with exact embedding ranking for a small collection

Computing cosine similarities against every stored, normalized embedding is simple and provides a useful reference ranking. It also makes it easier to inspect whether the model’s notion of similarity matches your duplicate definition. As data volume or query-rate requirements increase, measure whether this baseline still meets your latency and memory needs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use approximate retrieval when scale justifies its trade-offs

The near-duplicate tutorial is a teaching implementation of random-projection LSH, not a recommendation to write a production index from scratch. As the author puts it: “Crucially, you wouldn’t reimplement locality-sensitive hashing yourself when working with real world applications.” Keras material names ScaNN and Annoy, and also mentions Faiss and Vald in the context of approximate matching or LSH. Check each library’s current compatibility and deployment support for your environment before choosing it.

There is no controlled head-to-head benchmark in the cited Keras material for these libraries. Compare them on your own representative data using:

  • Recall and false-match rate at the number of candidates you can review.
  • Query latency at realistic collection sizes and query rates.
  • Embedding and index memory, including index-build cost.
  • Implementation and operational burden, including updates to the collection.
  • Your available CPU/GPU hardware and deployment environment.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the tutorial’s timings do—and do not—tell you

The tutorial uses the tf_flowers dataset and a 1,000-image subset for a short demonstration. It reports 54.1 seconds to build tables on a Tesla T4 GPU. Its displayed benchmark over 1,000 queries reports 54.359 seconds for the unoptimized model and 13.963 seconds for the TensorRT path. These are figures from that tutorial’s particular setup, not portable performance expectations or an independent library comparison.

The example’s TensorRT work uses a GPU runtime; that does not establish a GPU requirement for the core workflow of computing embeddings and searching them. The tutorial also mentions TensorFlow Lite for mobile or edge deployment, ONNX for commodity CPU servers, and Apache TVM for cross-platform compilation. Treat these as directions discussed in the tutorial, not guarantees of current compatibility for a particular model or environment.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning

Practical takeaway

Use Keras to create image embeddings, begin with exact cosine-similarity ranking when the collection is small, and move to approximate candidate retrieval when measured scale or latency needs warrant it. In either case, validate against labeled matches and hard negatives, rank and inspect candidates, and keep deletion decisions separate from approximate search.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.