To find near-duplicate images with Keras, turn each image into a feature embedding, then retrieve images with nearby embeddings. Keras’s official example demonstrates this with a pretrained BiT-ResNet, normalized features, and locality-sensitive hashing (LSH). Treat its results as candidate matches—not proof that two files are duplicates—and verify candidates before deleting or merging anything.
What the Keras workflow does
The official Keras near-duplicate image search tutorial, by Sayak Paul, builds an approximate retrieval system in four stages:
- Prepare each image. The example resizes demonstration images to 224 × 224 pixels.
- Extract features. A pretrained BiT-ResNet classifier produces a 2,048-dimensional representation for each image.
- Normalize and project. The example normalizes those representations and applies random projections to reduce them to bitwise hash values.
- Retrieve candidates. It queries multiple hash tables to find images that land in nearby buckets, then displays candidate matches for review.
Similar images can fall into different buckets because the projections are random, so querying multiple tables helps retrieval. The number of tables and the reduced dimensionality affect the trade-off between finding matches and keeping the index manageable. LSH is approximate: a true match can be missed, and a returned candidate can be wrong.
For an application, store each embedding with a stable image identifier and file path. Deduplicate hits returned from multiple tables and rank the remaining candidates with a similarity measure appropriate to your definition of “duplicate.” Don’t rely on a hash-bucket match alone.
Recommended Free Tools
#1 Best Overall
Choose the method that fits your duplicate definition
“Near-duplicate” can mean the same photograph after recompression, a resized copy, or two images that merely depict similar subjects. Choose a representation and comparison method for the kind of match you actually want.
| Approach | Useful for | Main limitation |
|---|---|---|
| Exact file or pixel hash | Finding byte-identical files, or files with identical pixel data when decoded consistently. | Recompression, resizing, cropping, or other edits can change the hash, so it does not solve visual near-duplicate search. |
| Perceptual or structural comparison | Checking whether two images remain visually alike after relatively light changes. Keras image operations include SSIM for comparing an image pair: Keras image ops. | Thresholds depend on image content and transformations. Pairwise comparison by itself is not an indexed retrieval system for a large collection. |
| Learned embeddings with exact ranking | A straightforward baseline for a modest collection: compute embeddings, normalize them, and rank by cosine similarity (equivalently, dot product for normalized vectors). | A model may rank semantically similar but distinct images highly. That is useful for visual search, but can be a false match for strict duplicate detection. |
| Learned embeddings with an approximate index or LSH | Retrieving candidates quickly as a collection grows, with the model supplying visual features and an index reducing the search work. | Approximate retrieval may miss matches; accuracy, latency, memory use, and operational effort depend on the model, index, and data. |
The Keras tutorial’s own example shows incorrect retrievals. Its author notes that better embeddings may help and mentions ArcFace and supervised contrastive learning as possible representation approaches. A separate Keras metric-learning example explores image similarity with metric learning; a Keras dual-encoder example demonstrates another learned-representation approach for image search. Neither makes one model or threshold universally correct for duplicate detection.
Build a trustworthy candidate-review workflow
- Define what counts as a match. Decide whether you want exact copies, the same underlying photo after edits, or broader visual similarity. This determines which transformations and false matches matter.
- Create representative labeled pairs. Include same-image variants such as resized, recompressed, cropped, color-adjusted, rotated, or watermarked images when those changes occur in your collection. Include hard negatives: distinct images that look alike.
- Set a retrieval and decision rule. For exact ranking, choose a similarity threshold or review a top-k list; for LSH or another approximate index, retrieve candidates first and then rank them. A threshold is not portable across models or datasets.
- Measure outcomes on held-out examples. Track precision and recall at the chosen threshold or top-k. Inspect false positives and missed matches, especially for image types that matter to your use case.
- Keep human review before destructive actions. Use candidate results to guide review; don’t automatically delete or merge files until the workflow has been evaluated against representative labeled pairs and its errors are acceptable.
These evaluation steps are practical safeguards, not a protocol specified by the tutorial. They address the central limitation it demonstrates: approximate search can return incorrect matches.
When to use exact search, LSH, or an ANN library
Start with exact embedding ranking for a small collection
Computing cosine similarities against every stored, normalized embedding is simple and provides a useful reference ranking. It also makes it easier to inspect whether the model’s notion of similarity matches your duplicate definition. As data volume or query-rate requirements increase, measure whether this baseline still meets your latency and memory needs.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteRank #3
Use approximate retrieval when scale justifies its trade-offs
The near-duplicate tutorial is a teaching implementation of random-projection LSH, not a recommendation to write a production index from scratch. As the author puts it: “Crucially, you wouldn’t reimplement locality-sensitive hashing yourself when working with real world applications.” Keras material names ScaNN and Annoy, and also mentions Faiss and Vald in the context of approximate matching or LSH. Check each library’s current compatibility and deployment support for your environment before choosing it.
There is no controlled head-to-head benchmark in the cited Keras material for these libraries. Compare them on your own representative data using:
Rank #4
- Recall and false-match rate at the number of candidates you can review.
- Query latency at realistic collection sizes and query rates.
- Embedding and index memory, including index-build cost.
- Implementation and operational burden, including updates to the collection.
- Your available CPU/GPU hardware and deployment environment.
What the tutorial’s timings do—and do not—tell you
The tutorial uses the tf_flowers dataset and a 1,000-image subset for a short demonstration. It reports 54.1 seconds to build tables on a Tesla T4 GPU. Its displayed benchmark over 1,000 queries reports 54.359 seconds for the unoptimized model and 13.963 seconds for the TensorRT path. These are figures from that tutorial’s particular setup, not portable performance expectations or an independent library comparison.
The example’s TensorRT work uses a GPU runtime; that does not establish a GPU requirement for the core workflow of computing embeddings and searching them. The tutorial also mentions TensorFlow Lite for mobile or edge deployment, ONNX for commodity CPU servers, and Apache TVM for cross-platform compilation. Treat these as directions discussed in the tutorial, not guarantees of current compatibility for a particular model or environment.
Best Value
- Use scikit-learn to track an example ML project end to end
- Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
- Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
- Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
- Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
Practical takeaway
Use Keras to create image embeddings, begin with exact cosine-similarity ranking when the collection is small, and move to approximate candidate retrieval when measured scale or latency needs warrant it. In either case, validate against labeled matches and hard negatives, rank and inspect candidates, and keep deletion decisions separate from approximate search.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




