Build a small semantic search prototype by embedding a handful of text passages, embedding a query with the same model, and ranking passages by vector similarity. The example below uses Sentence Transformers and a direct scan of the corpus: it is easy to inspect and needs no separate search database. Results are ranked candidates, not guaranteed answers.
How semantic search finds relevant passages
Semantic search represents corpus entries—such as sentences, paragraphs, or documents—and a query as vectors in a shared space, then retrieves nearby vectors. That can help find related wording when a query uses synonyms, abbreviations, or misspellings that do not literally appear in a passage. What counts as similar depends on the embedding model.
This example is asymmetric retrieval: a short query is matched against longer passages. Sentence Transformers recommends using encode_query for the query and encode_document for corpus entries when the selected model supports those methods. Some models apply different prompts or task routing for the two roles, so follow the model’s intended usage. For inputs of similar length, such as question-to-question search, retrieval is symmetric instead.
Build the minimal Python search engine
Install the library
Install Sentence Transformers in the Python environment you intend to use:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
pip install -U sentence-transformers
The code uses the documented API shape; check compatibility with your installed library version and the chosen model before relying on it in an application.
Encode and search a small corpus
Each passage has a stable ID and original text. Keeping them together makes it clear which text belongs to each embedding row and lets search return readable content rather than an index alone.
Rank #2
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2")
corpus = [
{"id": "p1", "text": "A semantic search system compares text embeddings."},
{"id": "p2", "text": "Cosine similarity compares vector directions."},
{"id": "p3", "text": "A bicycle uses two wheels."},
]
texts = [item["text"] for item in corpus]
# Compute passage vectors once, then reuse them for queries.
corpus_embeddings = model.encode_document(texts, convert_to_tensor=True)
query = "How can I compare the meaning of two passages?"
query_embedding = model.encode_query(query, convert_to_tensor=True)
scores = model.similarity(query_embedding, corpus_embeddings)[0]
k = min(3, len(corpus))
values, indices = scores.topk(k)
results = [
{"id": corpus[int(i)]["id"], "text": corpus[int(i)]["text"], "score": float(score)}
for score, i in zip(values, indices)
]
for result in results:
print(result["id"], result["score"], result["text"])
This is an illustrative adaptation of the official documented workflow, not a tested or benchmarked program. The official quickstart uses sentence-transformers/all-MiniLM-L6-v2; its example shows three text embeddings with shape [3, 384]. That shape is specific to that example and model, not a universal embedding size.
What the search code does
- Load a model: the model converts text into numeric vectors.
- Prepare the corpus: the text list and embedding rows stay in the same order, while IDs preserve stable references.
- Encode the passages once: reuse these vectors until the corpus changes rather than recomputing them for every query.
- Encode a query and score it: compare its vector with each stored passage vector.
- Return the top results:
k = min(requested_k, len(corpus))prevents asking for more results than there are passages.
How to interpret similarity scores
The sample ranks results by cosine similarity, which compares vector direction using a normalized dot product. A larger score means a closer match under the selected model and scoring method; it is not a calibrated probability that a passage is correct or relevant. Inspect the returned text, and assess quality with queries representative of your own corpus.
Sentence Transformers uses cosine similarity by default in its semantic-search utility. When vectors are normalized to unit length, dot product gives the same ranking as cosine similarity and can avoid repeated normalization. Scikit-learn also documents cosine similarity for document vectors, including sparse matrices. A sparse TF-IDF baseline can therefore use cosine similarity too, but TF-IDF measures lexical feature overlap rather than learned sentence-level meaning.
When a direct scan is enough—and when to index
For a tiny corpus, comparing one query vector against every stored vector is the simplest design: there is no index to build, and the exact scores can be inspected directly. Sentence Transformers documentation says manual exact search can be used for corpora “up to about 1 million entries.” Treat that as project guidance, not a capacity or latency guarantee. Model dimensions, available memory, batching, hardware, query rate, and response-time needs all affect what is practical.
For larger collections or demanding latency targets, approximate-nearest-neighbor (ANN) indexes such as FAISS, Annoy, and hnswlib can speed up retrieval. ANN search can miss the exact nearest neighbors; its settings involve a recall-versus-latency trade-off. Evaluate it on the intended corpus and query set, and choose an acceptable balance rather than assuming an index is automatically better.
Improve results with retrieve-and-rerank
If the first-stage ranking is not good enough, a two-stage system can use a bi-encoder to retrieve a shortlist and a cross-encoder to rescore each query–passage pair. Sentence Transformers describes cross-encoders as often more accurate but slower because they compute each pair, so applying one to a shortlist limits the extra work. Whether the improvement justifies that cost depends on your quality and latency requirements.
Best Value
Compare approaches using representative-query relevance, latency, memory use, index-building complexity, and—where ANN is involved—recall. Keep lexical search or filtering available when exact names, codes, or phrases matter; semantic similarity does not ensure those exact strings rank first.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Practical limits and next steps
- Keep mappings aligned: if corpus items and embedding rows drift out of order, the system can display the wrong passage for a highly ranked vector.
- Refresh vectors when text changes: stored embeddings represent the text used to create them; update the affected vectors when corpus content is edited.
- Choose the model deliberately: the model determines which relationships the vectors can capture, and its documented query/document usage should guide encoding.
- Validate before deployment: try representative queries, inspect retrieved passages, and measure performance against the application’s requirements. No accuracy or speed level is guaranteed by this minimal example.
The minimal workflow requires a text corpus and a compatible Sentence Transformers software setup. The cited documentation does not establish a dedicated hardware requirement or a paid database requirement for this prototype.
Quick Recap
Sources and further reading
- Sentence Transformers: Semantic Search — encoding roles, cosine scoring, exact search, and ANN options.
- Sentence Transformers: Quickstart — model-loading and embedding examples, including the documented output shape.
- scikit-learn: Cosine Similarity — cosine similarity for dense and sparse vectors.
- Natural Language Processing: A Textbook with Python Implementation by Raymond S. T. Lee — broader NLP further reading, including Python-based workshops on semantic analysis and word vectors using spaCy; it is not required for this implementation.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




