Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteTF-IDF gives a word more weight when it is frequent in one document but uncommon across the corpus. It multiplies a document-level term frequency (TF) by a corpus-level inverse document frequency (IDF). The exact numbers depend on choices such as the IDF formula, tokenization, and normalization, so implementations can differ while following the same basic idea.
What TF-IDF measures
TF-IDF stands for term frequency–inverse document frequency. It assigns a weight to each term in each document: TF reflects how often the term appears in that document, while IDF reduces the weight of terms found in many documents. A term that appears repeatedly in one document but in relatively few corpus documents can therefore be more useful for distinguishing that document.
As an Amazon Associate I earn from qualifying purchases.
Document frequency, written df(t), counts how many documents contain term t at least once—not the term’s total number of occurrences. IDF is calculated once per term from the corpus and then reused for each document. Stanford’s explanation of tf-idf weighting develops this distinction and intuition.
Calculate TF-IDF by hand
Consider three short documents, after consistently lowercasing and splitting on spaces:
#1 Best Overall
red apple applered berryberry tart
Use raw term counts for TF and the unsmoothed formula idf(t) = log(n / df(t)), where n is the number of documents. This is one convention, not a universal formula. Here n = 3. The term apple occurs twice in the first document and in no others, so its document frequency is 1. The term red appears in two documents, and berry also appears in two. tart appears in one.
| Term | TF in document 1 | Document frequency | IDF, using ln(3 / df) | TF-IDF in document 1 |
|---|---|---|---|---|
| apple | 2 | 1 | ln(3) ≈ 1.099 | 2 × 1.099 ≈ 2.197 |
| red | 1 | 2 | ln(1.5) ≈ 0.405 | 1 × 0.405 ≈ 0.405 |
| berry | 0 | 2 | ln(1.5) ≈ 0.405 | 0 |
| tart | 0 | 1 | ln(3) ≈ 1.099 | 0 |
The resulting document-1 vector over the vocabulary [apple, berry, red, tart] is approximately [2.197, 0, 0.405, 0]. Although red occurs in the document, it receives less weight than apple because it is shared more broadly. A logarithm base other than the natural logarithm rescales the values; it does not change this example’s ordering.
Implement a transparent TF-IDF vectorizer in Python
This small implementation uses raw counts, scikit-learn-style smoothed IDF, and optional L2 normalization. It deliberately uses a basic lowercase-and-whitespace tokenizer so the stages are visible; it is not a replacement for configurable text preprocessing.
Recommended Free Tools
import math
import re
from collections import Counter
def tokenize(text):
return re.findall(r"bw+b", text.lower())
def fit_tfidf(documents):
tokenized = [tokenize(doc) for doc in documents]
vocabulary = sorted({token for doc in tokenized for token in doc})
n_documents = len(documents)
document_frequency = {
term: sum(term in set(doc) for doc in tokenized)
for term in vocabulary
}
idf = {
term: math.log((1 + n_documents) / (1 + document_frequency[term])) + 1
for term in vocabulary
}
def transform(new_documents):
rows = []
for text in new_documents:
counts = Counter(tokenize(text))
row = [counts[term] * idf[term] for term in vocabulary]
length = math.sqrt(sum(value * value for value in row))
if length:
row = [value / length for value in row]
rows.append(row)
return rows
return vocabulary, idf, transform
corpus = ["red apple apple", "red berry", "berry tart"]
vocabulary, idf, transform = fit_tfidf(corpus)
rows = transform(corpus)
print(vocabulary)
print(rows[0])
The code’s fit_tfidf function learns the vocabulary and IDF values from the supplied corpus. Its returned transform function uses those same features and weights for any later documents. Terms not in the fitted vocabulary are ignored. If a document contains no fitted vocabulary terms, its vector remains all zeros rather than attempting to normalize a zero-length vector.
What each stage does
- Normalize and tokenize: Apply the same text rules to every document. Lowercasing, punctuation handling, stop-word removal, and n-gram inclusion affect which features exist.
- Build the vocabulary: Create one shared ordered list of terms and count term occurrences in each document.
- Count document frequency: For each term, add one for each document containing it, regardless of how often it occurs there.
- Compute IDF: The example uses
log((1 + n) / (1 + df(t))) + 1, which smooths the frequency and adds an offset. - Multiply TF by IDF: Here TF is the raw occurrence count. Each document gets a weighted vector over the shared vocabulary.
- Normalize: The code divides each nonzero vector by its Euclidean length, producing a unit-length vector.
- Reuse the fit: Transform new documents with the original vocabulary and IDF, rather than learning new ones for each batch.
With L2-normalized vectors, a dot product between two vectors corresponds to their cosine similarity. Normalization makes vector length less dependent on document length; it does not change which terms are in the vocabulary.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Why results differ from scikit-learn
Scikit-learn’s TfidfVectorizer combines count vectorization and TF-IDF transformation. Its documented defaults include raw-count TF, smoothed IDF, an additive offset, and L2 normalization. The implementation above uses the same stated IDF expression and normalization, but its simple tokenizer and vocabulary policy do not necessarily match the library’s defaults. For reproducible comparisons, align preprocessing and every formula choice as well as the corpus. See the scikit-learn text feature extraction documentation and the TfidfVectorizer API reference.
| Choice | Teaching implementation above | scikit-learn documented default |
|---|---|---|
| Term frequency | Raw occurrence count | Raw occurrence count; with sublinear_tf=True, use 1 + log(tf) instead |
| IDF smoothing and offset | log((1 + n) / (1 + df)) + 1 |
log((1 + n) / (1 + df)) + 1 when smooth_idf=True |
| Normalization | L2 for each nonzero vector | norm='l2' |
| Preprocessing and features | Lowercase and regex tokenization; all observed terms | Configurable preprocessing, tokenization, stop words, and n-gram range |
| Vocabulary and IDF for later data | Reuse values learned during fitting | Fit the vectorizer, then transform later data with the fitted vocabulary and IDF |
Scikit-learn describes smoothing as adding one to the numerator and denominator of IDF “as if an extra document was seen containing every term in the collection exactly once,” preventing division by zero. With smoothing, even a term present in every document has IDF 1 because of the additive offset; without smoothing, the formula and resulting values differ. The library’s options and exact behavior are documented in its TfidfTransformer reference.
Check the choices when matching values
- Confirm that tokenization, lowercasing, stop words, and n-gram range match.
- Check whether TF is a raw count, binary presence, or a logarithmically scaled count.
- Compare the IDF equation, including smoothing and any additive constant.
- Check whether vectors are unnormalized, L1-normalized, or L2-normalized.
- Use the same fitting corpus and the same fitted vocabulary and IDF weights.
For a deeper information-retrieval treatment, Stanford’s Introduction to Information Retrieval includes a chapter on term weighting.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




