October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk5 min

TF-IDF: Calculate Term Weights and Build a Python Vectorizer

TF-IDF weights terms by their frequency in a document and rarity across a corpus. See a hand calculation, a from-scratch Python implementation, and the choices that affect scikit-learn results.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

TF-IDF gives a word more weight when it is frequent in one document but uncommon across the corpus. It multiplies a document-level term frequency (TF) by a corpus-level inverse document frequency (IDF). The exact numbers depend on choices such as the IDF formula, tokenization, and normalization, so implementations can differ while following the same basic idea.

What TF-IDF measures

TF-IDF stands for term frequency–inverse document frequency. It assigns a weight to each term in each document: TF reflects how often the term appears in that document, while IDF reduces the weight of terms found in many documents. A term that appears repeatedly in one document but in relatively few corpus documents can therefore be more useful for distinguishing that document.

As an Amazon Associate I earn from qualifying purchases.

Document frequency, written df(t), counts how many documents contain term t at least once—not the term’s total number of occurrences. IDF is calculated once per term from the corpus and then reused for each document. Stanford’s explanation of tf-idf weighting develops this distinction and intuition.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Calculate TF-IDF by hand

Consider three short documents, after consistently lowercasing and splitting on spaces:

  • red apple apple
  • red berry
  • berry tart

Use raw term counts for TF and the unsmoothed formula idf(t) = log(n / df(t)), where n is the number of documents. This is one convention, not a universal formula. Here n = 3. The term apple occurs twice in the first document and in no others, so its document frequency is 1. The term red appears in two documents, and berry also appears in two. tart appears in one.

Term TF in document 1 Document frequency IDF, using ln(3 / df) TF-IDF in document 1
apple 2 1 ln(3) ≈ 1.099 2 × 1.099 ≈ 2.197
red 1 2 ln(1.5) ≈ 0.405 1 × 0.405 ≈ 0.405
berry 0 2 ln(1.5) ≈ 0.405 0
tart 0 1 ln(3) ≈ 1.099 0

The resulting document-1 vector over the vocabulary [apple, berry, red, tart] is approximately [2.197, 0, 0.405, 0]. Although red occurs in the document, it receives less weight than apple because it is shared more broadly. A logarithm base other than the natural logarithm rescales the values; it does not change this example’s ordering.

Implement a transparent TF-IDF vectorizer in Python

This small implementation uses raw counts, scikit-learn-style smoothed IDF, and optional L2 normalization. It deliberately uses a basic lowercase-and-whitespace tokenizer so the stages are visible; it is not a replacement for configurable text preprocessing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import math
import re
from collections import Counter

def tokenize(text):
    return re.findall(r"bw+b", text.lower())

def fit_tfidf(documents):
    tokenized = [tokenize(doc) for doc in documents]
    vocabulary = sorted({token for doc in tokenized for token in doc})
    n_documents = len(documents)

    document_frequency = {
        term: sum(term in set(doc) for doc in tokenized)
        for term in vocabulary
    }
    idf = {
        term: math.log((1 + n_documents) / (1 + document_frequency[term])) + 1
        for term in vocabulary
    }

    def transform(new_documents):
        rows = []
        for text in new_documents:
            counts = Counter(tokenize(text))
            row = [counts[term] * idf[term] for term in vocabulary]
            length = math.sqrt(sum(value * value for value in row))
            if length:
                row = [value / length for value in row]
            rows.append(row)
        return rows

    return vocabulary, idf, transform

corpus = ["red apple apple", "red berry", "berry tart"]
vocabulary, idf, transform = fit_tfidf(corpus)
rows = transform(corpus)

print(vocabulary)
print(rows[0])

The code’s fit_tfidf function learns the vocabulary and IDF values from the supplied corpus. Its returned transform function uses those same features and weights for any later documents. Terms not in the fitted vocabulary are ignored. If a document contains no fitted vocabulary terms, its vector remains all zeros rather than attempting to normalize a zero-length vector.

What each stage does

  1. Normalize and tokenize: Apply the same text rules to every document. Lowercasing, punctuation handling, stop-word removal, and n-gram inclusion affect which features exist.
  2. Build the vocabulary: Create one shared ordered list of terms and count term occurrences in each document.
  3. Count document frequency: For each term, add one for each document containing it, regardless of how often it occurs there.
  4. Compute IDF: The example uses log((1 + n) / (1 + df(t))) + 1, which smooths the frequency and adds an offset.
  5. Multiply TF by IDF: Here TF is the raw occurrence count. Each document gets a weighted vector over the shared vocabulary.
  6. Normalize: The code divides each nonzero vector by its Euclidean length, producing a unit-length vector.
  7. Reuse the fit: Transform new documents with the original vocabulary and IDF, rather than learning new ones for each batch.

With L2-normalized vectors, a dot product between two vectors corresponds to their cosine similarity. Normalization makes vector length less dependent on document length; it does not change which terms are in the vocabulary.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Why results differ from scikit-learn

Scikit-learn’s TfidfVectorizer combines count vectorization and TF-IDF transformation. Its documented defaults include raw-count TF, smoothed IDF, an additive offset, and L2 normalization. The implementation above uses the same stated IDF expression and normalization, but its simple tokenizer and vocabulary policy do not necessarily match the library’s defaults. For reproducible comparisons, align preprocessing and every formula choice as well as the corpus. See the scikit-learn text feature extraction documentation and the TfidfVectorizer API reference.

Choice Teaching implementation above scikit-learn documented default
Term frequency Raw occurrence count Raw occurrence count; with sublinear_tf=True, use 1 + log(tf) instead
IDF smoothing and offset log((1 + n) / (1 + df)) + 1 log((1 + n) / (1 + df)) + 1 when smooth_idf=True
Normalization L2 for each nonzero vector norm='l2'
Preprocessing and features Lowercase and regex tokenization; all observed terms Configurable preprocessing, tokenization, stop words, and n-gram range
Vocabulary and IDF for later data Reuse values learned during fitting Fit the vectorizer, then transform later data with the fitted vocabulary and IDF

Scikit-learn describes smoothing as adding one to the numerator and denominator of IDF “as if an extra document was seen containing every term in the collection exactly once,” preventing division by zero. With smoothing, even a term present in every document has IDF 1 because of the additive offset; without smoothing, the formula and resulting values differ. The library’s options and exact behavior are documented in its TfidfTransformer reference.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Check the choices when matching values

  • Confirm that tokenization, lowercasing, stop words, and n-gram range match.
  • Check whether TF is a raw count, binary presence, or a logarithmically scaled count.
  • Compare the IDF equation, including smoothing and any additive constant.
  • Check whether vectors are unnormalized, L1-normalized, or L2-normalized.
  • Use the same fitting corpus and the same fitted vocabulary and IDF weights.

For a deeper information-retrieval treatment, Stanford’s Introduction to Information Retrieval includes a chapter on term weighting.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.