What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Item-based collaborative filtering recommends items by finding other items that resemble things a user has already interacted with. The resemblance comes from patterns across users—not from titles, descriptions, genres, images, or other item metadata. If many users who watched The Matrix also watched Inception, the system can recommend Inception to someone who watched The Matrix.

This guide builds a transparent Python baseline that creates an item–user matrix, calculates item-to-item cosine similarity, generates personalized recommendations, filters consumed items, and evaluates results with a time-aware holdout.

What item-based collaborative filtering does

The typical pipeline is:

  1. Collect user–item interactions.
  2. Represent items by the users who interacted with them.
  3. Calculate similarity between item vectors.
  4. Aggregate similarities across a user’s history.
  5. Remove items the user has already consumed.
  6. Return the highest-scoring candidates.

The method can power “customers also liked,” “because you watched…,” related articles, similar courses, and personalized product or media shelves.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

It is not the same as recommending semantically similar items. Collaborative similarity reflects behavior. Two films can be behaviorally similar even if their genres differ, while two films in the same genre may have little co-consumption.

#1 Best Overall
Mark Twain Grades 5-8 General Science WorkBook, Solar System, Weather, Energy, Natural Disasters, and Biology Textbook, Classroom or Homeschool Curriculum (Volume 3)
  • Supports NSE standards
  • Students will gain extra practice with the skills they are learning in their physical, earth, space, and life science curriculums
  • Grades 5-8
  • Includes 96 pages

Item-based versus other recommenders

Method Finds similarity between Recommendation logic
User-based CF Users Recommend items preferred by similar users
Item-based CF Items Recommend items related to the user’s history
Content-based Item attributes Recommend items with similar metadata or embeddings
Matrix factorization Latent user and item vectors Rank items by predicted user–item scores

Item-based methods became influential because item relationships can often be precomputed and reused for many users. Academic work by Sarwar and colleagues examined item-based recommendation algorithms in 2001, while Amazon published its item-to-item collaborative-filtering architecture in 2003. These sources support the history of the approach, but do not justify the claim that Amazon invented the entire method: academic paper and Amazon’s paper.

What data you need

At minimum, store:

user_id, item_id, interaction, timestamp

Interactions may be explicit ratings—such as one to five stars—or implicit events such as purchases, clicks, saves, completed views, likes, or add-to-cart actions.

Explicit and implicit feedback

An explicit dataset might look like:

user_id,item_id,rating
u1,m1,5
u1,m2,3
u2,m1,4

Implicit data requires more caution. A purchase or completed view is positive behavioral evidence, but it does not prove that the user liked the item. More importantly, a missing interaction usually means “unknown,” not “negative.” Google’s recommendation documentation explains this distinction between explicit ratings and implicit signals: Google’s collaborative-filtering basics.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Repeated events should also be aggregated deliberately. Depending on the product, you might keep the latest rating, retain the maximum rating, average repeated ratings, or convert events into a weighted count. One simple policy for ratings is:

ratings = (
    ratings.sort_values("timestamp")
           .drop_duplicates(["user_id", "item_id"], keep="last")
)

For clicks or views, cap or transform counts so that one user’s repeated activity does not overwhelm every other signal:

weights = np.log1p(interaction_count)
# or
weights = interaction_count.clip(upper=5)

Build a cosine-similarity baseline in Python

Prerequisites

python -m venv .venv
source .venv/bin/activate        # macOS/Linux
# .venvScriptsactivate         # Windows
python -m pip install pandas numpy scipy scikit-learn

For a reproducible learning example, use a specific release of MovieLens from GroupLens. MovieLens releases differ in files and size, so record the exact variant you download rather than vaguely referring to “the MovieLens dataset.”

1. Load and normalize the data

import pandas as pd

ratings = pd.read_csv("ratings.csv")

ratings = ratings.rename(columns={
    "userId": "user_id",
    "movieId": "item_id"
})

ratings = ratings[["user_id", "item_id", "rating", "timestamp"]]
ratings = ratings.dropna(subset=["user_id", "item_id", "rating"])
ratings["user_id"] = ratings["user_id"].astype(int)
ratings["item_id"] = ratings["item_id"].astype(int)
ratings["rating"] = ratings["rating"].astype(float)
ratings["timestamp"] = pd.to_datetime(
    ratings["timestamp"], unit="s", errors="coerce"
)

print(ratings.shape)
print(ratings["user_id"].nunique())
print(ratings["item_id"].nunique())
print(ratings.isna().sum())
print(ratings["rating"].describe())

2. Construct the item–user matrix

A user–item matrix has users as rows and items as columns:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
User Item A Item B Item C Item D
User 1 1 1 0 0
User 2 1 0 1 0
User 3 0 1 1 1

Item-based filtering transposes that representation. Each item becomes a vector of users:

Item User 1 User 2 User 3
Item A 1 1 0
Item B 1 0 1
Item C 0 1 1
Item D 0 0 1

Items A and B have similar interaction patterns, so their item-to-item similarity is high.

For explicit ratings, a teaching implementation can use:

item_user = ratings.pivot_table(
    index="item_id",
    columns="user_id",
    values="rating",
    fill_value=0
)

However, zero usually means “missing,” not a genuine zero-star rating. For implicit feedback, a binary matrix is conceptually cleaner:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
interactions = ratings.assign(interaction=1)

item_user = interactions.pivot_table(
    index="item_id",
    columns="user_id",
    values="interaction",
    aggfunc="max",
    fill_value=0
)

Large datasets should use sparse matrices. A dense matrix with one row per item and one column per user can consume impractical amounts of memory even when most values are missing. See the SciPy sparse reference.

3. Calculate item-to-item cosine similarity

For item vectors i and j:

sim(i,j) = (i · j) / (||i|| ||j||)

Cosine similarity measures the angle between interaction vectors. It is easy to understand, works naturally with sparse data, and is a useful baseline—not a universally best metric.

from sklearn.metrics.pairwise import cosine_similarity
import numpy as np

item_similarity = cosine_similarity(item_user)
item_similarity = pd.DataFrame(
    item_similarity,
    index=item_user.index,
    columns=item_user.index
)

# An item should not recommend itself.
np.fill_diagonal(item_similarity.values, 0)

For binary vectors, popular items may share many users simply because they are popular. Minimum-support rules, popularity correction, and shrinkage are often needed before treating a similarity as trustworthy. The scikit-learn implementation is documented at cosine_similarity.

Other similarity metrics

  • Pearson correlation: useful for explicit ratings when users use different parts of the rating scale, but unstable with few co-ratings and less natural for one-way events.
  • Jaccard similarity: |A ∩ B| / |A ∪ B|; useful for binary adopter sets and ignores users who interacted with neither item.
  • Weighted or adjusted cosine: can downweight popular items, apply time decay, or give purchases more weight than clicks.

Generate personalized recommendations

For a user with history Hu, score candidate item j as:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

score(u,j) = Σ w(u,i) × sim(i,j)

Here, w(u,i) is the interaction weight. For an explicit-rating predictor, a normalized alternative is:

r̂(u,j) = Σ sim(i,j)r(u,i) / Σ |sim(i,j)|

For binary implicit interactions, a simple similarity sum is a good starting point.

def recommend_for_user(
    user_id,
    ratings,
    item_similarity,
    n_recommendations=10,
    min_similarity=0.0
):
    user_history = ratings[ratings["user_id"] == user_id]

    if user_history.empty:
        return pd.DataFrame(columns=["item_id", "score"])

    seen_items = set(user_history["item_id"])
    candidate_scores = {}

    for _, row in user_history.iterrows():
        source_item = row["item_id"]
        if source_item not in item_similarity.index:
            continue

        for candidate_item, similarity in item_similarity.loc[source_item].items():
            if candidate_item in seen_items or similarity <= min_similarity:
                continue

            weight = row.get("rating", 1.0)
            candidate_scores[candidate_item] = (
                candidate_scores.get(candidate_item, 0.0)
                + float(similarity) * float(weight)
            )

    return (
        pd.DataFrame(candidate_scores.items(), columns=["item_id", "score"])
          .sort_values("score", ascending=False)
          .head(n_recommendations)
          .reset_index(drop=True)
    )
recommendations = recommend_for_user(
    user_id=1,
    ratings=ratings,
    item_similarity=item_similarity,
    n_recommendations=10
)
print(recommendations)

The critical detail is seen_items. Without that exclusion, the system will often return items the user already rated, purchased, or watched.

Add item names and explanations

Keep metadata separate from the collaborative model unless you intentionally build a hybrid recommender:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
movies = pd.read_csv("movies.csv")

recommendations = recommendations.merge(
    movies.rename(columns={"movieId": "item_id"}),
    on="item_id",
    how="left"
)

Item-based recommendations can provide a useful explanation such as “Because you interacted with Item A, we recommend Item B.” More precisely, the explanation means that users who interacted with both items created a strong behavioral relationship. It is not a causal claim and does not prove that the user will like the recommendation.

Improve the baseline

Require interaction support

One shared user may create a misleading similarity. Shrink low-support scores toward zero:

Rank #4
Carson Dellosa The 100 Series: Biology Workbook—Grades 6-12 Science, Matter, Atoms, Cells, Genetics, Elements, Bonds, Classroom or Homeschool Curriculum (128 pgs)
  • Great extension activities for science and biology
  • Correlated to standards
  • Comprehensive biology vocabulary study
  • Fascinating true-to-life illustrations

shrunk_sim(i,j) = sim(i,j) × nij / (nij + λ)

nij is the number of users who interacted with both items, and λ controls the penalty. Also consider minimum item-interaction and shared-user thresholds.

Store only top-K neighbors

The naïve approach computes all item pairs, with approximate cost O(I²U), and stores an I × I matrix. That is often the real bottleneck. A production system usually stores only the strongest neighbors:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
item_id neighbor_id similarity
A B 0.82
A C 0.64
A D 0.51

Calculate only pairs with shared users where possible, retain top K neighbors per item, and consider approximate nearest-neighbor methods for very large catalogs. Scikit-learn’s nearest-neighbor utilities are documented here.

Weight behavior and recency

Purchases, saves, completed views, and clicks do not carry identical intent. Assign event weights, cap repeated actions, and optionally apply time decay:

wtime = e−γΔt

Recency can improve freshness but may hurt users whose long-term interests are stable. Similarities can be refreshed hourly or daily even when recommendations are served in real time; serving latency and model-update frequency are separate concerns.

Control popularity and diversity

Popular items can dominate because they co-occur with many other items. Mitigations include inverse-popularity weighting, category or brand caps, diversity constraints, freshness boosts, and blending niche candidates with popular ones. Evaluate whether recommendations are concentrated in a small part of the catalog.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Evaluate with a time-aware holdout

Do not use a random split by default. Randomly mixing past and future events can leak future co-occurrences into training. A simple temporal split holds out each user’s latest interaction:

ratings = ratings.sort_values(["user_id", "timestamp"])

test = ratings.groupby("user_id").tail(1)
train = ratings.drop(test.index)

Do not build the similarity matrix with test-period interactions. Users with only one interaction need a stated policy: exclude them from personalized evaluation, evaluate them separately as cold-start users, or retain them in training under a different protocol.

Useful metrics

  • Precision@K: the fraction of the top K recommendations that are relevant.
  • Recall@K: the fraction of held-out relevant items recovered in the top K.
  • Hit rate: the share of users with at least one held-out item in the top K.
  • NDCG@K: rewards relevant items more when they appear near the top.
  • Coverage: the proportion of the catalog the system can surface.
  • Diversity and novelty: detect repetitive lists and overreliance on popular items.
def precision_at_k(recommended_items, relevant_items, k):
    recommended = recommended_items[:k]
    relevant = set(relevant_items)

    if not recommended:
        return 0.0

    hits = sum(item in relevant for item in recommended)
    return hits / len(recommended)

Compare against a popularity baseline. Offline accuracy does not automatically predict revenue, retention, satisfaction, or long-term engagement because exposure, ranking position, novelty, margins, and feedback loops affect real outcomes. Coverage is commonly defined as the share of unique catalog items that may be recommended; Amazon documents this metric in its recommender evaluation guidance: AWS evaluation metrics.

Production architecture

event tracking
      ↓
data validation and aggregation
      ↓
offline similarity job
      ↓
top-K neighbor store
      ↓
online candidate generation
      ↓
business and safety filters
      ↓
ranking
      ↓
recommendation API or cache
      ↓
impression and outcome logging

A classroom pivot table does not solve identity, event quality, retraining, availability, monitoring, or policy enforcement. Production systems should use sparse storage, scheduled or incremental similarity updates, cached candidate neighborhoods, and post-score filters for consumed items, stock, geography, age restrictions, expired content, explicit rejections, and account policies.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Cold starts and failure modes

New users

A user with no history cannot receive personalized item-based recommendations. Use a fallback hierarchy such as regional or category popularity, editorial selections, contextual recommendations, onboarding preferences, or content-based results.

New items

A new item has no behavioral vector. Use metadata-based similarity, exploration traffic, editorial placement, popularity priors, or a hybrid model. Amazon Personalize documents that its Similar-Items recipe uses interaction co-occurrence and can incorporate item metadata; it may return popular items when the requested item is unknown: Similar-Items recipe.

Feedback loops and bias

If the system recommends only items that already receive interactions, those items gather more data while unseen items remain invisible. Use exploration quotas, randomized candidate injection, freshness boosts, and separate reporting for exposed and unexposed items. Historical interactions can also encode popularity, demographic, geographic, and editorial biases.

When cosine similarity is insufficient

Use Pearson correlation for carefully handled explicit ratings when rating-scale differences matter. Use Jaccard for certain binary-set problems. Move toward implicit-feedback models, matrix factorization, neural retrieval, or hybrid recommenders when behavior is extremely sparse, context changes rapidly, item meaning is important, or new-item cold start dominates.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The open-source Implicit library is designed for implicit-feedback models and large sparse data. Surprise is useful mainly for classic explicit-rating experiments and is not, by itself, a complete production foundation for large implicit event streams.

Self-hosted code or managed service?

Start with pandas, NumPy, SciPy, and scikit-learn when learning, prototyping, working with a small or medium catalog, or needing complete control over scoring and filtering. Move to sparse top-K computation as memory and latency become constraints, then consider a specialized library when interaction volume justifies additional complexity.

Amazon Personalize is a managed alternative for teams that value APIs, retraining workflows, real-time or batch recommendations, and reduced ML operations. Its Similar-Items API accepts an item ID for related-item recommendations: GetRecommendations API. It is unnecessary overhead for a small offline experiment whose goal is to understand cosine similarity. Service pricing, quotas, and free-tier terms are date- and region-sensitive, so consult the official pricing page before adopting it.

Implementation checklist

  • Are explicit and implicit signals modeled differently?
  • Is missing implicit data treated as unknown rather than dislike?
  • Are duplicate events aggregated deliberately?
  • Are interactions stored sparsely at scale?
  • Are low-support similarities suppressed?
  • Are consumed items excluded?
  • Are unknown users and new items handled with fallbacks?
  • Are future events excluded from evaluation?
  • Are popularity, coverage, diversity, and freshness monitored?
  • Are business, safety, availability, and policy filters applied?
  • Can each recommendation be traced to supporting history?
  • Are impressions and outcomes logged for unbiased improvement?

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.