Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PC×
Skip to content
caching

Python Cache: How to Speed Up Your Code With Effective Caching Techniques

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The fastest safe way to speed up repeated Python work is to cache only deterministic results, give every entry a key that represents all inputs, and choose a scope that matches your deployment. Start with functools.lru_cache for one-process memoization, move to Django’s cache framework for web responses, and use Redis or Memcached when workers or hosts must share entries. Add finite TTLs, explicit invalidation, stampede protection, and hit/miss monitoring before calling the job done.

What Python caching does

A cache stores the result of expensive or repeated work so a later request can reuse it instead of repeating the computation or I/O. The cached value is derived data: your database, API, or calculation remains the source of truth. A useful cache therefore has three properties:

  • The operation is safe to repeat and its result can be reused.
  • The key contains every input that can change the result.
  • There is a deliberate policy for expiry, invalidation, and failure.

Caching can reduce latency and backend load, but it is not an automatic speedup. Serialization, network hops, lock contention, oversized entries, low hit rates, and stale-data bugs can make an application slower than an uncached path.

Choose the cache scope first

Technique Scope and latency Best fit Main limitation
functools.lru_cache Process-local; lowest setup and read latency Deterministic functions called repeatedly with hashable arguments Entries are not shared by worker processes or hosts
Memoization library (for example, cachetools) Usually process-local collections or decorators Applications needing an eviction policy or interface beyond the standard decorator Additional dependency and policy choices to operate
Django cache framework Per-site, per-view, template-fragment, or low-level; backend-dependent Django responses and application data Keys, variation, timeout, and backend behavior still need design
Redis or Memcached Shared over the network Several workers or hosts sharing a working set Network, serialization, capacity, and operational failure become part of the request path

Use local memoization when each process can calculate independently. Use a shared backend when a cache miss would otherwise be repeated by every worker or when all hosts must see the same value. Django’s documentation recommends its included backends unless you have a compelling reason to do otherwise; they are tested and documented together with the framework.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Start with functools.lru_cache

A bounded function cache

lru_cache keeps up to maxsize recent calls. It is thread-safe, but two threads can still perform the same underlying call when they miss concurrently. Arguments must be hashable, so strings, numbers, and tuples of hashable values work; mutable lists and dictionaries do not.

from functools import lru_cache

@lru_cache(maxsize=512)
def convert_currency(amount_cents: int, source: str, target: str) -> int:
    """Pure conversion using a versioned, in-process rate table."""
    rates = {("USD", "EUR"): 0.92, ("EUR", "USD"): 1.09}
    rate = rates[(source, target)]
    return round(amount_cents * rate)

print(convert_currency(1000, "USD", "EUR"))
print(convert_currency.cache_info())  # hits, misses, maxsize, currsize
convert_currency.cache_clear()       # use after a rate/configuration change

Keep the function free of side effects such as sending email, charging a card, or mutating a record. Include a data or configuration version in an argument when changing the underlying source, or call cache_clear() as part of the update. A bounded cache limits memory; an unbounded cache is appropriate only when the argument space is known to be small and stable.

Designing keys that remain correct

Every value that changes the result belongs in the arguments: locale, units, feature flags, tenant, permission context, and a data version are common examples. Normalise equivalent inputs before the cached call (for example, canonicalise a hostname) so one logical item does not consume several entries. Never use a key that omits authentication or tenant identity for personalised data.

When local memoization is the wrong tool

  • Multiple gunicorn, uWSGI, or container workers need to see the same entry.
  • The working set is larger than a process can safely hold.
  • You need a shared TTL or explicit deletion from another service.
  • The function’s result depends on mutable state that you cannot version or invalidate.

Add TTLs and invalidation deliberately

A small TTL wrapper for Python functions

lru_cache is an eviction policy, not a freshness policy. If a value changes, add a time bucket to the key or use a backend with native expiry. This compact wrapper creates a finite lifetime while retaining LRU behavior:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from functools import lru_cache
from time import monotonic

TTL_SECONDS = 60

def ttl_cached(seconds: int, maxsize: int = 512):
    def decorate(func):
        @lru_cache(maxsize=maxsize)
        def timed(bucket, *args, **kwargs):
            return func(*args, **kwargs)

        def wrapper(*args, **kwargs):
            bucket = int(monotonic() // seconds)
            return timed(bucket, *args, **kwargs)

        wrapper.cache_clear = timed.cache_clear
        wrapper.cache_info = timed.cache_info
        return wrapper
    return decorate

@ttl_cached(60, maxsize=256)
def load_exchange_rates(region: str):
    return fetch_rates_from_source(region)  # your source-of-truth call

The bucket makes old entries unreachable, so memory can temporarily contain more than one period. Clear the cache after an urgent correction rather than waiting for the next boundary. For high-cardinality arguments, a backend TTL that deletes expired keys is usually easier to control.

Choose timeout values by freshness

Django documents a default backend timeout of 300 seconds, None for no expiry, and 0 for immediate expiry. Those are configuration semantics, not universal recommendations. Set a timeout from the business rule: prices and permissions usually need shorter windows than a slowly changing country list. Pair TTL with event-driven invalidation when stale data is unacceptable.

Use Django’s cache framework for web applications

Configure a backend

# settings.py
CACHES = {
    "default": {
        "BACKEND": "django.core.cache.backends.redis.RedisCache",
        "LOCATION": "redis://127.0.0.1:6379/1",
        "TIMEOUT": 300,
    }
}

Django supports local-memory, database, filesystem, Memcached, Redis, and custom backends. Local memory is thread-safe but private to each process and uses LRU culling, so it is not a shared cache. Filesystem, database, and local-memory backends expose capacity controls such as MAX_ENTRIES and CULL_FREQUENCY.

Cache an entire view

# urls.py
from django.urls import path
from django.views.decorators.cache import cache_page
from . import views

urlpatterns = [
    path("catalog/", cache_page(120)(views.catalog)),
]

Whole-view caching is appropriate for responses that are identical for every eligible visitor. If output varies by language, authentication, tenant, device, or request headers, ensure the variation is represented in the cache key and response Vary policy. URL-only caching can expose one user’s response to another.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Cache a low-level query

from django.core.cache import cache

def country_list():
    key = "countries:v3"
    value = cache.get(key)
    if value is None:
        value = list(Country.objects.values("code", "name"))
        cache.set(key, value, timeout=3600)
    return value

def publish_country_change():
    cache.delete("countries:v3")

Versioned keys make bulk invalidation simple: increment the version when the schema or source changes. Always treat a cache miss as a normal path unless your design explicitly guarantees a preloaded working set.

Move to Redis when workers must share entries

Prefetch a reference-data working set

Redis’s prefetch pattern loads reference data before traffic arrives, serves reads from Redis, synchronises mutations, deletes keys when records are deleted, and applies a safety-net TTL. The guide describes near-100% hit ratios for reference and master data and sub-millisecond lookup reads at peak traffic for that pattern; those are pattern-specific figures, not guarantees for every deployment.

import json
import redis

r = redis.Redis.from_url("redis://localhost:6379/0", decode_responses=True)
PREFIX = "product:v1:"


def prefetch(products):
    pipe = r.pipeline()
    for product in products:
        pipe.setex(PREFIX + str(product["id"]), 3600, json.dumps(product))
    pipe.execute()


def get_product(product_id):
    raw = r.get(PREFIX + str(product_id))
    if raw is not None:
        return json.loads(raw)
    # Choose fallback behavior explicitly: load from the source, or fail if
    # your service promises that the working set was preloaded.
    product = load_product_from_database(product_id)
    r.setex(PREFIX + str(product_id), 3600, json.dumps(product))
    return product


def delete_product(product_id):
    delete_product_from_database(product_id)
    r.delete(PREFIX + str(product_id))

Redis adds a network round trip and serialization cost, so compare the complete miss and hit paths rather than assuming it is faster than local memory. Set memory limits and an eviction policy that match your working set, and decide whether a backend outage should fall back to the source of truth. The Redis prefetch design intentionally treats an absent key as an error because its contract is “every read is a cache hit”; that is a design choice, not a general rule.

Prevent stampedes, stale reads, and unsafe values

Coalesce concurrent misses

When a popular key expires, many requests can perform the same expensive load. Use a per-key lock, request coalescing, or a single-flight pattern so one request populates the value while others wait. Add randomized TTL jitter for large fleets so thousands of keys do not expire simultaneously. Measure duplicate loads; Python’s own documentation warns that concurrent misses can call the wrapped function more than once.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Protect serialization boundaries

Django’s filesystem backend serializes values with pickle. Protect cache directories from writes by untrusted users: a malicious cache file can falsify trusted HTML or execute code when unpickled. Treat serialized cache data as untrusted at every boundary, and do not put secrets in keys or values unless the backend and access controls are designed for them.

Make fallback behavior explicit

For non-critical derived data, a cache outage should normally degrade to the source of truth with bounded timeouts and rate limits. For a preloaded reference-data service, failing closed may be safer than serving incomplete data. Document this choice per cache rather than relying on a generic retry loop.

Measure whether caching helped

Record hit rate, miss rate, load latency, eviction count, key cardinality, memory use, backend errors, duplicate loads, and stale-read incidents. Compare p50 and tail latency for hits and misses separately. A high hit rate can still hide a slow cache or oversized serialization; a low hit rate may indicate a key that includes an unnecessary timestamp or user-specific dimension. There is no universal percentage speedup for Python caching: the result depends on the work avoided, data size, backend, and traffic pattern.

Troubleshooting checklist

Symptom Likely cause Fix
Repeated calls never hit Arguments differ in type, normalisation, or an omitted dimension Inspect cache_info(), canonicalise inputs, and include every result-changing input.
Memory grows continually Unbounded cache or high-cardinality keys Set a finite maxsize, reduce key cardinality, and monitor entry size.
Users see another user’s response Shared response key ignores authentication, tenant, language, or Vary Separate personalised and public caches and add the missing dimensions.
Fresh updates appear late TTL is longer than the freshness requirement or invalidation was skipped Delete/version keys on writes and choose a finite timeout.
CPU spikes at expiry Many requests miss one hot key simultaneously Add locking or single-flight loading and jitter expirations.
Redis cache is slower than the function Network and JSON/pickle serialization cost exceed saved work Keep tiny hot values local, batch operations, or cache a more expensive computation.
Cache errors break requests No defined fallback or overly broad retry policy Set connection timeouts, choose fail-open or fail-closed per data class, and rate-limit source reloads.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup:

If the expensive operation in your pipeline is rendering a webpage, you can cache the resulting asset while letting ScreenshotNeo handle capture. It accepts one GET request and returns PNG, JPEG, WebP, or PDF. Before capture it accepts cookie/consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

See the ScreenshotNeo API documentation for output and option details. The same request with cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo includes full-page and selector capture, dark mode, device presets, retina scale, PDF controls, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, user agent, timezone, geolocation, resizing, configurable caching TTLs, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage APIs, and an OpenAPI specification. Every feature is on every plan: 1,000 shots per month are free with no card; paid plans start at $5 for 3,000 shots, with yearly billing providing two months free. Create a free ScreenshotNeo account.

Frequently Asked Questions

How should I test that a cache is correct rather than merely fast?

Run the same request before and after each relevant source update, vary every key dimension (such as tenant and language), and assert that an invalidation or version change returns the new value. Include tests for misses, expiry, backend outages, and concurrent first requests.

Should I cache exceptions?

Usually no: transient failures can become persistent outages when cached. If a negative result is valid, cache it with a short, explicitly documented TTL and distinguish it from an operational error.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When is a stale value acceptable?

Define the tolerance with the owner of the data. If stale reads can affect authorization, payments, or legal content, use short TTLs plus write-triggered invalidation or bypass the cache for the critical decision.

The Bottom Line

Use lru_cache for bounded, process-local memoization; Django’s framework for response and application caches; and Redis or Memcached for shared working sets. Correct keys, finite TTLs, invalidation, stampede control, safe serialization, and measurement matter more than the brand of backend.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.