October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PCOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
aiolimiter

How to Rate Limit Async Requests in Python

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use a time-based limiter to control how many requests start over time; use an asyncio.Semaphore separately if you also need to cap how many requests are in flight. They solve different problems. For asyncio code, aiolimiter provides an AsyncLimiter you can place around each outbound request. Set its rate and burst capacity to match the API provider’s current quota, not an arbitrary example.

Rate limits and concurrency limits are different

A rate limit measures requests over time, such as 60 requests per minute. A concurrency limit measures simultaneous work, such as no more than 10 requests in flight. An async program can have low concurrency but still send requests too quickly in a short burst; it can also maintain a low average request rate while allowing too many slow requests to remain open.

asyncio.Semaphore limits simultaneous holders: acquiring it decrements its counter, and releasing it increments the counter. It does not enforce requests per second or minute. Python’s asyncio documentation recommends using a semaphore with async with so release occurs when the block exits.

For a time-based limit, aiolimiter supplies AsyncLimiter, which gates entry to a section of async code. Its algorithm is a leaky bucket: max_rate is also the maximum initial burst. That burst behavior matters when an API allows a sustained quota but restricts how many calls can arrive together.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Install aiolimiter and set the quota

Install the library in the Python environment that runs your async application:

python -m pip install aiolimiter

Choose the limiter’s values from the API provider’s current documentation. Limits may differ by endpoint, credential, account tier, or operation cost. The values below are examples, not a recommendation for any particular API:

import asyncio
from aiolimiter import AsyncLimiter

# Example only. Replace with the API's documented quota and burst policy.
requests_per_minute = 60
limiter = AsyncLimiter(requests_per_minute, 60)

# Optional: independently cap simultaneous in-flight requests.
concurrency = asyncio.Semaphore(10)

async def fetch(client, url):
    async with limiter:
        async with concurrency:
            response = await client.get(url)
            response.raise_for_status()
            return response

Put the limiter around the actual outbound operation so each request consumes capacity. The semaphore is optional and addresses a separate concern. If the remote service permits 60 requests per minute but your responses take a long time, a concurrency cap can keep your application from accumulating too many open operations.

Choose the acquisition order deliberately

The example above acquires rate capacity before waiting for a concurrency slot. If every concurrency slot is occupied, a task can pass through the limiter and then wait; by the time a slot opens, its rate capacity may have been consumed earlier than the network request actually starts. That makes this ordering simple but not ideal for every workload.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Reversing the order avoids reserving rate capacity before a concurrency slot is available, but it holds that slot while waiting for the limiter. That reduces the slots available to other tasks. Select based on which resource you need to protect most, and measure behavior under your actual workload rather than assuming either order is universally optimal.

async def fetch_slot_first(client, url):
    async with concurrency:
        async with limiter:
            response = await client.get(url)
            response.raise_for_status()
            return response

For many producers, fairness, or explicit backpressure, a queue-based dispatcher can be a better design: producers enqueue work, while a controlled set of workers obtains rate capacity and performs requests. A queue also gives you a place to bound pending work instead of allowing an unbounded number of tasks to wait on a limiter or semaphore.

Control bursts and pacing

With AsyncLimiter(max_rate, time_period), the first argument sets the maximum amount of capacity available and therefore the maximum initial burst. For example, AsyncLimiter(60, 60) can permit an initial burst of up to 60 entries, then replenishes capacity over the period. This is not equivalent to a strict rule that spaces every request exactly one second apart.

If the API does not permit a burst and you want one entry per interval, the aiolimiter documentation shows using AsyncLimiter(1, interval_seconds). For example, AsyncLimiter(1, 1.5) spaces entries by about 1.5 seconds. Confirm that this pacing model matches the provider’s definition of its quota.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some operations may have different costs. aiolimiter supports acquiring a weighted amount, but do this only when the API assigns different quota costs to operations. Its documentation warns that smaller-capacity requests can be favored over larger ones near capacity, so mixed weights may not behave like a fair first-in, first-out queue.

Keep the limiter aligned with the event loop and the whole application

Create an AsyncLimiter for the event loop that uses it. The aiolimiter documentation says reuse across event loops is unsupported and can produce undefined behavior. Avoid putting a module-global limiter in code that is shared across separately created loops, test cases, or application lifecycles unless its ownership is carefully controlled.

A local limiter only controls calls routed through that limiter instance. If multiple processes, containers, or machines use the same API credential, each local instance can independently spend capacity; the combined traffic may exceed a shared provider quota. The cited in-process limiter documentation does not establish distributed coordination. A global quota therefore requires a separately designed shared-state or centralized coordination mechanism.

Likewise, a limiter does not automatically handle HTTP 429 responses, provider-directed retry delays, or transient network errors. Consult the API provider’s own current rate-limit and retry documentation. Treat any Retry-After header according to that provider’s documented behavior rather than assuming a universal policy.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Keep async waits non-blocking

Use async-aware waiting and await network operations. Do not call blocking time.sleep() in an event-loop task to implement pacing: it blocks the event loop and can stall unrelated coroutines. A library limiter handles waiting asynchronously. If you implement a custom limiter, use monotonic timing, account for cancellation correctly, and test boundary conditions; rolling your own scheduling logic adds edge cases that a library avoids.

Other limiter algorithms to consider

The asynciolimiter documentation describes three approaches: Limiter accounts for delays such as CPU-heavy work; LeakyBucketLimiter supports a maximum capacity and initial burst; and StrictLimiter prevents bursts and keeps the resulting rate below its configured rate. Its documentation suggests the regular Limiter if unsure. That documentation is older than the aiolimiter and Python references, so check the current package documentation and API version before relying on its installation or usage examples.

Choose based on the provider’s actual rule: permitted burst size, whether missed time should be compensated, whether strict pacing is necessary, whether operation costs vary, and whether rate state must be shared across workers. Do not treat a local algorithm as a cross-process quota controller.

Troubleshoot common rate-limiting problems

  • You still receive HTTP 429 responses. The configured limit may not match the provider’s current quota, endpoint-specific rules, or shared credential usage across other processes. Check the provider’s documentation and response guidance; a local limiter only sees traffic that uses its instance.
  • Requests arrive in a burst at startup. With aiolimiter, max_rate is the maximum initial burst. Lower it to match the provider’s burst policy; for one-at-a-time pacing, use a maximum rate of 1 and the desired interval.
  • Throughput is lower than expected. Check whether the semaphore is saturated, whether the limiter’s period and burst match your intended quota, and whether you acquire a semaphore slot while waiting for rate capacity. Slow upstream responses can make a concurrency cap the bottleneck even when rate capacity remains.
  • Tasks wait longer than expected after cancellation or errors. Keep limiter and semaphore scopes in async with blocks, and await requests. These context managers release their acquired resources when execution leaves the block, including when an exception unwinds it.
  • Tests behave inconsistently across event loops. Do not reuse one limiter across separate event loops. Construct it within the loop or application lifecycle that will use it.
  • A home-grown limiter drifts or stalls unrelated tasks. Avoid blocking sleeps; use an async library. If custom logic is necessary, use monotonic time and test cancellation, exact interval boundaries, bursts, and simultaneous callers.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your async workflow also needs website screenshots, ScreenshotNeo is a screenshot API and MCP server from Yorker Media. A single GET request takes a URL and returns an image or PDF; the Python example below runs inside an async function using a thread so the synchronous requests call does not block the event loop. This screenshot call is separate from the limiter pattern above: wrap it in the same kind of rate limiter only if your own ScreenshotNeo usage policy or workload calls for it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

See the ScreenshotNeo API documentation for request parameters.

import asyncio
import requests

async def capture():
    def request_screenshot():
        r = requests.get(
            "https://api.screenshotneo.com/v1/shot",
            params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
            timeout=90,
        )
        r.raise_for_status()
        with open("shot.webp", "wb") as image:
            image.write(r.content)

    await asyncio.to_thread(request_screenshot)

asyncio.run(capture())
  • Cookie and consent banners, newsletter popups, and chat widgets are removed before capture; each cleanup step can be turned off.
  • Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; responses identify the page verdict and billing status in headers.
  • An MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs.
  • The free plan includes 1,000 screenshots a month without a card; paid plans start at $5 for 3,000 shots.

Sign up for ScreenshotNeo’s free plan to try it with 1,000 screenshots a month and no card.

Frequently Asked Questions

Does an asyncio semaphore rate limit requests?

No. A semaphore limits simultaneous holders; use a time-based limiter such as aiolimiter for entries over time.

Can one aiolimiter instance be shared between event loops?

No. aiolimiter documents cross-loop reuse as unsupported; create the limiter for the loop that uses it.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does a local limiter enforce a quota across multiple processes?

No. Each local instance controls only work routed through it; coordinating a shared quota requires a separate shared mechanism.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.