Hardware FixRecommendedDevice not working? Your driver may be the problemCheck updates for common hardware issues.Fix DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run Scan×
Skip to content
World desk7 min

Why Production Microservices Need Circuit Breakers (Plus a 50-Line Example)

Circuit breakers limit repeated calls to failing dependencies. Learn the closed, open, and half-open states, when to pair breakers with retries, and what production implementations need beyond a short code sample.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A circuit breaker stops a service from repeatedly waiting on a remote dependency that is failing or responding too slowly. It does not fix that dependency: it limits the damage while your application returns a controlled error or uses a safe fallback. Below, you’ll see how the three-state pattern works, how it differs from retries, and a small Python example that illustrates the core logic without pretending to be production-ready.

What is the circuit breaker pattern in microservices?

A circuit breaker is a guard around a potentially failing operation, such as an HTTP request to another service. It tracks selected failures and changes whether calls are allowed through. The name describes the behavior: after enough evidence of trouble, the breaker opens and blocks calls for a period instead of letting every caller wait on an operation likely to fail.

This matters because a slow or unavailable dependency can consume time and resources throughout a system. Waiting requests may occupy threads or connections; retries can add more traffic and network contention. AWS Prescriptive Guidance describes this as a risk of failure amplification, including pressure on database thread pools. A breaker reduces repeated waiting and calling. It does not restore the dependency or guarantee that the rest of the system stays healthy.

Closed: pass calls through

In the normal state, the breaker forwards calls and records the failures you designate as evidence of dependency trouble. A timeout or connection failure may qualify. An ordinary business response—such as “account not found” or a validation rejection—usually should not: it says something about the request, not necessarily the health of the dependency. Choose this classification deliberately.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Open: fail fast

When the configured failure rule is met, the breaker opens. It rejects subsequent calls promptly for the configured break period, rather than calling the struggling dependency. The caller must decide what that rejection means for the user or workflow: show an error, serve suitable cached data, use an alternate service, or defer work if the operation allows it.

Half-open: test recovery cautiously

After the break period, the breaker allows a limited test of the dependency. Success supplies evidence to close the breaker; failure opens it again and begins another recovery wait. Limiting those tests helps prevent a recovering service from receiving a sudden flood of traffic. The exact number of permitted tests and the evidence required to close are policy choices, not universal rules.

How do I implement a circuit breaker?

The following Python 3 sketch uses asyncio and a consecutive-failure threshold. It is a teaching example: callers supply the operation and decide which exceptions count as dependency-health failures. It has a single half-open probe at a time, but it is not a drop-in resilience library. The state is local to one process, and the snippet omits production telemetry, configuration, and comprehensive tests.

import asyncio
import time

class CircuitOpen(Exception):
    pass

class CircuitBreaker:
    def __init__(self, threshold=5, break_seconds=30):
        self.threshold = threshold
        self.break_seconds = break_seconds
        self.failures = 0
        self.open_until = 0.0
        self.half_open = False
        self.probing = False
        self.lock = asyncio.Lock()

    async def call(self, operation, is_failure):
        async with self.lock:
            now = time.monotonic()
            if self.open_until and now < self.open_until:
                raise CircuitOpen("dependency circuit is open")
            if self.open_until:
                self.half_open = True
            if self.half_open:
                if self.probing:
                    raise CircuitOpen("recovery probe already in progress")
                self.probing = True

        try:
            result = await operation()
        except Exception as error:
            async with self.lock:
                failed = is_failure(error)
                if failed:
                    self.failures += 1
                if self.half_open and self.probing:
                    if failed:
                        self.open_until = time.monotonic() + self.break_seconds
                    else:
                        self.open_until = 0.0
                        self.failures = 0
                    self.half_open = False
                    self.probing = False
                elif self.failures >= self.threshold:
                    self.open_until = time.monotonic() + self.break_seconds
            raise
        else:
            async with self.lock:
                self.failures = 0
                if self.half_open and self.probing:
                    self.open_until = 0.0
                    self.half_open = False
                    self.probing = False
            return result

For example, an application could pass a predicate that recognizes its HTTP client’s timeout and connection exceptions, but not application-level responses. The example re-raises the original operation exception; only a call rejected before the operation runs raises CircuitOpen. A caller should handle both outcomes intentionally.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The defaults of five failures and 30 seconds are illustrative values, not recommended settings. Microsoft’s .NET documentation gives a Polly example that opens after five consecutive qualifying faults for 30 seconds. That is a library-specific configuration example, not evidence that these values suit another service or that this Python sketch behaves like Polly.

How should you choose the breaker policy?

Thresholds and recovery behavior should reflect the protected dependency’s behavior and the cost of a failed call. A threshold that is too sensitive can block calls during ordinary noise; one that waits too long can allow callers to pile up while the dependency is failing.

Choose what counts as failure

Separate health-related faults from expected business outcomes. Timeouts, connection failures, and overload responses may be relevant, but they need not share a policy. For example, an overload status could merit a different threshold from a connection refusal. A breaker should not turn a valid “not found” response into evidence that a service is unavailable.

Choose how to measure failures

A simple breaker may open after consecutive failures. Other policies use failures in a time window or a failure ratio after a minimum number of calls. These approaches respond differently to traffic volume and intermittent faults, so their numbers are not interchangeable. Current Polly strategy documentation illustrates a sampling duration of two seconds, minimum throughput of two, and failure ratio of 0.5; those are example settings, not production recommendations.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose how recovery is tested

A timed half-open probe is a common recovery check. If a dependency’s recovery is unusually variable, an explicit health check or an operator-controlled reset may be more appropriate. Whichever approach you choose, constrain concurrent half-open tests so the check itself does not burden the recovering service.

Scope the breaker to the right resource

Track failures for the dependency or resource whose health you are measuring. Combining independent providers, shards, or endpoints can cause failures in one to block healthy calls to another. Also decide who owns the policy: an application library or an infrastructure layer such as a service mesh. Overlapping breakers with unclear ownership make behavior harder to understand.

When should you use a circuit breaker instead of retry?

Retry and circuit breaking address different situations. Retry makes a bounded repeat attempt when a fault may be transient. A circuit breaker suppresses calls after recent failures suggest that more attempts are likely to fail. Microsoft’s Azure Architecture Center makes this distinction explicitly.

They can be composed: a caller may make a small, bounded number of retries, while a breaker prevents further attempts when the dependency appears unhealthy. But retries must stop when the breaker reports an open circuit. Unbounded retries, or retries layered unknowingly in multiple components, can add load precisely when a service is least able to handle it. In message-driven systems, existing retry and dead-letter behavior may already serve part of this role, so account for it before adding another layer.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the breaker does not provide

Fallback is a separate application decision

A breaker can reject a call; it cannot decide what substitute is correct. A cache might be a reasonable fallback for a read where slightly stale data is acceptable. It may be unsafe for an update or command, where returning a default could imply work succeeded when it did not. Define fallback behavior according to the operation’s meaning.

Bulkheads limit a different kind of pressure

A bulkhead caps concurrent work or queued requests, limiting resource use before a pattern of failures has accumulated. A breaker reacts to failure evidence and blocks calls for a period. The mechanisms can complement each other, but neither is a substitute for understanding the other.

Exception handling is still necessary

A breaker does not replace error handling. Callers still need to handle operation failures and the open-circuit outcome, and decide what to expose to users or downstream workflows. Microsoft’s guidance also cautions that the implementation should not block concurrent requests or add excessive overhead to each call.

Production review checklist

  • Match the policy to the dependency: select a failure measure, observation period, threshold, and break duration based on its failure and recovery behavior. A long open period can keep a recovered service unavailable to callers; a short one can repeatedly probe a service still recovering.
  • Protect concurrency-sensitive state: keep calls nonblocking, make state changes safe under concurrency, and limit half-open probes. A compact example that does not provide these properties should not be treated as production-ready.
  • Plan the caller response: specify what callers do with an open-circuit result and whether a fallback is valid for this operation.
  • Instrument behavior: record successes, qualifying failures, and state transitions. Tracing helps connect a caller’s failure to the downstream operation; make state visible enough for operators to understand and, where justified, control recovery.
  • Check existing protection: account for retries, dead-letter handling, infrastructure, and service-mesh policies so the new breaker does not duplicate or conflict with another owner.
  • Test the transitions: cover threshold crossing, fast rejection while open, recovery success, recovery failure, non-qualifying exceptions, and concurrent calls around the half-open transition.

For a real deployment, prefer a maintained resilience library or a clearly owned infrastructure policy unless you have a strong reason to maintain a custom implementation. Verify the API generation and version you use: Microsoft’s older Polly/IHttpClientFactory example and current Polly resilience-strategy documentation describe different generations of guidance.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.