DriversRecommendedOutdated drivers can make a good PC feel brokenScan driver issues before chasing fixes manually.Scan NowOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PC×
Skip to content
World desk6 min

When Should You Actually Worry About a Growing Replication Queue? A PostgreSQL Decision Guide

A growing replication queue needs attention when standby delay exceeds what your users or failover process can tolerate, or when retained WAL threatens disk headroom. Here is how to tell the difference in PostgreSQL.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Worry when the standby’s delay has grown past what its users or failover process can tolerate, or when WAL retained for replication is consuming disk headroom and the trend is still rising. A backlog that is growing but still inside a defined freshness objective, with plenty of free disk, is a trend to watch rather than an incident.

This guide uses PostgreSQL physical streaming replication as its concrete example. Metric names, field meanings, and alert thresholds do not carry over unchanged to MySQL, Kafka, or managed database migration services, so check their own documentation before applying anything here to those systems.

What the “queue” is in PostgreSQL

PostgreSQL does not keep a visible queue of pending replication work. The backlog is the write-ahead log (WAL) that the primary has generated but the standby has not yet written, flushed, or replayed. You see it in two forms: as time-based lag values in the pg_stat_replication view, and as a byte difference between WAL positions. Both are useful, and they answer different questions.

One scope limit matters when you read that view. pg_stat_replication on a primary lists only standbys connected directly to it. A cascaded standby that streams from another standby appears in the view on the intermediate server, not on the original primary.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the lag columns tell you, and what they do not

According to the PostgreSQL Global Development Group’s monitoring documentation for PostgreSQL 19, the lag fields describe recent WAL progress. They are write_lag, flush_lag, and replay_lag, and each measures how long recent WAL took to reach that stage on the standby. For an asynchronous standby, the documentation says that replay_lag approximates the delay before recent transactions became visible to queries. That is the number most closely tied to read freshness.

The lag is not a catch-up estimate

The same documentation states: “The reported lag times are not predictions of how long it will take for the standby to catch up with the sending server assuming the current rate of replay.” Do not divide the backlog by the current replay speed and present the result as a finish time. A standby can be replaying faster than WAL arrives, or slower, and a single lag reading does not reveal which.

NULL can mean caught up

When a standby has fully caught up and the primary is idle, the lag columns can eventually show NULL instead of zero. Treat NULL as “no recent WAL to measure” rather than as a broken monitor, and confirm it against the byte positions described below.

Set the threshold from your objective, not a default number

No official source gives a universal number of seconds or bytes at which replication should page someone. The PostgreSQL documentation explains the signals and the risks, but it does not prescribe an alert level, and no published benchmark in the reviewed material establishes one for general use. A threshold has to come from two local inputs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Freshness or recovery objective. How stale can a read replica be before users see wrong results? How long can a failover candidate lag before promotion becomes unacceptable? Write this down as a number with an owner.
  • Storage budget. How much free space on the volume holding pg_wal can you afford to lose to retained WAL before other work is affected?

Once you have the storage budget, a simple calculation shows how much time a stalled consumer leaves you. Using hypothetical figures: if 200 GB is free for WAL and the primary generates about 2 GB of WAL per hour, a fully stalled slot would fill that space in roughly 100 hours. If the same primary generates 20 GB per hour during a batch job, the same headroom lasts about 10 hours. Replace these figures with measurements from your own system, and remember that WAL volume varies with workload.

How to decide whether a growing backlog matters

  1. Confirm the impact. Identify what depends on the standby: read traffic, failover, analytics, or change data capture. Compare the observed delay with that dependency’s explicit tolerance. If nothing depends on the standby’s freshness and disk is safe, the backlog is a maintenance item.
  2. Check whether data is arriving. Run the query below on the primary several times over a few minutes. If sent_lsn and write_lsn are advancing but replay_lsn is not, the standby is receiving WAL and replaying it slowly. If sent_lsn itself has stopped moving, the problem is on the connection or the standby’s receiver side.
    SELECT application_name, state, sent_lsn, write_lsn, flush_lsn, replay_lsn,
           write_lag, flush_lag, replay_lag,
           pg_wal_lsn_diff(sent_lsn, replay_lsn) AS replay_backlog_bytes
    FROM pg_stat_replication;
  3. Compare generated WAL with replayed WAL over time. The rising byte difference between pg_current_wal_lsn() and the standby’s replay position is more meaningful than one lag reading. A backlog that rises during a bulk load and falls afterward is normal. A backlog that keeps rising after the workload has settled is not.
  4. Check replication slots and disk. Retained WAL can grow through slots even when the lag columns look modest. Use the slot query in the next section.
  5. Check the retention cap. If max_slot_wal_keep_size is set, confirm how close each slot is to it and what recovery path exists if the cap is crossed.

Replication slots and WAL retention

A replication slot ensures that the primary keeps WAL a consumer still needs. That protects standbys from losing required WAL during short outages, which is the point of slots. The cost appears when a consumer disconnects or stalls. The primary keeps the retained WAL, and according to PostgreSQL’s documentation, slots can retain enough WAL to fill the primary’s pg_wal space.

Check retention per slot with this query on the primary:

SELECT slot_name, slot_type, active, restart_lsn,
       pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn) AS retained_bytes
FROM pg_replication_slots;

An inactive slot with a steadily growing retained_bytes is the clearest warning sign in this area. It means WAL is accumulating for a consumer that is not currently reading it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

max_slot_wal_keep_size: a cap with a cost

The max_slot_wal_keep_size setting can bound the WAL that slots retain, and PostgreSQL applies that limit at checkpoint time. A cap protects the primary’s disk. It also introduces a failure mode: if required WAL is removed because a slot fell too far behind, the standby may no longer be able to continue replication from that slot and may need to be rebuilt.

Approach Disk risk on the primary Continuity risk for the standby What you must have in place
No cap Highest: a stalled slot can fill pg_wal Lowest while the standby stays reachable Alerting on retained_bytes and free space well before the volume fills
Cap set, with monitoring Bounded by the configured limit Moderate: a standby that exceeds the cap can lose its slot’s WAL Monitoring of how close each slot is to the cap, plus a documented rebuild procedure
Cap set, without a recovery plan Bounded High: a lost slot can force an unplanned rebuild during an incident Not recommended as a standalone setting

Treat the cap as a deliberate trade between storage and recoverability, not as cleanup. Decide which standby you are willing to rebuild, and set the cap so that a rebuild is an acceptable outcome for that standby.

Signals that call for action

Observation Usually means Response
Lag rising, replay_lsn still advancing Standby is replaying slower than WAL arrives Check standby load, I/O, and long-running replay conflicts against your freshness objective
sent_lsn not advancing Data is not reaching the standby Check connectivity, the walreceiver on the standby, and authentication before tuning replay
Lag NULL on an idle primary Standby is caught up with no recent WAL to measure Confirm with byte positions; no action if they match
Inactive slot with rising retained_bytes A consumer is gone or stalled while the primary keeps its WAL Restore the consumer or drop the slot after confirming nothing needs it
Free space on the pg_wal volume trending toward your budget Retained WAL is consuming disk faster than it is released Act now, not at the threshold; the time to exhaustion is the decision input

What this guide does not cover

The guidance above applies to PostgreSQL physical streaming replication. Logical replication, cascaded topologies, and managed PostgreSQL services can expose different views and retention behavior, so verify their own metrics before setting thresholds.

For authoritative definitions of each field and setting, use the PostgreSQL documentation for your exact major version, especially the monitoring and replication sections. The field semantics quoted here come from the PostgreSQL 19 documentation, and older versions may word them differently.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

“

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.