Worry when the standby’s delay has grown past what its users or failover process can tolerate, or when WAL retained for replication is consuming disk headroom and the trend is still rising. A backlog that is growing but still inside a defined freshness objective, with plenty of free disk, is a trend to watch rather than an incident.
This guide uses PostgreSQL physical streaming replication as its concrete example. Metric names, field meanings, and alert thresholds do not carry over unchanged to MySQL, Kafka, or managed database migration services, so check their own documentation before applying anything here to those systems.
What the “queue” is in PostgreSQL
PostgreSQL does not keep a visible queue of pending replication work. The backlog is the write-ahead log (WAL) that the primary has generated but the standby has not yet written, flushed, or replayed. You see it in two forms: as time-based lag values in the pg_stat_replication view, and as a byte difference between WAL positions. Both are useful, and they answer different questions.
One scope limit matters when you read that view. pg_stat_replication on a primary lists only standbys connected directly to it. A cascaded standby that streams from another standby appears in the view on the intermediate server, not on the original primary.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →#1 Best Overall
What the lag columns tell you, and what they do not
According to the PostgreSQL Global Development Group’s monitoring documentation for PostgreSQL 19, the lag fields describe recent WAL progress. They are write_lag, flush_lag, and replay_lag, and each measures how long recent WAL took to reach that stage on the standby. For an asynchronous standby, the documentation says that replay_lag approximates the delay before recent transactions became visible to queries. That is the number most closely tied to read freshness.
The lag is not a catch-up estimate
The same documentation states: “The reported lag times are not predictions of how long it will take for the standby to catch up with the sending server assuming the current rate of replay.” Do not divide the backlog by the current replay speed and present the result as a finish time. A standby can be replaying faster than WAL arrives, or slower, and a single lag reading does not reveal which.
Rank #2
NULL can mean caught up
When a standby has fully caught up and the primary is idle, the lag columns can eventually show NULL instead of zero. Treat NULL as “no recent WAL to measure” rather than as a broken monitor, and confirm it against the byte positions described below.
Set the threshold from your objective, not a default number
No official source gives a universal number of seconds or bytes at which replication should page someone. The PostgreSQL documentation explains the signals and the risks, but it does not prescribe an alert level, and no published benchmark in the reviewed material establishes one for general use. A threshold has to come from two local inputs.
Rank #3
- Freshness or recovery objective. How stale can a read replica be before users see wrong results? How long can a failover candidate lag before promotion becomes unacceptable? Write this down as a number with an owner.
- Storage budget. How much free space on the volume holding
pg_walcan you afford to lose to retained WAL before other work is affected?
Once you have the storage budget, a simple calculation shows how much time a stalled consumer leaves you. Using hypothetical figures: if 200 GB is free for WAL and the primary generates about 2 GB of WAL per hour, a fully stalled slot would fill that space in roughly 100 hours. If the same primary generates 20 GB per hour during a batch job, the same headroom lasts about 10 hours. Replace these figures with measurements from your own system, and remember that WAL volume varies with workload.
How to decide whether a growing backlog matters
- Confirm the impact. Identify what depends on the standby: read traffic, failover, analytics, or change data capture. Compare the observed delay with that dependency’s explicit tolerance. If nothing depends on the standby’s freshness and disk is safe, the backlog is a maintenance item.
- Check whether data is arriving. Run the query below on the primary several times over a few minutes. If
sent_lsnandwrite_lsnare advancing butreplay_lsnis not, the standby is receiving WAL and replaying it slowly. Ifsent_lsnitself has stopped moving, the problem is on the connection or the standby’s receiver side.SELECT application_name, state, sent_lsn, write_lsn, flush_lsn, replay_lsn, write_lag, flush_lag, replay_lag, pg_wal_lsn_diff(sent_lsn, replay_lsn) AS replay_backlog_bytes FROM pg_stat_replication; - Compare generated WAL with replayed WAL over time. The rising byte difference between
pg_current_wal_lsn()and the standby’s replay position is more meaningful than one lag reading. A backlog that rises during a bulk load and falls afterward is normal. A backlog that keeps rising after the workload has settled is not. - Check replication slots and disk. Retained WAL can grow through slots even when the lag columns look modest. Use the slot query in the next section.
- Check the retention cap. If
max_slot_wal_keep_sizeis set, confirm how close each slot is to it and what recovery path exists if the cap is crossed.
Replication slots and WAL retention
A replication slot ensures that the primary keeps WAL a consumer still needs. That protects standbys from losing required WAL during short outages, which is the point of slots. The cost appears when a consumer disconnects or stalls. The primary keeps the retained WAL, and according to PostgreSQL’s documentation, slots can retain enough WAL to fill the primary’s pg_wal space.
Check retention per slot with this query on the primary:
SELECT slot_name, slot_type, active, restart_lsn,
pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn) AS retained_bytes
FROM pg_replication_slots;
An inactive slot with a steadily growing retained_bytes is the clearest warning sign in this area. It means WAL is accumulating for a consumer that is not currently reading it.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchmax_slot_wal_keep_size: a cap with a cost
The max_slot_wal_keep_size setting can bound the WAL that slots retain, and PostgreSQL applies that limit at checkpoint time. A cap protects the primary’s disk. It also introduces a failure mode: if required WAL is removed because a slot fell too far behind, the standby may no longer be able to continue replication from that slot and may need to be rebuilt.
| Approach | Disk risk on the primary | Continuity risk for the standby | What you must have in place |
|---|---|---|---|
| No cap | Highest: a stalled slot can fill pg_wal |
Lowest while the standby stays reachable | Alerting on retained_bytes and free space well before the volume fills |
| Cap set, with monitoring | Bounded by the configured limit | Moderate: a standby that exceeds the cap can lose its slot’s WAL | Monitoring of how close each slot is to the cap, plus a documented rebuild procedure |
| Cap set, without a recovery plan | Bounded | High: a lost slot can force an unplanned rebuild during an incident | Not recommended as a standalone setting |
Treat the cap as a deliberate trade between storage and recoverability, not as cleanup. Decide which standby you are willing to rebuild, and set the cap so that a rebuild is an acceptable outcome for that standby.
Signals that call for action
| Observation | Usually means | Response |
|---|---|---|
Lag rising, replay_lsn still advancing |
Standby is replaying slower than WAL arrives | Check standby load, I/O, and long-running replay conflicts against your freshness objective |
sent_lsn not advancing |
Data is not reaching the standby | Check connectivity, the walreceiver on the standby, and authentication before tuning replay |
| Lag NULL on an idle primary | Standby is caught up with no recent WAL to measure | Confirm with byte positions; no action if they match |
Inactive slot with rising retained_bytes |
A consumer is gone or stalled while the primary keeps its WAL | Restore the consumer or drop the slot after confirming nothing needs it |
Free space on the pg_wal volume trending toward your budget |
Retained WAL is consuming disk faster than it is released | Act now, not at the threshold; the time to exhaustion is the decision input |
What this guide does not cover
The guidance above applies to PostgreSQL physical streaming replication. Logical replication, cascaded topologies, and managed PostgreSQL services can expose different views and retention behavior, so verify their own metrics before setting thresholds.
For authoritative definitions of each field and setting, use the PostgreSQL documentation for your exact major version, especially the monitoring and replication sections. The field semantics quoted here come from the PostgreSQL 19 documentation, and older versions may word them differently.
Quick Recap
“
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




