Track BullMQ failures at both the job and worker levels, but do not treat queue recovery as a rollback of application data. A processor error can be retried, a lost lock can cause a job to run again, and a Postgres transaction only covers the database operations inside it. To make a worker safer, capture enough context to diagnose each failure, distinguish retryable errors from permanent ones, shut workers down gracefully, and design application writes and external effects for the possibility of repeated execution.
What “rollback-safe” means for a BullMQ worker
There is no single transaction that automatically covers every part of a job. Keep the boundaries clear:
As an Amazon Associate I earn from qualifying purchases.
- BullMQ queue state: BullMQ manages job transitions and recovery using its configured backend. Its optional PostgreSQL backend documents queue-state transitions implemented as SQL functions within transactions.
- Application Postgres data: a transaction in your application covers only the database operations you include in that transaction. BullMQ’s queue transaction does not automatically include arbitrary application-table writes.
- External effects: an email, payment request, or other external call is not made atomic with a Postgres transaction merely because a worker initiated it.
So “rollback-safe” should describe a verified application design, not an assumed BullMQ feature. Before claiming that a failed job leaves no partial effects, identify which operations commit together and how the handler behaves if it runs again after a retry or stalled-job recovery.
Redis queue plus application Postgres
If BullMQ uses Redis while the worker writes to application Postgres, treat queue state and application data as separate systems. A job can be retried or recovered independently of the fate of a previous database transaction or external call. Document the boundary and verify how duplicate execution is handled in your application; the BullMQ queue’s recovery mechanisms do not themselves prove that application side effects will happen exactly once.
#1 Best Overall
BullMQ’s optional PostgreSQL backend
BullMQ’s official PostgreSQL backend documentation specifies PostgreSQL 13 or newer, recommends 14 or newer, and requires the pg package. It describes transactional queue-state changes, not a general distributed transaction spanning queue state, application tables, and external systems.
What should error tracking capture?
Record enough structured context to connect an exception to the job and the work that initiated it. A practical log entry can include:
Rank #2
- Queue name, job name, and stable job identifier.
- Attempt information, if available in the processor context.
- Error class, message, stack, and timestamp.
- A correlation identifier linking the job to its initiating request or domain record.
This is an implementation recommendation, not a schema prescribed by BullMQ. Avoid logging the whole job payload by default: BullMQ’s production guidance says job data is stored in clear text. Keep sensitive values out of payloads where possible, or encrypt sensitive fields before enqueueing. Retain failed-job records for diagnosis according to your retention policy, while considering the sensitivity and storage implications of keeping them.
Which BullMQ errors need separate handling?
A processor exception, a BullMQ operational error, and a stalled job are different signals. Treating them as one generic “job failed” event makes it harder to diagnose whether work needs retrying, infrastructure needs attention, or the worker stopped renewing a lock.
Rank #3
| Situation | What it indicates | How to handle it |
|---|---|---|
| Processor throws an ordinary error | The job’s processing attempt failed and may follow the configured retry behavior. | Log the job context and error; classify whether retrying is appropriate. |
Processor throws UnrecoverableError |
BullMQ treats the failure as permanent and bypasses the configured retries. | Use it when the processor has determined that another attempt should not be made; the job goes to the failed set. |
Worker or queue emits error |
An operational signal that can include connection issues; it is not necessarily a processor exception. | Attach an event handler and route the error to structured logs or monitoring. |
| Job becomes stalled | The worker did not renew the active-job lock as expected; the job may return to waiting or eventually enter the failed set after its allowed stalls. | Investigate worker health and event-loop responsiveness separately from ordinary processor retries. |
Attach handlers to both Worker and Queue
BullMQ’s production guidance recommends handling error events on both objects. For example, route each error to your application logger rather than leaving the event unhandled:
worker.on('error', (err) => logger.error({ err }, 'BullMQ worker error'));
queue.on('error', (err) => logger.error({ err }, 'BullMQ queue error'));
Adapt the fields to your logger’s structured-error format. These listeners are operational safeguards; they do not replace processor-level error classification, failed-job retention, or alerts that tell an operator when intervention is needed.
How do retries and stalled-job recovery differ?
Normal retries respond to processor failures according to the retry configuration. A stalled job is a lock-renewal problem: BullMQ expects an active worker to keep signaling that it is processing the job. If the lock is not renewed, BullMQ can treat the job as stalled and return it to waiting, or fail it after the allowed number of stalls.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteThat distinction matters for diagnosis. A permanent input or business-rule failure may call for UnrecoverableError; a worker that is blocking the event loop may instead need a code or workload change. Stalled-job recovery can cause work to run again, so application side effects must be safe to encounter on a subsequent execution.
Keep the Node.js event loop available
Long synchronous, CPU-heavy work can block the event loop and interrupt lock renewal. BullMQ’s stalled-job documentation warns that workers need to return control to the Node.js event loop often enough to avoid this problem. Use work that yields appropriately, or isolate CPU-heavy processing using a suitable process or thread design. Confirm the approach against the BullMQ version and worker implementation you actually deploy; do not assume a specific isolation API without verifying it.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How should a worker shut down?
On termination, close the worker as part of service cleanup. BullMQ documents that await worker.close() stops the worker from taking new jobs and waits for active work to finish or fail. The method has no built-in timeout, so the deployment’s termination grace period and the maximum duration of active jobs need to be compatible.
process.on('SIGTERM', async () => {
await worker.close();
process.exit(0);
});
Adapt signal handling to your service lifecycle and ensure cleanup errors are handled by the application. Graceful shutdown reduces the likelihood of jobs being interrupted; stalled-job recovery remains a separate safeguard for ungraceful shutdowns or workers that stop renewing locks.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsShould queue state live in Redis or PostgreSQL?
BullMQ’s official PostgreSQL backend documentation presents PostgreSQL as an option for teams that want job state alongside relational data or want to avoid operating a separate Redis service. It also says: “The Redis backend remains the default and the most battle-tested option.” Choose based on your operating environment and measured workload, not on an assumption that the PostgreSQL backend makes application writes atomic with queue acknowledgement.
| Operation in BullMQ’s published illustrative benchmark | PostgreSQL | Redis |
|---|---|---|
Sequential add() |
About 7,000 jobs/s | About 7,500 jobs/s |
Concurrent add() |
About 15,000 jobs/s | About 38,000 jobs/s |
Concurrent bulk addBulk() |
About 45,000 jobs/s | About 52,000 jobs/s |
| Processing, one worker at concurrency 1 | About 2,300 jobs/s | About 6,000 jobs/s |
| Processing, concurrency 8–32 | About 11,000 jobs/s | About 18,000 jobs/s |
These are BullMQ-published rough illustrations from an Apple Silicon laptop with local PostgreSQL, trivial no-op jobs, and default durable settings, on the current documentation page accessed in 2026. They are not production guarantees or independent benchmark results. BullMQ cautions that results depend on hardware, PostgreSQL configuration, and where workers and the database are networked. Benchmark your own job mix and deployment topology before using throughput figures for capacity planning.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




