A job queue is the easiest buy decision on the list. BullMQ is MIT licensed, installs in a second, and arrives with retries, exponential backoff, priorities, concurrency limits and a rate limiter already written. Nobody should be writing those in 2026.
What it does not arrive with is a job that is safe to run twice. And it will run some of your jobs twice.
We run zeroutil’s video, PDF and image pipeline on BullMQ and Redis. The library does the queueing. The parts we wrote are the parts that decide what a second run means, what happens to a job nobody can process, and how long an id stays unique. That is the half you are still buying with your own time, so it is worth knowing its shape before you cost the work.
A retry re-runs your handler, not your job
BullMQ’s retry calls your processor again with the same job data. It does not undo anything the first attempt did. Every file written, row inserted, email sent or byte deleted on attempt one is still there on attempt two, and the library has no idea any of it happened.
Retries are off until you ask for them: you set attempts above 1, and with
exponential backoff the delay before each retry is 2 ^ (attempts - 1) * delay
milliseconds.[3] zeroutil runs three attempts with a 5 second base, so
a failing job retries 5 seconds after the first failure and 10 seconds after the
second. Fifteen seconds of waiting, three runs of the work itself, and the job is
either finished or dead.
Here is the part that cost us. Our processor deletes the uploaded input file once it has produced an output, because a video conversion service that keeps every upload is a disk bill waiting to happen. The first version did that cleanup in the error path too. So attempt one failed for a real reason, deleted the input, and attempts two and three then failed instantly on a missing file. Three attempts, one real error, two lies in the log.
logger.error({ jobId, op, attempt: job.attemptsMade + 1, err: msg }, 'job failed')
void unlink(outputPath).catch(() => {})
// Only unlink inputs on the last attempt; otherwise BullMQ retry needs the file.
if (job.attemptsMade + 1 >= (job.opts.attempts ?? 1)) {
cleanupInputs(inputPath, params)
}Note the + 1. Inside the processor, attemptsMade counts attempts that have
already finished, so the one currently running is attemptsMade + 1. In the
worker’s failed event the same field means something different, because by then
BullMQ has incremented it: the Lua that moves a job to failed runs HINCRBY on
the counter and only then compares it against the limit.[4] Our own
failed handler, which writes the dead-letter record, therefore asks
job.attemptsMade >= maxAttempts with no + 1 at all.
Two readings of one field, both correct, in one file. Nothing in the types tells you which one you are holding.
A stalled job is re-run, and nothing failed
If your worker stops renewing a job’s lock, BullMQ gives the job to a different worker. The docs are plain about it: a stalled job “is moved back to the waiting status and will be processed again by another worker”, or moved to the failed set once it has stalled too many times.[2]
The defaults decide how easily that happens, and they are tighter than they look. A worker takes a 30 second lock, renews it at half that interval, checks for stalled jobs every 30 seconds, and tolerates exactly one stall before failing the job.[1] Miss two renewals and your job is someone else’s.
Missing a renewal is not exotic. The renewal is a timer on the Node event loop, and BullMQ’s own advice is to make sure workers “return control to the NodeJS event loop often enough”.[2] zeroutil’s processor shells out to ffmpeg and pdf tooling, runs two jobs concurrently by default, and lives in a container that can be CPU-starved by the job next to it. A slow event loop, a paused container during a deploy, a Redis blip: any of them and the same job is running in two places.
This is where “the queue gives you retries” stops being a feature you bought and starts being a property you have to hold. At-least-once delivery is the guarantee, not a bug in it. So every side effect in the handler has to survive being done twice.
Ours does, because output paths are derived from the job id rather than from a timestamp or a counter:
const outputPath = imageMode
? buildImageOutputPath(jobId, ext)
: isPdfOp(op)
? buildPdfOutputPath(jobId, ext)
: buildOutputPath(jobId, ext)A second run writes the same bytes to the same path and the user gets one file. That is not luck and it is not a detail. It is the single design decision that turns a stall from an incident into a non-event, and it has to be made before the first line of the handler, not patched in after the first duplicate.
There is no dead-letter queue
BullMQ has a failed set, not a dead-letter queue. A job that exhausts its
attempts lands there and stays until something removes it.[3] For us
that something is our own retention setting: failed jobs carry
removeOnFail: { age: 3600 }, so anything nobody looked at within the hour is
gone.
An hour is short. It is short on purpose, because the failed set lives in the same memory as the working queue. The consequence is that the record of what broke outlives the breakage by sixty minutes, which is not long enough to notice a pattern.
So we wrote a dead-letter queue: a second BullMQ queue whose jobs are records
rather than work, written from the failed handler when attempts run out, with
retention set to never auto-remove. It is about forty lines including the
read-side.
The honest cost of those forty lines: it is a second queue on the same Redis, so it competes for the same memory budget it exists to outlive. Nothing alerts on it. Reading it means a human calling an admin endpoint, and nobody does that on a Tuesday. It is a place for failures to be found once you already suspect something, which is better than the hour we had and worse than it sounds.
Deduplication has an expiry date
A custom job id makes a job unique, and BullMQ will ignore a second add with an
id it already holds. But that protection ends when the job is removed: jobs
dropped by removeOnComplete “will not be considered as duplicates”, so the same
id can be added again.[5]
Which means the real dedupe window is whatever your retention happens to be. In zeroutil it is this:
defaultJobOptions: {
removeOnComplete: { age: config.fileExpiryMin * 60 + 300 },
removeOnFail: { age: 3600 },
attempts: config.queue.maxAttempts,
backoff: { type: 'exponential', delay: config.queue.backoffMs },
}fileExpiryMin is 15, so completed jobs are kept for 20 minutes and a job id is
unique for 20 minutes. Nobody chose 20 minutes as a dedupe window. It is the
lifetime of the output file plus five minutes of slack, picked so that the job
record outlives the file it points at. The dedupe behaviour is a side effect of a
retention decision made for a different reason.
That is fine here, because a user re-submitting the same file 21 minutes later genuinely should get a fresh job. It would not be fine in a payment pipeline, and the trap is that it looks identical in both: one line of retention config, no mention of deduplication anywhere near it.
The queue is exactly as durable as your Redis, and we had that wrong
Your jobs are Redis keys. Whatever your Redis is configured to do to keys, it
does to your jobs. BullMQ’s production guide says it in one sentence: noeviction
is “the only setting that guarantees the correct behavior of the queues”.[6]
We were not running noeviction. Until the day this article was written, both of
our BullMQ services ran Redis like a cache:
redis:
image: redis:7-alpine
command: redis-server --appendonly yes --maxmemory 256mb --maxmemory-policy allkeys-lruallkeys-lru evicts the least recently used keys once the limit is reached,
whether or not they carry an expiry.[8] A queued job is a Redis hash
and an entry in a list. Under memory pressure at 256mb, Redis would have deleted
them like any other cold key. The job would not fail, would not stall, would not
reach the dead-letter queue we built precisely so that nothing would be lost. It
would simply stop existing, and every error path we wrote would have had nothing
to report.
Redis’s own eviction documentation explains why this is a footgun rather than a mistake in the docs: it frames eviction as safe on the reasoning that cache entries are copies of data stored somewhere durable.[8] For a cache that is true. For the only copy of a queued job it is not, and nothing in the config says which one you are holding.
Two of the three settings on that page we had right. AOF was on, and
maxRetriesPerRequest: null was set on every connection handed to a Queue or
Worker, which BullMQ requires or the worker breaks in ways that look like
something else.[6] The third was wrong and nothing ever told us.
Worth knowing about the one we had right, too: AOF with the default
appendfsync everysec means you can still lose about a second of writes in a
disaster, and RDB snapshots alone mean being prepared to lose the latest minutes
of data.[7] A second of accepted-but-lost jobs is a real number, not
zero. Whether it matters is a product question. Ours convert a file the user
still has open in a browser tab, so it does not. If a job is money leaving an
account, it does.
What this actually costs
Buy the queue. Writing a reliable priority queue with backoff, rate limiting and distributed locks is not where a product gets built, and BullMQ has more operational hours behind it than anything you would write this quarter.
Then budget the other half, because it is not small and it does not appear in any comparison table:
- Handlers whose side effects survive running twice, designed in rather than patched on.
- A dead letter you can actually read, plus something that tells you to read it.
- A deduplication window you picked on purpose, sitting somewhere other than a retention setting.
- A Redis configured as a queue rather than as a cache, which is one word and no effort once someone has thought to check it.
On zeroutil that half is countable: 42 lines of dead-letter queue, one naming
rule for output paths, one + 1 in a cleanup guard, and one word of Redis
config. None of it is hard. All of it was added a piece at a time, after the
failure mode it guards against had already happened, and the last item on the
list was wrong in production until the week this was written.
That is the pattern worth taking away. A queue you bought does not fail loudly. It fails by going quiet, and the code that would notice is code nobody sold you.
Sources
- bullmq/src/classes/worker.ts
Supports: A BullMQ Worker's defaults are concurrency 1, lockDuration 30000 ms, stalledInterval 30000 ms and maxStalledCount 1, and lockRenewTime falls back to lockDuration / 2.
- Stalled Jobs
Supports: A worker locks a job while processing it and must periodically notify BullMQ that it is still working; a stalled job is moved back to waiting and processed again by another worker, or moved to the failed set once it has reached its maximum number of stalls.
- Retrying failing jobs
Supports: Automatic retries need the attempts option set above 1, exponential backoff retries after 2 ^ (attempts - 1) * delay milliseconds, and a processor that throws moves the job to the failed set.
- bullmq/src/commands/moveToFinished-14.lua
Supports: On the failure path BullMQ increments the job's attempt counter with HINCRBY on the atm field and only then compares it against the attempts limit, so the count includes the attempt that has just failed.
- Job Ids
Supports: Adding a job with an existing id means that job is ignored and not added at all, but jobs already removed by settings such as removeOnComplete are not considered duplicates, so the same id can be added again.
- Going to production
Supports: BullMQ calls noeviction the only maxmemory-policy setting that guarantees correct queue behaviour, recommends enabling AOF persistence, and requires maxRetriesPerRequest set to null on ioredis connections used by Workers.
- Redis persistence
Supports: With appendfsync everysec you may lose one second of data in a disaster, while RDB snapshotting alone means being prepared to lose the latest minutes of data if Redis stops without a correct shutdown.
- Key eviction
Supports: The allkeys-lru policy evicts the least recently used keys regardless of whether they carry an expiry, and Redis frames eviction as safe on the assumption that cache entries are copies of persistently-stored data.