Stripe Billing will retry a failed subscription payment on a schedule its own model picks, email the buyer after every failed attempt, and flip the subscription to a status you choose once the retries run out. That is dunning, it is a set of Dashboard checkboxes, and turning it on is the right call. What it does not decide is what your product does during the days between the first failed charge and the cutoff, or what closing the account means when the account is a private Telegram channel.
We know the shape of that gap because we built volta’s subscription layer on Stripe Billing and had to write the half Stripe leaves to you. volta sells access to creators’ private channels: a creator connects a bot, a buyer subscribes, and every renewal is a Stripe invoice. When one of those invoices fails, Stripe’s retries are one part of the response. The other part is a state machine we own and have to cancel.
What Stripe Billing’s dunning actually gives you
Three tools and a status flag. Smart Retries reattempts the charge at times an AI model picks from signals like how many devices have recently used the card and which country issued it, for a number of attempts you cap inside a window of one week to two months, with a recommended default of eight tries over two weeks.[1] Failed-payment emails go out after each attempt, every one carrying a link to a Stripe-hosted portal or a page of yours where the buyer can replace the card.[4] Automatic card updates pull a fresh number from the network when the issuer reissues, so a share of failures never happen.[3] None of it needs code.[3]
The status flag is the handoff. While retries run, the subscription sits in
past_due.[2] When they are exhausted, Stripe moves it to whatever
you set: canceled, unpaid, or left at past_due.[1] Stripe’s own
instruction for the unpaid case is one sentence: revoke access to your
product, because the payments were already tried.[2] That sentence is
the seam. Stripe tells you the moment to act. Acting is yours to build.
The grace window is a product decision, and now it is yours to run
volta’s window is 96 hours. That number is not Stripe’s retry schedule and it is in no Stripe setting. It is a churn-versus-goodwill call we made: a card that fails on renewal day is usually a rotation or a temporary hold, not a buyer leaving, so they keep channel access for four days while we retry. Too short and you ban paying customers over a bank hold. Too long and you are giving the product away.
Because the window is ours, the retry cadence is ours too. On
invoice.payment_failed we move the subscriber into a grace state, stamp
graceUntil 96 hours out, and schedule four delayed jobs in Redis: a retry at
24, 48 and 72 hours that each call stripe.invoices.pay, and a ban at 96. The
first retry also DMs the buyer a link to update their card, because volta’s
checkout is a Telegram Mini App and the Stripe email points at a page volta does
not use.
// on invoice.payment_failed
await db.update(subscribers)
.set({ status: 'grace', graceUntil: new Date(Date.now() + 96 * 3600_000) })
.where(eq(subscribers.id, sub.id));
const pfx = `${sub.id}:grace:${invoice.id}`;
for (const [attempt, hours] of [[1, 24], [2, 48], [3, 72]] as const) {
await renewalRetryQueue.add('retry',
{ subscriberId: sub.id, invoiceId: invoice.id, attempt },
{ delay: hours * 3600_000, jobId: `${pfx}:retry:${attempt}`, attempts: 1 });
}
await subscriberBanQueue.add('ban',
{ subscriberId: sub.id },
{ delay: 96 * 3600_000, jobId: `${pfx}:ban`, attempts: 3,
backoff: { type: 'exponential', delay: 30_000 } });Driving our own invoices.pay cadence means Stripe’s Smart Retries has to be
switched off in the Dashboard. Leave both on and the same invoice is retried by
two schedules that do not know about each other, and the buyer can be charged in
the same minute we remove them.
Recovery has to unwind four scheduled jobs, and the cancel can lose a race
When a retry clears, Stripe fires invoice.paid and returns the subscription to
active.[2] Our handler moves the subscriber back to active and
then has to find the jobs it scheduled and delete them, because a delayed job in
Redis does not know the invoice is now paid. Each job was keyed
<subscriber>:grace:<invoice>:..., so recovery looks each one up by id and
removes it.
Two details make this more than a loop. First, job.remove() fails if the job
has already begun executing, so the removal is wrapped and allowed to throw: the
72-hour retry can fire in the same second the card finally clears. Second, and
the reason the first is survivable, every worker re-reads state before it does
anything. The retry worker exits unless the subscriber is still grace. The ban
worker exits if the subscriber is already banned or churned. Cancelling the
jobs is the fast path; the status check inside each job is what makes a missed
cancel harmless. Neither is enough alone.
// on the recovered charge
const pfx = `${sub.id}:grace:${invoiceId}`;
for (const suffix of [':retry:1', ':retry:2', ':retry:3']) {
const job = await renewalRetryQueue.getJob(`${pfx}${suffix}`);
await job?.remove().catch(() => {}); // may have already started
}
const banJob = await subscriberBanQueue.getJob(`${pfx}:ban`);
await banJob?.remove().catch(() => {});
// inside the ban worker, before it bans anyone
if (sub.status === 'banned' || sub.status === 'churned') return;There is a redelivery case this handles for free. A repeated
invoice.payment_failed for the same invoice re-adds jobs with ids that already
exist, and BullMQ drops them as duplicates.[5] That covers Stripe
sending the webhook twice. It does not cover a different invoice failing while
the subscriber is still in grace from the first: those jobs carry a different
invoice id and you get two overlapping cycles. volta has not hit it, because an
unpaid subscription pauses collection and later invoices stay in draft rather
than being charged,[2] so a second failure does not arrive. The code
does not guard it, and that assumption is worth writing down before it changes.
What it costs to own this
You are running a scheduler now, and its memory is Redis. Four delayed jobs per failed invoice sit in a queue for up to four days. If that data is lost, or the queue is renamed or cleared during a deploy, the 96-hour ban never fires and the buyer keeps access for nothing. Stripe’s dunning has no matching failure mode, because that schedule lives on Stripe’s side. We wrote a guard that blocks the destructive queue operations in our admin tooling, a piece of code that exists only because the scheduler became load-bearing.
The 96-hour window is itself a running cost: every subscriber in grace is using
the product with no completed payment behind them. And each stripe.invoices.pay
call is an attempt we chose to make instead of Smart Retries, so we also gave up
the model that picks better retry times than a fixed 24 / 48 / 72.[1]
That was a deliberate trade for one clock we control. It is still a trade.
| Stripe Billing handles | Stays yours | How you find out it is not done |
|---|---|---|
| Retrying the card at model-picked times, for a capped number of attempts | The grace window: how long access survives a failed renewal, and why that number | You banned a paying customer over a two-day hold, or gave the channel away for a month |
| Emailing the buyer a link to fix the card | Where that link goes when checkout is not a web page | The buyer gets a Stripe email pointing at a portal your product does not use |
Flipping the subscription to unpaid when retries run out | Removing the buyer from the channel at that moment | The subscription reads unpaid in Stripe and the buyer is still posting |
| Not retrying one invoice twice | Cancelling your own scheduled ban when a retry finally succeeds | A recovered subscriber is kicked at hour 96 anyway |
None of this argues for hand-rolling billing. Stripe’s retry model recovers more
than a schedule you would write, the emails are fine, and the card updater is
free revenue. The point is narrower: invoice.payment_failed arrives, unpaid
gets set, Stripe says revoke access, and every line after that is yours to
build, schedule, and cancel. Budget it as a state machine, not a webhook
handler.
Sources
- Automate payment retries
Supports: Smart Retries uses an AI model fed signals such as how many devices have recently presented a payment method and which country issued the card to pick retry times; you cap the number of attempts inside a window of 1 week, 2 weeks, 3 weeks, 1 month or 2 months, with a recommended default of 8 tries within 2 weeks; a custom schedule allows up to three retries; after recovery fails the subscription transitions to canceled, unpaid, or stays past_due according to your Dashboard setting; the invoice.payment_failed webhook carries attempt_count and next_payment_attempt; and hard decline codes are not retried.
- How subscriptions work
Supports: A subscription whose latest finalized invoice failed sits in past_due while retries run; once retries are exhausted it moves to canceled, unpaid, or stays past_due per your failed-payment settings; Stripe's instruction for the unpaid case is to revoke access to your product because payments were already attempted and retried while past_due; a successful payment fires invoice.paid and returns the subscription to active; and an unpaid subscription pauses collection with subsequent billing periods' invoices held in draft rather than charged.
- Revenue recovery
Supports: Stripe Billing's revenue recovery toolset is Smart Retries, automatic failed-payment and card-expiry customer emails, automatic card updates when an issuer reissues a number, and recovery analytics, and none of the features require you to write code.
- Automate customer emails
Supports: When failed-payment emails are enabled Stripe sends the customer an email after each failed payment, and the email carries a link to either a Stripe-hosted Customer Portal or your own subscription-management page where they can update the card; that hosted link is invalidated once the subscription status changes to unpaid, cancelled or incomplete_expired.
- Job Ids
Supports: Setting an identical custom jobId makes BullMQ treat the job as a duplicate and not add it to the queue, and a job removed on completion no longer counts as existing for that check.