Warren
Guide

How to replay dead-lettered messages in RabbitMQ

2026-09-28 · 9 min read · Viktor Baumann

A dead-letter queue is filling up, the bug that caused it is fixed, and now 300 messages have to go back to where they came from. This is what your options are, what each one gets wrong, and the details that decide whether the replay works or produces a second incident.

What you are actually dealing with

RabbitMQ dead-letters a message when a consumer rejects it with requeue=false, when its TTL expires, when the queue is over its length limit, or (quorum queues) when the delivery limit is reached. The broker republishes the message to the queue's dead-letter exchange, and on the way it adds an x-death header: which queue it came from, why, when, the original exchange and routing keys. (Reading x-death is its own article.)

That header is the key to a correct replay. It tells you where the message was supposed to go. It is also the thing that trips up most home-made replays, because the header stays on the message unless you remove it, and consumers with retry logic often read it.

Three questions before you move anything:

  1. Is the cause fixed? Replaying into a consumer that still fails just cycles the messages back into the DLQ, with a bigger x-death count each time.
  2. Where do they go back to? To the original exchange with the original routing key (so bindings decide again), or directly to the queue they died in? The first is correct if the topology has changed; the second is what most people mean.
  3. All of them, or some? A DLQ usually mixes several failure causes. Replaying everything replays the poison messages too.

Option 1: the management UI's "Move messages"

With the rabbitmq_shovel and rabbitmq_shovel_management plugins enabled, every queue page in the management UI has a Move messages panel: type a destination queue name, click, done. Under the hood it creates a temporary shovel that drains the queue into the destination via the default exchange.

rabbitmq-plugins enable rabbitmq_shovel rabbitmq_shovel_management

Good for: one queue, all messages, right now, you trust every message in there.

Not good for: anything selective. It moves everything, including the ten messages that are the actual poison. It does not strip x-death, so consumers that count retries may reject them on arrival. It publishes with no confirms visible to you and leaves no record of who moved what. And it only targets a queue by name, never the original exchange.

Option 2: a shovel you define yourself

A dynamic shovel gives you a bit more control than the button. src-delete-after: queue-length makes it stop after moving the messages that were in the queue when it started, so new dead letters arriving during the replay are not dragged along.

rabbitmqctl set_parameter shovel replay-orders \
  '{"src-protocol": "amqp091", "src-uri": "amqp://", "src-queue": "orders.dlq",
    "dest-protocol": "amqp091", "dest-uri": "amqp://",
    "dest-exchange": "orders", "dest-exchange-key": "order.created",
    "src-delete-after": "queue-length", "ack-mode": "on-confirm"}'

ack-mode: on-confirm is the important line: the shovel acknowledges a source message only after the destination confirmed it, so a crash mid-way duplicates at most one message and loses none.

Limits: one destination for the whole batch. If your DLQ collects messages from five different routing keys (common when one DLX serves a whole service), a shovel sends them all to the same place. And again: no selection, no header cleanup, no audit beyond the shovel's own log line.

Option 3: a script

This is what most teams end up with at 3 a.m., and it is fine if it gets four things right. In Python with pika:

import pika, json

conn = pika.BlockingConnection(pika.URLParameters("amqp://user:pass@rabbit/%2F"))
ch = conn.channel()
ch.confirm_delivery()                     # 1. publisher confirms

while True:
    method, props, body = ch.basic_get("orders.dlq", auto_ack=False)   # 2. never auto-ack
    if method is None:
        break
    headers = dict(props.headers or {})
    death = (headers.get("x-death") or [{}])[0]
    exchange = death.get("exchange", "")
    routing_key = (death.get("routing-keys") or [method.routing_key])[0]

    if not looks_replayable(body, headers):            # 3. select (your own check)
        ch.basic_nack(method.delivery_tag, requeue=True)
        continue

    for k in list(headers):                             # 4. strip death bookkeeping
        if k.startswith("x-death") or k.startswith("x-first-death") or k.startswith("x-last-death"):
            del headers[k]
    props.headers = headers

    ch.basic_publish(exchange, routing_key, body, props, mandatory=True)
    ch.basic_ack(method.delivery_tag)                   # only after the confirm above returned

conn.close()

The four things, in order of how often they go wrong:

  1. Confirms before acks. confirm_delivery() turns basic_publish into a call that raises when the broker did not accept the message. Ack the source message only after that. Ack first and a broker hiccup deletes messages you never republished.
  2. No auto-ack. With auto_ack=True the message is gone the moment you fetched it. If the script dies between fetch and publish, so is the message.
  3. Select, do not drain. Messages you skip are released with requeue=True. Because unacknowledged messages stay outstanding on the channel, the next basic_get returns the next message rather than the same one again. Closing the channel releases everything you skipped.
  4. Strip the death headers. Otherwise the consumer sees a message with x-death[0].count = 3 and, if it counts retries, gives up immediately. Also look for application-level retry headers such as x-retry-count or Spring's x-exception-*; a replay is supposed to be a fresh start.

mandatory=True makes the broker return the message (basic.return) instead of silently dropping it when the routing key no longer matches any binding, which happens more often than you would think after a deploy renamed a queue. Register a return callback or the confirm will look like a success.

Flooding. Three hundred messages arriving in the same second at a consumer that was just restarted is how you get a second outage. Sleep between publishes, or replay in batches of fifty and watch the consumer's error rate in between.

Same idea in Java

With the RabbitMQ Java client: channel.confirmSelect(), channel.basicGet(queue, false), publish with channel.basicPublish(exchange, key, true, props, body), then channel.waitForConfirmsOrDie(), then channel.basicAck(tag, false). Spring AMQP users: RabbitTemplate with publisher-confirm-type: correlated and a ConfirmCallback, and take the messages with rabbitTemplate.receive(queue) inside a transaction or with manual acks via execute(ChannelCallback).

Option 4: do not replay by hand at all

Some failures are transient by nature: a downstream API was down for two minutes. For those, a retry topology handles the replay before anyone is paged: the consumer rejects into a wait queue with a per-queue TTL and no consumers, whose dead-letter exchange points back at the work queue. After N cycles (read x-death[0].count) the consumer routes the message to a final parking lot DLQ that a human looks at.

This does not remove the need for a manual replay; it removes the manual replays that were never necessary. The parking lot still fills up with the messages that failed for a real reason, and those you replay after the fix.

Checklist

Move buttonShovelScript
Select which messagesnonoyes
Original exchange and key per messagenono (one target)yes
Strip x-death and retry headersnonoif you write it
Loss-free (confirm before ack)yeswith on-confirmif you write it
Throttlenonoif you write it
Who moved what, whennonoif you log it
Edit a message before replaynonoif you write it

Whatever you use: replay a handful first, watch the consumer, then the rest. And write down what you did. The postmortem will ask.