The fix is deployed, 143 messages go back from the dead-letter queue to their original exchange, the broker confirms every one of them, and the incident is closed. Ten minutes later the dead-letter queue holds 5 messages again. Are they five of the 143, or five new failures? Without a little preparation nobody can tell, because a publish confirm only says the broker took the message, not that the consumer processed it. Here is how to make a replay answer that question, measured on RabbitMQ 4.3.6.
A replay done right reads each message from the dead-letter queue, publishes it to its original route with publisher confirms, and acknowledges the original only after the confirm. That guarantees nothing is lost on the way. It says nothing about what happens next:
The management UI's "Move messages", a shovel, or a script that loops basic.get and basic.publish all stop at the confirm. To see the second case you have to recognise a message when it comes back.
Give every replayed message a header that names the replay, and remove the death headers so the message starts fresh:
import uuid, pika
DEATH = ("x-death", "x-first-death-", "x-last-death-", "x-delivery-count", "x-acquired-count")
replay_id = str(uuid.uuid4())
ch.confirm_delivery()
while (got := ch.basic_get("orders.dlq", auto_ack=False)) != (None, None, None):
method, props, body = got
death = props.headers["x-death"][0] # broker dead-lettering; Spring or MassTransit error queues keep the route elsewhere
headers = {k: v for k, v in (props.headers or {}).items() if not k.startswith(DEATH)}
headers["x-replay-id"] = replay_id
props.headers = headers
ch.basic_publish(death["exchange"], death["routing-keys"][0], body, props, mandatory=True)
ch.basic_ack(method.delivery_tag) # confirm_delivery made the publish wait for the broker
print("replay", replay_id)
Keep message_id and the body as they are: they are what lets you match a returned message to the one you sent, also when two replays of the same message overlap.
We published a message with an application header, rejected it into a dead-letter queue, replayed it with x-replay-id (once with the death headers stripped, once with them kept), and rejected it again. Classic and quorum queues behaved the same:
| After the second dead-lettering | Result |
|---|---|
your own headers (x-app, x-replay-id) | kept, unchanged |
x-death when the replay stripped it | a fresh entry, count 1 |
x-death when the replay kept it | also a fresh entry, count 1: the broker does not continue a count a publisher sent along |
x-delivery-count (quorum) | starts again at 1 |
Two things follow. Application headers ride along through dead-lettering, so the replay id is still there when the message comes back. And x-death cannot count replays: whatever the message went through before, after a republish RabbitMQ starts from scratch. Inside a broker-side retry loop (reject into a wait queue whose TTL dead-letters back to the work queue) the count does go up, 1, 2, 3, as before; it is only the republish that resets it. If you want to know how often a message has been replayed, count it in a header of your own, for example x-replay-count, increased on every replay.
A message that came back is therefore one that carries your replay id and a death: a new x-death from the broker, or, for consumers that republish failures themselves (Spring's RepublishMessageRecoverer, which copies the original headers), fresh exception headers. The id alone is not enough, because a replayed message still waiting in the work queue carries it too.
Returned messages arrive at the tail of the dead-letter queue. A peek through the management API takes the head, so look early, while the queue is short, or read deep enough:
rabbitmqadmin -f raw_json get queue=orders.dlq ackmode=ack_requeue_true count=500 \
| jq --arg id "$REPLAY_ID" '[.[] | select(.properties.headers["x-replay-id"] == $id
and .properties.headers["x-death"] != null)] | length'
ack_requeue_true puts the messages back. They are marked redelivered afterwards, and on a quorum queue of RabbitMQ 3.13 every such peek counts as a delivery against the queue's delivery limit; on 4.x a requeue no longer counts.
When to look: returns usually show up within minutes, because the consumer fails on the message as soon as it gets it. Look once after a few minutes and again after an hour or two. Failures that depend on time (a nightly job, a token that expires, a downstream that is only down at peak) need a second look the next day.
x-death restarts after a republish.