A queue with messages and zero consumers is the quietest outage there is. Nothing errors. Producers publish happily. The dashboard shows the broker green. Then the TTL runs out, or the queue hits its limit, and an hour of orders is in a dead-letter queue. Here is how to get paged before that.
"No consumers" alone is not the alert. Plenty of queues have no consumers by design: retry-wait queues, parking-lot DLQs, queues of a batch job that runs at night. The signal is the combination:
consumers == 0messages_ready > 0 (something is waiting)Two neighbours of this alert are worth setting up at the same time, because they catch the cases "consumers exist but do nothing":
consumers > 0, messages_ready > 0, and the ack rate is zero for five minutes. The consumer process is alive, holds the connection, and is deadlocked or stuck on a downstream call.messages_ready grew by more than N in the last M minutes. Consumers are there but cannot keep up.The management plugin exposes everything you need per queue. Ask for only the columns you want, otherwise the response for a broker with a thousand queues is huge:
curl -s -u monitor:secret \
'http://rabbit:15672/api/queues?columns=vhost,name,consumers,messages_ready,messages_unacknowledged,message_stats.ack_details.rate'
Relevant fields:
| Field | Meaning |
|---|---|
consumers | Number of consumers currently attached to the queue. |
messages_ready | Messages waiting to be delivered. This is the "backlog". |
messages_unacknowledged | Delivered but not yet acked. High and flat means consumers hold messages without finishing them. |
message_stats.ack_details.rate | Acks per second, averaged over the last sampling interval. Zero with a backlog and consumers present is the stuck-consumer signal. |
consumer_utilisation | Fraction of time consumers could take new messages. Low with a backlog means consumers are the bottleneck. |
idle_since | Set when the queue had no activity at all since that time. Useful to ignore truly dead queues. |
The monitoring user needs the monitoring tag, nothing more. Do not use an administrator account for a poller.
Since RabbitMQ 3.8, rabbitmq_prometheus ships with the broker. By default /metrics is aggregated: you get totals per node, not per queue, which is useless for this alert. Per-queue values come from either:
/metrics/per-object, every metric for every object. Fine for a few hundred queues, heavy for thousands./metrics/detailed?family=queue_coarse_metrics&family=queue_consumer_count (3.9+), only the families you ask for. Metrics from this endpoint are prefixed rabbitmq_detailed_.prometheus.return_per_object_metrics = true in rabbitmq.conf to make /metrics per-object.rabbitmq-plugins enable rabbitmq_prometheus
curl -s 'http://rabbit:15692/metrics/detailed?family=queue_coarse_metrics&family=queue_consumer_count' | grep orders
With the detailed endpoint scraped as its own job, three rules cover the three signals above. Adjust the prefix (drop detailed_) if you use per-object metrics.
groups:
- name: rabbitmq-queues
rules:
- alert: RabbitMQQueueNoConsumers
expr: |
rabbitmq_detailed_queue_consumers == 0
and on (vhost, queue) rabbitmq_detailed_queue_messages_ready{queue!~".*(\\.wait|\\.dlq|\\.parking)$"} > 0
for: 5m
labels: {severity: page}
annotations:
summary: "{{ $labels.vhost }}/{{ $labels.queue }} has {{ $value }} messages and no consumer"
- alert: RabbitMQQueueConsumersStuck
expr: |
rabbitmq_detailed_queue_consumers > 0
and on (vhost, queue) rabbitmq_detailed_queue_messages_ready > 0
and on (vhost, queue) rate(rabbitmq_detailed_queue_messages_acked_total[5m]) == 0
for: 5m
labels: {severity: page}
- alert: RabbitMQQueueGrowing
expr: |
delta(rabbitmq_detailed_queue_messages_ready[15m]) > 1000
for: 0m
labels: {severity: warn}
The regex on queue excludes the queues that have no consumers on purpose; more on that below. Because the exclusion sits on the messages_ready side of the and, those queues never produce a result, whatever their consumer count.
Route severity: page to your on-call and warn to a channel. A queue growing at 2 a.m. may be a batch job; a queue with no consumer at 2 a.m. is not.
A twenty-line script on a box that can reach the management API does the job for a small setup. This one posts to a Slack incoming webhook when a queue has a backlog and no consumer, and keeps a marker file so it does not repeat itself every minute:
#!/usr/bin/env bash
set -euo pipefail
API=http://rabbit:15672; AUTH=monitor:secret
HOOK=https://hooks.slack.com/services/T000/B000/xxxx
IGNORE='\.(wait|dlq|parking)$'
STATE=/var/tmp/rabbit-noconsumer; mkdir -p "$STATE"
# "vhost/queue<TAB>ready" for every queue that is in trouble right now
ALERTING=$(curl -sf -u "$AUTH" "$API/api/queues?columns=vhost,name,consumers,messages_ready" | jq -r --arg ign "$IGNORE" '.[] | select(.consumers == 0 and .messages_ready > 0 and (.name | test($ign) | not))
| "\(.vhost)/\(.name) \(.messages_ready)"')
# notify once per incident: a marker file per queue
while IFS=$' ' read -r q ready; do
[ -n "$q" ] || continue
key="$STATE/$(printf '%s' "$q" | md5sum | cut -c1-32)"
[ -e "$key" ] && continue
printf '%s
' "$q" > "$key"
curl -sf -X POST -H 'content-type: application/json' "$HOOK" -d "$(jq -nc --arg t ":rotating_light: $q has $ready messages and no consumer" '{text: $t}')"
done <<< "$ALERTING"
# recovered: drop markers for queues that are no longer alerting
for f in "$STATE"/*; do
[ -e "$f" ] || continue
grep -qF "$(cat "$f")" <<< "$ALERTING" || rm -f "$f"
done
Run it every minute from cron. The "for 5 minutes" part is missing here; add it by requiring the marker to be older than five minutes before posting, or accept the occasional deploy-time notification. Teams users swap the Slack payload for an Adaptive Card or a plain {"text": ...} to an incoming webhook.
Do this once and write it down, otherwise the alert gets muted the first week:
Name your queues so a regex can tell these apart. orders.process, orders.retry.wait, orders.dlq is a convention that makes every rule on this page a one-liner.
The no-consumer alert is the early warning. The late warning is messages_ready > 0 on a queue matching \.dlq$, because by then messages have already failed. Give that one a lower threshold than you think: one dead letter in a queue that is normally empty is news, three hundred is a postmortem. If you want fewer false alarms, alert on the rate of arrivals into the DLQ (message_stats.publish_details.rate on the DLQ, or rate(rabbitmq_detailed_queue_messages_published_total[5m])) rather than on the count.
Whatever alerts you build, keep the per-queue time series for at least a week. The question after a no-consumer page is always "since when", and the management UI only shows the last hour at any useful resolution. Prometheus gives you this for free. The cron script does not; if you go that route, at least append the numbers to a file per run.