Skip to main content
/ops/health says how deep each queue is. That is enough to know something is stuck and not enough to do anything about it: 4 000 messages waiting on vendor1 reads identically whether it is one customer’s misfired batch or every account on the box.
A bare number is read as a limit rather than a worker name, so queue dump 100 widens everything. The default is 20 per queue, ceiling 500 — past that the question belongs to cdr_submit rather than to a live gateway.

The four kinds

retries is the field that answers the question

A queue full of zeroes is a volume problem — more arrived than the vendor can take.A queue full of non-zero retries is a vendor problem — the same messages keep coming back.Those call for opposite responses. Adding capacity to a retry storm makes it worse; suspending a vendor that is merely busy moves the backlog somewhere more expensive.
depth and returned are different questions. depth is the whole queue; returned is how much was listed. Comparing them is how you tell whether you are seeing the problem or a corner of it.

No message text, and no flag that adds one

cdr_submit excludes the body on principle — destination and source are kept because rating and disputes need them, the text is not because operating a gateway never does. A queue dump inherits that rule rather than becoming the one operator surface that hands back customer traffic in bulk. body_chars with encoding tells a one-part OTP from a stuck six-part blast, which is the diagnostic question actually being asked.

Nothing is consumed and nothing is reordered

Each queue is walked with its own iterator while the gateway keeps draining it. So:
  • A message listed may already be gone by the time you read it.
  • depth is sampled after the walk and can disagree with entries by a few.
  • ordered says whether the list is the order the gateway will send in — false for a priority queue and for the delay queue.
That is the honest trade. The alternative is stopping traffic in order to look at it.

A broker queue is reported but never walked

AMQP has no way to look at a queued message without taking it. basicGet plus a requeue would reorder the queue and reset delivery counts on paying customers’ messages to satisfy a diagnostic. So the depth is reported and the dump says so. A broker that is not answering reports the depth as unknown rather than as zero — which matters, because zero is a claim and unknown is not.

There is deliberately no queue flush

Discarding a queue destroys messages that were accepted, billed and promised to a customer.To take traffic off a stuck vendor without losing it, use suspend — it drains that vendor’s queue back to the router so routing can pick another.

Access

The read token, not the admin one. A dump takes nothing and moves nothing, so a monitoring system holding the watching secret can ask for it.

Worker control

Suspend, hold, and what each does to a queue.

Queues and retries

Where the retries in that column come from.