1. Ask the database, not the log
Three of the four likely causes are not routing at all, and the database says which in one query.2. Is a vendor bound at all?
A perfect catch-all pointing at a vendor that is not connected delivers nothing, and looks like a routing bug from every angle.3. Did the client ever get told?
No. This is worth internalising, because it shapes every conversation with a customer. The submit response is sent before routing runs. So a routing failure can never surface as an error on the customer’ssubmit_sm or their HTTP call — it can only appear later, as a delivery
receipt and a cdr_rejected row.
4. Is the log quiet because it is fixed, or because it already said so?
routing.rule.broken is logged once per rule, not once per message — deliberately, since
unguarded it would be a line per message per attempt on exactly the traffic already going wrong.
The consequence is a trap: a quiet log is not proof the problem is gone. The counter resets only
when a new routing table is published.
So after fixing a rule, watch for the line to reappear, not for it to stay absent. The reliable
form of the same question is the number:
What to capture before you change anything
If you are going to ask for help, or you expect to be explaining this later, take these four now — they are all cheap and three of them are destroyed by a restart:1
The rejection breakdown
The
cdr_rejected query above, with its counts.2
/ops/health, whole
last_error, last_submit_error, routing_failed_queue and routing_retry_queue in one
snapshot.3
The WARN and ERROR lines
Not the whole log —
routing.rule.broken, message.dropped, listener.bind.failed,
message.submit.response.4
One message serial that failed
Everything downstream is easier when you can trace one real message rather than a class of them.
Related
Nothing is delivered
When it really is routing.
Refusals
Every
cdr_rejected reason, and what fixes it.