fireflo_, so one match finds all of it:
The metrics
All gauges, read at scrape time, so none of them costs anything per message.
The four worth an alert
fireflo_routing_failed_queue
Non-zero means messages are gone — no CDR and no receipt — unless
outSms.enqueueFailedRouting is set.fireflo_cdr_failed
Revenue you cannot invoice. Silent everywhere else.
fireflo_rating_unrated
Traffic delivered and charged nothing. Grows quietly and never self-corrects.
amqp.bridge.degraded
In the log rather than a metric. REST durability is gone and nothing else will tell you.
routing_rules_broken is how you tell fixed from already-reported
The log says each broken rule once per rule, not once per message — deliberately, since unguarded
it would be a line per message per attempt on exactly the traffic already going wrong.
The counter resets only when a new routing table is published. So a quiet log is not proof the
problem is gone, and this gauge is the reliable form of the same question.
whitelist_would_reject before you enforce
Report-only mode counts what would have been refused without refusing it.
Per-worker queue depth is not here
It would need a meter per worker, registered and retired as workers start and stop./ops/health
reports it per worker in the meantime, so a per-vendor backlog alert has to come from there rather
than from Prometheus.
Scrape configuration
Related
Live health
The state that is not in Prometheus.
Tokens and access
Which endpoints need which token.