My dashboard gets 404 from /ops/health
My dashboard gets 404 from /ops/health
smsg.ops.token disables the endpoint and answers 404. A wrong token gets the same
404 rather than a 401, deliberately — so a caller cannot tell whether there is something here to
find a token for.Prometheus stopped collecting after an upgrade
Prometheus stopped collecting after an upgrade
/q/metrics from 8080 to 9000. A job still pointed at 8080
gets 404s rather than failing loudly, so the dashboard reads as a quiet system rather than a broken
monitor.Change the scrape target at the same time as the upgrade.I changed a setting in the panel and nothing happened
I changed a setting in the panel and nothing happened
config_sources on /ops/health. It reports, per domain, whether properties, credentials,
routing and rates come from database or file.A panel change does nothing if the gateway is reading a file for that domain. The symptom sends
people looking at the edit rather than at the source.If the source is right, the poll is up to thirty seconds behind — ./scripts/fireflo reload applies
it now.I corrected an account's balance and it kept sending
I corrected an account's balance and it kept sending
./scripts/fireflo credit release <account> returns the unspent remainder so the next message
reserves against the corrected figure. It takes nothing away.A worker vanished from /ops/health
A worker vanished from /ops/health
disable removes a worker from health entirely, which looks like a crash if you did
not do it yourself.suspend is the one that keeps the worker visible while taking traffic off it.Counting the sessions array gives the wrong number
Counting the sessions array gives the wrong number
/ops/health, whatever sessions_total says — health
is polled constantly and a listener with five hundred binds was sending five hundred rows each time.Use bound, or sessions_total. To see the whole set, page the sessions endpoint.tps_measured is missing rather than zero
tps_measured is missing rather than zero
Is a growing queue always a problem?
Is a growing queue always a problem?
retries means the
same messages keep coming back, which is a vendor problem and gets worse if you add capacity.Should I raise maxAttempts when messages are dropping?
Should I raise maxAttempts when messages are dropping?
maxAttempts 20 → 60 with a 500 ms
backoff changed loss from 217 to 227 messages and made throughput slightly worse.Pace to the vendor instead. Setting that vendor’s tps to what it can actually take took the same
run to zero lost, zero ESME_RMSGQFUL, and exactly 1.00 submits per message.Can I clear a stuck queue?
Can I clear a stuck queue?
queue flush. Discarding a queue destroys messages that were accepted,
billed and promised to a customer.Use suspend — it drains that vendor’s queue back to the router so routing can pick another vendor.Delivery receipts are being sent to a carrier as new messages
Delivery receipts are being sent to a carrier as new messages
type == 18 routing rule. Receipts fall to the catch-all, reach a vendor worker, and are
submitted upstream as brand new outbound traffic — which you pay for.conf/routingTable.conf has always carried this rule; the database did not, and in database mode
the database wins.How much logging is safe to leave on?
How much logging is safe to leave on?
message.trace.mode = all produced 226 MB of log for one 20 000-message run. Excellent for one
investigation, ruinous as a default.The same applies to log.pdus — and it contains bind passwords and message bodies. Use a session
capture instead: bounded, in memory, and expiring in fifteen minutes.