The baseline
The corrected burst is the one to compare future work against: 752 msg/s delivered, 98.9% success,
26.3 s for 20 000 messages.
The router was never the bottleneck
fireflo_router_queue read 0 for the whole 20 000-message burst and peaked at 18 across every
other run.
Messages were being routed as fast as they arrived; they then piled up in the vendor worker’s
queue, which reached 7 987. Everything limiting throughput was downstream of the routing code.
The one line that cut loss by 92%
The first burst lost 2 528 messages, every onemessage.dropped reason=maxAttempts attempts=20 limit=20, with 151 985 submit attempts for 24 000 accepted messages — about six each.
The vendor connection never dropped. This was not connectivity. The test SMSC stopped answering
fast enough under burst, submits expired, the worker retried, and messages that used all 20 routing
attempts were abandoned.
Re-running at transceivers = 1 — same messages, same window, same router threads:
The queue still reached 8 042 and the router still never backed up, so the shape of the run was
identical. What changed is that a single session does not provoke whatever the test SMSC does when
over-sessioned.
The lever that looked obvious and did nothing
An earlier version of the baseline said the fix for the residual 1.1% was to raiseoutSms.routing.maxAttempts and outSms.routing.retryBackoffMillis, reasoning that a message
survives only about two seconds of a vendor not accepting.
That was measured and it is wrong.
No improvement, and marginally slower. More retries against a vendor that is not accepting is more
load on the thing already refusing.
What did fix it was
vendor1.tps = 300 — pacing the gateway to what the vendor could actually
take. Zero lost, zero ESME_RMSGQFUL, and exactly 1.00 submits per message.
What to measure on your own deployment
1
Router queue against vendor queue
Flat router, growing vendor queue = the vendor is the limit. Both growing = ingress is outrunning
routing.
2
Submits per message
1.00 is healthy. Anything materially above it is retries, and retries are wasted vendor capacity.
3
ESME_RMSGQFUL count
Non-zero means you are pushing harder than the far end accepts.
4
Loss at the attempt ceiling
message.dropped reason=maxAttempts is the end state of all of the above.Caveats
- A test SMSC is not a carrier. Both burst runs were bounded by the simulator, not by FireFlo.
- Run-to-run spread was larger than the effect of the window setting. Raising
windowSize100 → 500 and transceivers 1 → 4 produced 337 and 910 msg/s on two runs of identical configuration, against 500 and 840 for the baseline. This rig cannot resolve that effect — which is a statement about the measurement, not evidence that the window does not matter. message.trace.mode = allproduced 226 MB of log for one 20 000-message run. Useful for one investigation, ruinous as a default.
Related
Queues
Telling a volume problem from a vendor problem.
Generating traffic
Producing load to measure against.