Skip to main content
These are real numbers from a real run, not projections. Read the caveats before quoting any of them — the rig used a test SMSC, and the most useful finding is about that.

The baseline

The corrected burst is the one to compare future work against: 752 msg/s delivered, 98.9% success, 26.3 s for 20 000 messages.

The router was never the bottleneck

fireflo_router_queue read 0 for the whole 20 000-message burst and peaked at 18 across every other run. Messages were being routed as fast as they arrived; they then piled up in the vendor worker’s queue, which reached 7 987. Everything limiting throughput was downstream of the routing code.
This is the single most useful diagnostic shape to recognise. Router queue flat while a vendor queue grows means the gateway is fine and the vendor is the constraint — adding router threads or tuning routing changes nothing.

The one line that cut loss by 92%

The first burst lost 2 528 messages, every one message.dropped reason=maxAttempts attempts=20 limit=20, with 151 985 submit attempts for 24 000 accepted messages — about six each. The vendor connection never dropped. This was not connectivity. The test SMSC stopped answering fast enough under burst, submits expired, the worker retried, and messages that used all 20 routing attempts were abandoned. Re-running at transceivers = 1 — same messages, same window, same router threads: The queue still reached 8 042 and the router still never backed up, so the shape of the run was identical. What changed is that a single session does not provoke whatever the test SMSC does when over-sessioned.
The first burst measured the test SMSC’s limits, not the gateway’s. It is kept in the record because the contrast is the most useful thing the exercise produced — and because “we lost 12% of a burst” is exactly the kind of number that gets quoted without its cause.

The lever that looked obvious and did nothing

An earlier version of the baseline said the fix for the residual 1.1% was to raise outSms.routing.maxAttempts and outSms.routing.retryBackoffMillis, reasoning that a message survives only about two seconds of a vendor not accepting. That was measured and it is wrong. No improvement, and marginally slower. More retries against a vendor that is not accepting is more load on the thing already refusing. What did fix it was vendor1.tps = 300 — pacing the gateway to what the vendor could actually take. Zero lost, zero ESME_RMSGQFUL, and exactly 1.00 submits per message.
The general lesson: when a vendor is the constraint, pace to it rather than pushing harder at it. A submits-per-message ratio above 1 is the number that tells you which situation you are in.

What to measure on your own deployment

1

Router queue against vendor queue

Flat router, growing vendor queue = the vendor is the limit. Both growing = ingress is outrunning routing.
2

Submits per message

1.00 is healthy. Anything materially above it is retries, and retries are wasted vendor capacity.
3

ESME_RMSGQFUL count

Non-zero means you are pushing harder than the far end accepts.
4

Loss at the attempt ceiling

message.dropped reason=maxAttempts is the end state of all of the above.

Caveats

  • A test SMSC is not a carrier. Both burst runs were bounded by the simulator, not by FireFlo.
  • Run-to-run spread was larger than the effect of the window setting. Raising windowSize 100 → 500 and transceivers 1 → 4 produced 337 and 910 msg/s on two runs of identical configuration, against 500 and 840 for the baseline. This rig cannot resolve that effect — which is a statement about the measurement, not evidence that the window does not matter.
  • message.trace.mode = all produced 226 MB of log for one 20 000-message run. Useful for one investigation, ruinous as a default.

Queues

Telling a volume problem from a vendor problem.

Generating traffic

Producing load to measure against.