What is already cluster-safe
More than you would expect. The send path was built without a cluster in mind and is nonetheless correct across instances.What breaks
Money and records
Message loss
All receipt correlation lives in a node-local store.- A receipt arriving at an instance that did not submit the message finds nothing, is dropped with a
single log line, and produces no customer receipt, no webhook and no
cdr_finalrow. With N instances that is roughly (N−1)/N of all receipts. - The store takes an exclusive file lock. The second instance to start on the same path fails, the failure is caught, logged at WARN, and it silently continues with no persistence at all — so “two instances sharing a data directory” looks identical to “healthy”.
- Unpushed receipts are parked locally, so a customer reconnecting to a different instance never receives them.
- Bound sessions are tracked per JVM, so an instance cannot tell a customer is connected elsewhere.
Concatenated messages, concretely
Segments arriving on different instances never reassemble. For a three-segment message split two-and-one across two instances:- each instance answers
ESME_ROKand charges for the segments it received; - each waits out
reassembling.timeoutMillis, 30 seconds by default; - each then transmits its segments as orphaned fragments carrying a concatenation header for a message that will never complete.
message.parts.incomplete warning per instance, neither
of which knows the other exists.
A customer’s SMPP session is normally sticky at the TCP level, so this bites on reconnect mid-message,
on load-balancer rebalance, and on clients holding several binds that spread segments across them.Content whitelisting still applies to what each instance sends — the timeout path is content-checked —
so the failure is delivery and billing, not enforcement.
Degraded but not lost
Vendor TPS, ingress rate limits, throughput sampling and routing-group cursors are all per process. N instances send up to N × the contracted rate to a carrier. That is a guard against a runaway sender, not a distributed quota — see Limits.What it would take
Broadly: move receipt correlation and reassembly into shared storage, givecdr_final a unique key and
an upsert, and make the sweeper leader-elected or idempotent. Each is tractable; none is done.
Until then, scale vertically, and use several independent instances where the traffic can be
partitioned by customer.