Skip to main content
Status: a decision record, not a roadmap item. Nothing on this page is implemented. It exists so the next person to ask “can we run two of these?” gets a real answer instead of rediscovering the constraints from the source.
No, not today. Two instances against one database will drop most delivery receipts, and did until recently double-credit every prepaid top-up. What follows is why, and what it would take. For independent gateways on one host — which does work — see Several instances.

What is already cluster-safe

More than you would expect. The send path was built without a cluster in mind and is nonetheless correct across instances.

What breaks

Money and records

Message loss

All receipt correlation lives in a node-local store.
  • A receipt arriving at an instance that did not submit the message finds nothing, is dropped with a single log line, and produces no customer receipt, no webhook and no cdr_final row. With N instances that is roughly (N−1)/N of all receipts.
  • The store takes an exclusive file lock. The second instance to start on the same path fails, the failure is caught, logged at WARN, and it silently continues with no persistence at all — so “two instances sharing a data directory” looks identical to “healthy”.
  • Unpushed receipts are parked locally, so a customer reconnecting to a different instance never receives them.
  • Bound sessions are tracked per JVM, so an instance cannot tell a customer is connected elsewhere.

Concatenated messages, concretely

Segments arriving on different instances never reassemble. For a three-segment message split two-and-one across two instances:
  • each instance answers ESME_ROK and charges for the segments it received;
  • each waits out reassembling.timeoutMillis, 30 seconds by default;
  • each then transmits its segments as orphaned fragments carrying a concatenation header for a message that will never complete.
The handset holds them until its own timer expires and typically shows nothing. The customer is billed for all three. The only trace is one message.parts.incomplete warning per instance, neither of which knows the other exists.
A customer’s SMPP session is normally sticky at the TCP level, so this bites on reconnect mid-message, on load-balancer rebalance, and on clients holding several binds that spread segments across them.Content whitelisting still applies to what each instance sends — the timeout path is content-checked — so the failure is delivery and billing, not enforcement.

Degraded but not lost

Vendor TPS, ingress rate limits, throughput sampling and routing-group cursors are all per process. N instances send up to N × the contracted rate to a carrier. That is a guard against a runaway sender, not a distributed quota — see Limits.

What it would take

Broadly: move receipt correlation and reassembly into shared storage, give cdr_final a unique key and an upsert, and make the sweeper leader-elected or idempotent. Each is tractable; none is done. Until then, scale vertically, and use several independent instances where the traffic can be partitioned by customer.