Skip to main content
Three different problems wear the same face. Tell them apart first, because the fixes share nothing.

Telling “refused” from “never reached”

bound: 0 alone cannot distinguish “nobody connected right now” from “no socket was ever opened”. last_error is what separates them, and it clears the moment a bind succeeds.
Dials once and tells you which, without waiting for the reconnect cycle.
test-bind is refused while the vendor is already bound, deliberately. A second bind on the same systemId can cost you the live session, and many carriers allow only one. Use it before traffic depends on the vendor, or after suspending it.

A vendor will not bind

  • host, port, username and password are correct.
  • connections.transceivers, connections.transmitters or connections.receivers is greater than zero — a vendor with all three at zero binds nothing and reports no error.
  • TLS settings match what the provider requires.
  • Firewall allows outbound traffic to the provider.
  • The provider has allow-listed your source IP, if they require it.
Bound but nothing going out? Compare bound with bound_transmittable. A session carries traffic only if it is a transceiver, or a transmitter on a client session. A receiver-only vendor binds perfectly and sends nothing.

A customer’s bind is rejected

  • The credential is type: SMPP. An HTTP credential cannot bind, and the refusal does not say so.
  • The client is using the configured systemId and password.
  • No connection limit is exceeded: srv.maxConnections, srv.maxConnectionsPerIP, conf.maxConnectionsPerUser.default, or the per-user limit. Since 0.9.20 the refusals are counted by reason — fireflo_smpp_bind_refused_total on the metrics endpoint, and the Servers page shows a notice when any have happened. A customer fleet behind one NAT counts as one source IP against the per-IP limit.
  • If allowedIps is configured, the bind source IP is on it.

The listener never opened

Different from every case above: not a rejected bind, but a socket that was never created. The gateway runs normally — routing, rating, billing and vendor connections all up, systemd reporting active — while no customer can connect. The gateway does not exit when a listener cannot take its port. That is deliberate: one busy port must not take a whole gateway off the air, and TLS or proxy listeners are often not configured at all. So the evidence is in three places rather than in the exit status. On /ops/health:
In the log:
And the port names its holder:

The usual cause is two gateways

Either a second instance configured on the same port, or the same instance running twice under two unit names — which is what an update given the wrong --service-name used to produce.
1

List every unit

systemctl list-units 'fireflo*' --all
2

Compare the managed PID against the one holding the port

systemctl show <unit> -p MainPID — an orphan JVM holding the port is one systemd no longer manages.
3

Stop the service, kill the orphan by PID, start it again

By PID, never by pattern. Other JVMs may be on the host, and a pattern kill has taken out unrelated services before.
A listener that failed to bind retries on the next configuration poll or reload — it does not need a restart once the port is free. (Before 0.3.4 it did: the failed server was retained, so every later attempt logged not (re)starting and returned.)
A clean start is not proof every listener is up. Since the gateway logs and carries on, the absence of an error at boot means nothing. Check /ops/health instead.

Servers

Listener settings and what each one does.

Vendors

Outbound connections, TPS and holding.

Blocking abusive clients

A customer whose connection is refused before it reaches the gateway may be blocked at the firewall. Check this before anything else.