Skip to content
SpanDock

Satellites and failover

One server is enough for most teams. When you want to spread client connections and delivery across machines, add satellites: extra servers that join your first server, the hub. Clients are assigned to servers automatically and fail over to another server by themselves when theirs stops answering, without losing data. No external database is needed.

How it fits together

  • The hub keeps the settings, the client list, and the full history in its embedded database. It assigns clients to servers and serves clients like any server.
  • A satellite accepts clients, checks them against a copy of the hub's client list, forwards telemetry to your destinations itself, and sends every record to the hub for storage.
  • Clients connect to one server at a time: their assigned primary, with an ordered fallback list.

Settings are edited on the hub only. A satellite's dashboard shows its own status, its connected clients and live activity, with a note that settings are managed by the hub.

Add a satellite

  1. On the hub: open Settings → Scaling → Add server and copy the join code. It works once, for 30 minutes.
  2. On the new machine: install SpanDock and choose Join it as a satellite on the first launch, then paste the code.

On a machine without a screen, start it once with the join code in the environment:

SPANDOCK_JOIN_CODE='…' spandock -role=server -open=false -menubar=false

Pass the code in the environment variable, not as a command-line argument, so it stays out of the process list.

Once a satellite has joined, Clients → Add client on the hub asks which server the new client connects to.

Server assignment

The hub gives every client a primary server and a fallback list, and pushes them to the client over the encrypted tunnel.

  • Sticky: a client whose primary is healthy keeps it. Moving costs a reconnect, so SpanDock avoids it.
  • Spread: a client without a healthy primary goes to the least loaded server. Fallback lists are rotated per client, so when a server fails its clients spread across the others instead of all landing on one.
  • Gradual rebalancing: if one server becomes much busier than the average, a small share of clients is moved at a time, spread out, so there is never a reconnect stampede.
  • Pin to server: in a client's details you can pin it to one server whatever the load. A pinned client still fails over if that server stops answering.
  • Drain: choose Drain for a satellite under Settings → Scaling to stop sending it clients and remove it from fallback lists. Its clients move within a minute. Disconnect it once it has none.

Automatic failover

When a client's server stops answering, the client first tries to heal the connection to the same server, so a short restart or update doesn't bounce it around. If that hasn't worked after about 20 seconds, it moves to the next server in its fallback list, and on down the list.

  • No data loss: while the client is moving, telemetry goes into its on-disk queue. After the switch, the queue is replayed to the new server. Replays are deduplicated, so nothing is stored twice.
  • Going back: a client on a fallback checks its primary every 30 seconds and returns only after a minute of continuous health, with a small random delay so a recovered server doesn't receive all its clients at once.
  • When nothing answers (for example, the laptop itself is offline), the client keeps queueing and waits longer between attempts, up to 5 minutes.
  • Long outages: if a server stays down, the hub makes each of its clients' current server their new primary.

The client's Settings → Server connection shows which server it is on, whether that is its primary or a fallback, and when it last failed over.

The cluster diagram

The Cluster page on the hub is a live diagram of your servers and clients. Servers are coloured by state: green online, orange draining, red down. A star marks the hub. Lines from each client to its current server move with its live traffic, and when a client fails over you see its line move to the new server. An event log lists satellites joining, going offline and coming back, and clients failing over or returning. A table view lists the same information. The Clients page also tags each client with the server it is on.

Limits

  • The hub is a single point of failure for settings and history. It is not replaced automatically.
  • While the hub is down, satellites keep accepting telemetry from their clients and keep forwarding it to your destinations. Records meant for the hub's history are held on each satellite and sent in order when the hub returns. Settings are read-only in the meantime: no new pairings, approvals or revocations. Client history pages show that history is on the hub.
  • Storage throughput is limited to what the hub's embedded database can take. Connections and destination delivery scale out across satellites.

The Cluster page shows how many batches satellites are holding for the hub. A backlog that keeps growing is the signal that you have outgrown one hub.

High availability

If you need a high-availability cluster with no single point of failure, it is part of the Enterprise plan. Contact sales@spandock.com to discuss it.