Reconnection and Recovery
Understand how Celeris reconnects, what it restores, when it stops, and what your application must recover.
A reconnection opens a new socket after a working connection is lost. Recovery is the work that follows: restoring subscriptions and reconciling any updates missed during the outage.
The SDK handles the connection. Your application still owns its data. A connected socket is not proof that every missed message has arrived.
Before this page, read Channels and Segments. For JavaScript event handlers you can use in your app, see Connecting and lifecycle.
1. Know when recovery starts
Automatic recovery starts only after a channel has reached connected. A socket close or transport error starts it.
| Event | What happens |
|---|---|
| A working socket closes or encounters a transport failure | Enter reconnecting and schedule an attempt |
The first connect() fails | Reject that call and enter failed; your application decides whether to call connect() again |
| A server permission denial or rate limit | Report the server error; keep the connection open |
| A frame cannot be decoded | Drop that frame, report the error, and keep the connection open |
| Your message listener throws | Contain and report the listener error; keep the connection open |
Your application calls close() | Cancel recovery and close permanently |
idle -- connect() --> connecting -- success --> connected
| |
failure socket lost
| |
v v
failed <------------ reconnecting
| | |
explicit connect() retry success
| | |
+--> connecting +-----------+--> connected
close() from any active state --> closing --> closed
failed means automatic attempts have stopped, not that the channel is permanently unusable. closed is permanent: create a new channel handle to connect again.
2. Wait between attempts
The SDK uses exponential backoff: the maximum wait increases after failed attempts. Full jitter means it chooses a random delay anywhere between zero and that maximum. This spreads reconnecting clients out instead of making them all retry together.
| Failed reconnect attempts already used | Next delay range |
|---|---|
| 0 | 0 to 0.5 seconds |
| 1 | 0 to 1 second |
| 2 | 0 to 2 seconds |
| 3 | 0 to 4 seconds |
| 4 | 0 to 8 seconds |
| 5 | 0 to 16 seconds |
| 6 to 9 | 0 to 30 seconds |
The formula is random() * min(30 seconds, 500 milliseconds * 2^index). These are delay ceilings, not fixed waits.
There is a budget of ten failed reconnect attempts. The counter increases when a reconnect attempt fails with Transport or Timeout. A successful reconnect does not consume another failure, and does not reset failures already used.
At the next disconnect, the budget resets only if the connection just stayed up for at least 60 seconds. Brief successful connections retain the earlier failure count. An explicit connect() from failed also starts a fresh budget.
Do not use these timings as an application timeout guarantee. Each attempt also takes time, browser timers may be delayed, and repeated successful connections change the overall timeline.
3. Fetch fresh credentials and open a socket
For every attempt the SDK calls your credential provider again. During recovery the request includes:
| Field | Meaning |
|---|---|
channelReference | The channel being opened |
reason | "reconnect" |
disconnectedAt | Number: when this outage began, in Unix milliseconds |
replayLookbackMs | Number: suggested recent-history window in milliseconds |
signal | Cancellation signal to pass to your credential request |
The suggested lookback is the outage duration, rounded up to milliseconds, plus a five-second overlap. It is capped at 4,294,967,295 milliseconds. Failed attempts do not move the start of that outage, so later attempts request a larger window.
A suggestion is not permission. Your application's credential endpoint, or trusted-server claims callback, must authorize replay and explicitly sign replay: { lookbackMs }. It may cap or refuse the requested window. The SDK cannot grant history access by itself.
One configurable deadline, 15 seconds by default, covers credential acquisition and the WebSocket handshake together. Forward the request's signal to fetch so closing the channel also cancels an outstanding credential request. Late results from an abandoned attempt cannot establish the current connection.
See Authentication for server-side replay policy.
4. Restore subscriptions
When the new socket opens, the SDK restores the interests your application still holds:
- Restore message subscriptions, in registration order, except for
default, which Celeris joins automatically. - Restore presence subscriptions, in registration order, including any presence subscription to
default. Automatic message membership does not subscribe to presence. - Report the state as
connected. - Emit a recovery event with
retryIndex,possibleGaps: true, andpossibleDuplicates: true.
Here, retryIndex is the zero-based failure count used for that successful attempt, not a lifetime count of reconnections.
Your listeners stay registered; do not add a second copy after every reconnect. Restoring interests means sending the subscription commands, not receiving permission receipts or waiting for replay to finish. A server can still reject a subscription through the error handler.
Only subscriptions are restored. Previously published messages are never resent.
5. Handle calls made during an outage
| Your operation | While reconnecting |
|---|---|
| Publish | Rejects with NotConnected; nothing is queued |
| Query presence | Rejects with NotConnected |
| Subscribe to messages or presence | Records local interest for restoration |
| Cancel a subscription | Removes local interest so it will not be restored |
Call connect() again | Rejects with OperationInProgress |
Call close() | Cancels the retry timer and any active attempt, then closes |
A presence query already in flight when the connection drops rejects with Transport. The SDK does not retry that query; your application can issue a new one after recovery.
A practical UI disables sending while disconnected and displays the current connection state. Do not automatically retry a publish with uncertain delivery: the recipient may already have received it.
6. Understand when automatic attempts stop
During reconnection, Transport and Timeout failures consume the retry budget and schedule another attempt until ten failures have occurred.
A socket that refuses a subscription-restoration write is closed and counts as a Transport failure. Other attempt failures stop recovery immediately, such as malformed credential-provider output (Configuration) and cancellation. An explicit close instead ends in closed.
When recovery fails, the SDK invalidates old attempts and timers, reports the terminal error through onError, and enters failed. Your application may offer a Retry button that calls connect(). That explicit call resets the retry count, outage information, and duplicate window; its credential request has reason: "initial" and no suggested recovery lookback.
A rejected WebSocket handshake appears as Transport, including when credentials caused the rejection. The SDK cannot inspect the handshake's HTTP status, so it cannot classify that as an authentication failure. During recovery it retries; on the first connection attempt it rejects the caller.
7. Reconcile application data
Replay is recent, bounded history on the server, applied per segment join. It is not a durable log, a public sequence cursor, or a guarantee that every event during the outage is available.
The client remembers 1,024 recently delivered message IDs per channel. That window survives automatic reconnects and filters repeated IDs before delivering them to listeners. An older duplicate can still arrive after its ID leaves the window. An explicit connect() clears the window.
After the recovery event:
- Fetch important state from your application's authoritative API.
- Reconcile live updates with that snapshot using application-owned versions or identifiers.
- Refresh presence when a current connection list matters.
- Keep irreversible actions idempotent in your application.
For a chat app, reload stored messages from your database. For an order dashboard, fetch the latest order status. Celeris transports updates; it does not replace that stored state.
There is no offline queue, delivery receipt, replay-complete marker, or global ordering guarantee, and a reconnect never resends a publish. Read Message Ordering for ordering boundaries, then Connecting and lifecycle to wire recovery into JavaScript.