AI & agent infrastructure
Keep every agentic session continuous and in sync. Binary-first means streaming tokens efficiently without JSON overhead, built for the latency AI sessions demand.
The problem
AI sessions generate high-frequency, low-latency traffic. Every token streamed from a model is a small payload, but JSON encoding each one, opening and closing HTTP connections, or polling for updates adds up fast when users expect instant, continuous delivery.
Most realtime platforms weren't designed for this pattern. They're optimised for occasional messages between humans, not for hundreds of small binary frames per second from a model inference server.
How Celeris solves it
Celeris is a persistent WebSocket transport. Once the session is open, it stays open. No reconnection overhead per token, no HTTP request round-trip per chunk. Binary-first means each token frame is as small as it can be: MessagePack encodes it efficiently, no JSON wrapping required.
Each agent session maps naturally to a channel segment. Multiple users watching the same session, or multiple agents collaborating on a shared context, write and read to the same segment without extra coordination logic.
p95 latency at 32ms means the transport layer isn't the bottleneck between the model and the user.
Opening a streaming session
const client = createClient({ credentialProvider });
const channel = client.channel("agent-sessions");
const session = channel.segment(`session-${sessionId}`);
session.onMessage((payload) => {
appendToken(readJson(payload).token);
});
session.subscribe();
await channel.connect();
// From your inference server, publish each token as it arrives
await session.publish({
payload: jsonPayload({ token: nextToken, done: false }),
});
The session stays live until you close it. Reconnection logic is handled by the SDK.
Relevant features
- Binary-first messaging
token frames, no JSON overhead per chunk - Channels and segments
segment per session, multiple observers supported - p95 32ms latency
transport doesn't add perceived latency to generation - Persistent WebSocket
reconnect overhead between tokens