Why I Built Celeris
A founding story about Firebase's sharding problem, a Rust learning project that got serious, and the platform I wish had existed.
There's a specific kind of frustration that comes from knowing your architecture is wrong and not being able to fix it cleanly. Not a bug you can track down. Not a performance problem you can profile and optimise. A structural problem baked into a platform you don't control, that surfaces under load, at the worst possible times.
I lived with that feeling for longer than I should have. This is the story of what I did about it.
The Firebase problem
A few years ago I was working on a product that needed realtime functionality at serious scale. High concurrency, high message throughput, many simultaneous connections. The kind of thing where the realtime layer isn't a feature. It's load-bearing.
We were using Firebase. Specifically Firebase Realtime Database.
RTDB is genuinely good for a lot of use cases. Small teams, moderate concurrency, quick prototypes: it works. The problem is that it has no concept of horizontal sharding. There's one database. When you fill it up or overwhelm it with concurrent connections, you don't get a graceful degradation path. You get a ceiling.
We hit that ceiling.
The naive solution (just put everything in one RTDB instance) was off the table. So I started designing a workaround. The approach I landed on used Firestore as the primary datastore and selectively replicated the data that needed realtime access into multiple RTDB instances. Users were routed to whichever RTDB shard held their data. In theory, this distributed the load. In theory.
In practice, we had data living in two places simultaneously. The routing logic had to stay in sync with where data lived. State was scattered across Firestore and however many RTDB shards were currently active. If a user's shard was under load, their experience degraded, but nothing in the architecture surfaced that cleanly. We were stitching together two products that weren't designed to work this way, and every seam showed.
Every time traffic spiked, we found a new failure point. The architecture worked until it didn't, and "until it didn't" happened more often than I was comfortable with. I spent time I should have spent building product features instead thinking about which shard a user belonged to and why a connection was dropping.
The thing that sticks with me isn't the technical embarrassment of the workaround. It's the feeling of being responsible for a system built on a foundation you don't trust. You check your dashboards more than you should. You dread traffic spikes instead of hoping for them.
Looking for something better
After that experience I started looking properly at the alternatives.
The requirements were clear enough: high throughput, high concurrency, multi-region, and critically, a way to handle realtime delivery without the database becoming a liability under load.
That last one needs unpacking, because it shaped everything that came after.
The pattern most teams follow is: receive a realtime event, write it to Postgres (or whatever relational database is at the center of the stack). This works at modest scale. At serious scale, a burst of incoming messages becomes a burst of write pressure directly on the database. Relational databases are not designed to absorb that kind of load without careful protection. Connection pools fill. Latency climbs. If you're unlucky, things cascade.
I'd been close enough to this problem to know it wasn't hypothetical. What I wanted was a delivery layer that didn't require the database to be on the hot path. Receive the event, deliver it to subscribers, and let the application persist it at its own pace, not at the pace of the traffic spike.
I looked at Ably, Pusher, and PubNub. They're all solid products in different ways. But none of them are solving this at the architecture level. They're managed convenience layers over WebSocket delivery. They'll get your messages to clients, reliably, with good tooling. What they don't do is give you a Kafka-backed delivery layer with a write-only export connector so you can receive messages into your own queue and persist to your database when you're ready. That design decision, the one that would have changed everything in that previous job, wasn't there.
So I decided to build it.
The honest origin story
Here's the part I could dress up but won't: the initial reason I picked Rust was that I wanted to learn Rust.
I'd been writing mostly TypeScript and Go. Rust had been on my list for a while, for its reputation for correctness, its memory model, and the absence of a garbage collector. I wanted to build something non-trivial in it. Not a toy project. Something with real concurrency, real IO, real pressure. A WebSocket server seemed like the right shape of problem.
The realtime infrastructure problem I'd been thinking about gave me a reason to make it real rather than hypothetical. So I started building a WebSocket platform in Rust, using Actix Web and Tokio, backed by Kafka, with Redis for connection state.
The learning-project framing didn't survive contact with the results.
The number that changed things
At some point during development I ran a proper load test. Not a quick smoke test. A sustained benchmark. I pointed k6 at a single AWS c8g.large instance and let it run for 25 minutes.
The numbers: 96,114 messages per second. 102MB of memory. 56% CPU.
I looked at that output for a while.
I knew Rust was efficient in theory. Knowing it and seeing 96,000 messages per second at 102 megabytes of memory on a single mid-range EC2 instance are different things. That's not a carefully tuned result extracted from ideal conditions. That's a sustained test, running for 25 minutes, including the overhead of the k6 clients that were generating the load.
What this told me was that the economics of this platform were different from what I'd assumed. A single node could handle what would require a cluster of less efficient alternatives. The memory footprint meant you could run many more instances on the same hardware budget. The CPU headroom meant you still had room to scale vertically before you needed to scale horizontally.
At that point this stopped being a learning project. If a single node could do this, then developers who needed high-throughput realtime infrastructure deserved access to it as a managed service, without having to build the underlying platform themselves.
The binary gap
While I was building, I kept noticing something about the existing platforms: they're all built around JSON.
Not as a criticism, exactly. JSON is universal, human-readable, and easy to work with. But it's also expensive. Verbose on the wire. Slow to parse at high volume. And for certain categories of application (games, trading infrastructure, IoT, anything where payload size and parse latency are genuine constraints) the cost is real. Not philosophical. Bandwidth costs money. CPU cycles at scale cost money. Serialization overhead on every single message adds up.
MessagePack and Protobuf exist. They're genuinely better formats for high-frequency binary data. But in the existing managed platforms they're either entirely absent or treated as afterthoughts. Ably documents MessagePack support; it's not the default and it's not a first-class design decision. Protobuf is something you handle yourself. The transport layer is fundamentally built around text, and you pay the tax regardless.
I wanted a platform that treated binary formats as first-class from the start. Not a plugin or a compatibility mode. The default path for teams that know what their data looks like and don't want to serialize it into strings first.
What crystallised
Somewhere between the first working prototype and making it multitenant, a set of principles emerged that I wasn't willing to compromise on.
Pure WebSocket underneath. The SDK is a convenience layer, not a proprietary protocol. If you want to connect with raw WebSocket and skip the SDK entirely, you can. The protocol is documented. You're not trapped.
Your data goes to infrastructure you own. BYOD: bring your own destination, coming soon. Write-only export connectors to Kafka, RabbitMQ, S3, HTTP. From the moment Celeris accepts a message, you'll be able to have it flowing into your own queue. You consume from it at your own pace. Your database is never on the spike path.
Honest pricing. Three dimensions: messages, connection minutes, channel minutes. Published on the website. No sales call required to find out if you can afford it. The ten-year option bills the base plan fee upfront and renews automatically every ten years unless canceled. Usage charges remain separate.
A transport layer, not an application layer. Celeris does not want to be your chat SDK or your collaborative editing engine. It wants to be the thing you build those on top of. One job. Done well.
What it is now
Celeris is in public beta. Two regions, US and EU, with global routing designed to reduce network delays and keep connections more consistent. A JavaScript and TypeScript SDK today, with Java, Rust, Python, and Flutter coming soon. A dashboard where you can send test messages and verify everything works before you write a line of client code. Kafka-backed per-channel ordering. Binary-first transport. BYOD export connectors are next.
It's v1. There's more to build. But the foundation, the part that actually mattered, the part that would have changed things at that previous job, is there and it's working.
This is the platform I wish had existed. Now it does.