Errors and Troubleshooting
Catch operation failures, handle asynchronous server errors, and diagnose common SDK problems.
Before you start: follow Connecting and lifecycle. Register an error listener once per channel, and also catch exceptions from the operations you await.
Every error the SDK raises or reports is a CelerisError with a stable code. Match on code, never on message text. Messages name the field and the rule that failed, but never repeat your input, a credential or server text.
1. Catch the operation you started
from useceleris_client import CelerisConnectionError, text_payload
try:
await chat.publish(text_payload("Hello"))
except CelerisConnectionError as error:
if error.code == "NotConnected":
print("Wait until connected before sending.")
else:
print(error.code, error)
Catch connect(), publish() and presence_list() where you await them. A failed initial connect raises to its caller without also reaching on_error, and so does a failed presence query.
2. Handle server errors separately
from useceleris_client import ChannelError, ServerError
def handle(error: ChannelError) -> None:
if not isinstance(error, ServerError):
print(error.code, error)
return
print(error.type, error.sub_type, error)
if error.type == "PermissionDeniedError":
print("Check the server-issued channel and segment permissions.")
elif error.type == "MessageSizeLimitError":
print("Reduce the payload size.")
elif error.type == "RateLimitError":
# The SDK pauses and resends recent commands by itself.
print("Rate limited; publish less often if this repeats.")
else:
print("The server reported an error.")
stop_errors = channel.events().on_error(handle)
# Later:
stop_errors()
A ServerError carries type, sub_type (the command it answers, or None), its message, and resource (what that command names, such as the segment). Its code is "Server". Error types are open-ended, so keep a default branch for types a newer server adds. A resource can be None, a str, an int, or a tuple of these; check its type before using it.
A permission denial names a command and a segment, not which of several identical publishes caused it. Presence query errors are matched to their query and raise from presence_list() instead. Server error text is for people: display it as text, never as raw HTML. Server errors leave the connection open.
3. Choose the right response
code | Response |
|---|---|
Configuration | Fix the invalid argument or credential; retrying unchanged input will not help |
NotConnected | Wait for a connection; there is no offline queue |
Timeout | Check the credential endpoint and network, or the query deadline |
Cancelled | The channel was closed during the operation; do not restart cancelled work blindly |
Transport | Check network and credential setup; this does not prove an authentication failure |
Backpressure | Slow down: 64 publishes are already waiting, or a presence query found the writer full or paused |
OperationInProgress | Avoid simultaneous connects or presence queries on one channel |
DeliveryUnknown | Acceptance is uncertain; do not blindly resend an irreversible operation |
ProtocolError | One frame was dropped; the connection remains usable |
Server | An error the server sent; see step 2 |
Cancelling a task you own raises asyncio.CancelledError, not an SDK error. A listener's exception is contained and reported through on_error as a listener failure; catch errors in tasks your listeners start yourself.
4. Diagnose common symptoms
| Symptom | First checks |
|---|---|
| First connect fails and nothing retries | Expected: call connect() again from failed |
| Socket connects but messages are absent | Matching channel and segment, a held subscription, and read permission |
| Sender sees no received copy | allow_echo defaults to False |
| Publish returns, then an error arrives | Expected for server refusals: returning means only local acceptance |
Registering a listener raises ConfigurationError | The listener is an async def; use a plain function that starts a task |
| Presence events miss a user | Events are node-local; query a cluster-wide listing |
| Credentials were denied but the error says Transport | The handshake status is not available to the SDK |
| Updates are missing after recovery | Replay is bounded; reload authoritative application state |
Presence query raises Configuration | Use int: page 1 to 2147483647 and per_page 1 to 100 |
| A closed channel cannot reconnect | Obtain a new channel from the client |
See Reconnection and Recovery for which failures retry. Do not run a competing retry loop while the channel is already reconnecting.
Next: Server-side usage, or the Python client API for the full error surface.