Omrylo Blog

Reliable SSE Reconnection and Idempotency

Use ordered events, Last-Event-ID, idempotent writes, and final-message persistence to handle disconnection, duplication, and completion in streamed AI responses.

SSE makes a token-by-token AI response easy to demonstrate, but “the stream works” and “the response remains correct after interruption” are separate milestones. Mobile-network changes, sleeping browsers, proxy timeouts, and page refreshes all close connections. Automatic reconnection can then deliver an event more than once.

Reliability cannot depend on the client continuing to append strings. The system must know which ordered events belong to one execution, how far the client has confirmed, how duplicate delivery is recognized, and which persisted object represents the completed answer.

Treat an Event as a replayable fact

Each Agent execution produces an ordered event stream. An event needs at least a stable agentId, increasing eventId, type, payload, and creation time. Start, text delta, tool call, error, and completion are events rather than temporary network packets.

If an event exists only in the memory of the current connection, disconnection erases history. A recoverable design writes events to a short-lived store before or alongside delivery, indexed by agentId and eventId. Retention can be limited, but it has to cover the product's promised reconnection window.

Minimal event shape
{
  agentId: "A-999",
  eventId: 42,
  type: "text_delta",
  payload: { text: "..." },
  createdAt: "..."
}

Last-Event-ID is a recovery entry point, not the whole design

A browser reconnecting to SSE can send `Last-Event-ID`. The server should query events after that number, replay them in order, and then join the live stream. The important properties are stable numbering and an exact query boundary, not merely the presence of the header.

Recovery usually uses a strict greater-than rule: confirmation through 42 resumes from 43. If the execution cannot be found, the event window has expired, or the identifier is invalid, the server should return an explicit unrecoverable state. The client can then read a final message or offer a new run instead of silently rebuilding from the beginning.

Place idempotency at every state change

With at-least-once delivery, duplicates are expected. A client can deduplicate on `(agentId, eventId)`, but server-side tool execution, message persistence, and completion handling need their own keys. Otherwise, one retry may charge twice, send two notifications, or create duplicate final messages.

A practical boundary uses eventId for event append, toolCallId for a tool side effect, and agentId for the final message. A write checks whether the same key has already completed before performing work or returning the existing result. Idempotency does not remove repeated requests. It prevents them from changing the final facts twice.

Persist streaming deltas and final messages for different jobs

A text delta serves the immediate experience; a final message serves durable conversation history. When execution finishes, the backend writes one authoritative final message and then publishes a completed event. A refreshed page reads that message rather than reconstructing the answer from every historical delta.

If the connection closes before completion, the client first attempts event replay. If the event window is gone, it queries Agent status. A completed run returns the final message, a failed run returns an understandable failure, and an active run can be subscribed to again. This fallback is more controlled than reconnecting forever.

Recovery flow
Connection closes
  → Reconnect with Last-Event-ID
  → Replay missing events
  → If replay is unavailable, query Agent status
  → completed: read final message
  → failed: show failure and allow a retry

Define failure before selecting infrastructure

A small system can begin with a database event table and in-process broadcast. Higher concurrency may justify Redis Streams, a queue, or a dedicated event store. The decision should follow concurrency, retention, cross-instance delivery, and recovery objectives.

Whatever the components, the system must answer four questions: are events ordered, are duplicates safe, who writes completion, and what happens after replay expires? SSE reliability does not come from the long-lived HTTP connection itself. It comes from the explicit, replayable, and testable state model behind that connection.

Tests should target failure paths: disconnect after a random event, deliver one eventId twice, execute completion twice, expire the replay window, and then compare the page and database with the expected final answer. Until those scenarios are repeatable, reconnection remains an optimistic behavior rather than a reliability guarantee.