Omrylo Blog

Designing a Simple, Evolvable AI Backend

A practical breakdown of sessions, agents, events, streaming, connection management, persistence, and progressive implementation.

An AI chat demo is easy to start. The boundaries become less obvious once the product needs history, streaming output, tool calls, reconnection, and more than one model.

My preference is to complete the smallest useful loop first, then let the architecture grow with observed problems. Version one does not need a complex agent platform, but it should distinguish a few lifecycles from the beginning.

Separate sessions, agents, and events

These identifiers represent different lifetimes. Collapsing them makes retries produce ambiguous duplicates and leaves reconnection without a clear point of recovery.

  • A Session is the durable conversation container that holds messages and user context.
  • An Agent is one execution created by a request, responsible for model calls, tools, and the state of that run.
  • An Event is an ordered fact produced during execution: start, text delta, tool call, error, or completion.

Why streaming often starts with SSE

Most AI chat traffic flows in one direction while an answer is being produced. Server-Sent Events fit that HTTP-based model with less machinery than WebSockets, work well with existing gateways, and support reconnection with Last-Event-ID.

SSE is not the universal answer. WebSockets remain useful for high-frequency bidirectional communication, collaborative state, or complex binary transfer. The communication model should choose the transport.

A small, explicit API surface

The client creates a run, receives an agentId, and then subscribes to its events. The application server owns authentication, context, and model calls. A connection manager only owns subscriptions and delivery.

Minimal API surface
POST /api/v1/agent/chat
{ sessionId, prompt } -> { agentId, status }

GET /api/v1/events/subscribe/:agentId
Accept: text/event-stream

Add capability in the order of risk

Every stage should produce a verifiable user experience. Reliably completing one answer is usually more valuable than introducing multi-agent orchestration too early.

  • V0.1: one request, one streamed response, and an explicit completion state.
  • V0.2: persist the Session and final answer so history survives a refresh.
  • V0.3: sequence Events and handle reconnection, duplicate consumption, and timeouts.
  • V0.4: add model routing, tool calls, quotas, logs, and observability.

The engineering boundaries that create stability

An evolvable AI backend is not the one with the most components. It is the one with clear lifecycles, ownership, and failure modes. Once those boundaries hold, models and agent behavior are much easier to extend.

  • Idempotency: retries do not create untraceable duplicate results.
  • Recovery: a disconnected client knows which Event to resume from.
  • Persistence: streaming deltas serve the live experience; final messages serve durable history.
  • Observability: Session, Agent, Event, model calls, and errors can be correlated.
  • Security: identity, permissions, input size, and tool scope are checked before a model call.