WebSockets vs Server-Sent Events (SSE): Building Low-Latency AI Token Streaming Interfaces
Architecting real-time streaming interfaces for generative AI: Server-Sent Events vs WebSockets, handling network drops, and consuming HTTP event streams in React.
By Uttam Thapa · · Backend
⚡ Executive Summary (TL;DR)
LLM token streaming is one prompt out and several hundred tokens back — strictly one-directional, which is exactly the shape
Server-Sent Events was designed for. SSE is plain HTTP, so it needs no separate server, passes through proxies and
corporate firewalls unchanged, and reconnects on its own. Reach for a WebSocket when the client also needs to stream, not by default.
Figure 1: Tokens arriving incrementally. The perceived latency is the first token, not the last.
Match the Protocol to the Traffic Shape
WebSockets are the reflex answer to anything described as "real time", and for a collaborative whiteboard or a multiplayer game that reflex is right — both
sides transmit continuously. Generative AI does not look like that. The client sends one prompt and then goes quiet while the server produces a long stream. A
full-duplex connection to carry traffic in one direction is capability you pay for and never use.
|
Server-Sent Events |
WebSockets |
| Direction | Server → client only | Both ways |
| Protocol | Plain HTTP — no upgrade | Upgrade to ws:// |
| Reconnection | Automatic, with event replay | You implement it |
| Proxies and firewalls | Ordinary HTTP response | Often blocked or stripped |
| Payload | UTF-8 text | Text or binary |
| Infrastructure | Your existing HTTP server | A socket server plus an adapter to scale |
The Express Endpoint
SSE is a content type and a text format. There is no library involved.
app.get('/api/ai/stream', (req, res) => {
res.setHeader('Content-Type', 'text/event-stream');
res.setHeader('Cache-Control', 'no-cache');
res.setHeader('Connection', 'keep-alive');
res.setHeader('X-Accel-Buffering', 'no'); // stop nginx buffering the stream
const tokens = ['Analyzing ', 'weather ', 'telemetry ', 'for ', 'Mumbai...'];
tokens.forEach((token, idx) => {
setTimeout(() => {
res.write('data: ' + JSON.stringify({ token }) + '\n\n');
if (idx === tokens.length - 1) res.end();
}, idx * 100);
});
// A client that navigates away must not leave the generation running.
req.on('close', () => res.end());
});
🚨 Three details that break SSE in production
- The double newline is the delimiter. One
\n and the client waits forever for a message that already arrived.
- Reverse proxies buffer by default. Without
X-Accel-Buffering: no, nginx holds the stream and delivers it in one lump — indistinguishable from no streaming at all.
- Handle
req.on('close'). Otherwise a closed tab leaves an inference running and billing.
Consuming the Stream in React
EventSource is the built-in client and is excellent — but it only issues GET requests and cannot set headers, so it cannot carry an
Authorization header or a POST body. For an authenticated prompt submission, read the response body as a stream instead:
const response = await fetch('/api/ai/stream', {
method: 'POST',
headers: { 'Content-Type': 'application/json', Authorization: token },
body: JSON.stringify({ prompt }),
signal: controller.signal, // lets the user cancel a generation
});
const reader = response.body?.getReader();
const decoder = new TextDecoder();
let buffer = '';
while (true) {
const { value, done } = await reader!.read();
if (done) break;
// Chunks split anywhere, including mid-message. Buffer until a delimiter.
buffer += decoder.decode(value, { stream: true });
const parts = buffer.split('\n\n');
buffer = parts.pop() ?? '';
for (const part of parts) {
if (part.startsWith('data: ')) appendToken(JSON.parse(part.slice(6)).token);
}
}
The buffering is not optional. A network chunk boundary falls wherever TCP decides, frequently in the middle of a JSON payload, so parsing each chunk directly
works in development and throws intermittently in production.
💡 Serverless timeouts are the real constraint
A streamed response holds a function invocation open for its whole duration, and most serverless platforms cap that. A long generation can be cut off
mid-sentence by the platform rather than by your code. Check the limit on your plan before choosing where the streaming endpoint lives — this catches far
more teams than any protocol detail.
Resuming Rather Than Restarting
Native EventSource reconnects automatically, and if you emit an id: field with each message the browser sends it back as
Last-Event-ID on reconnect. Track which token index a client has received and resume from there — a dropped connection then costs a hiccup rather
than the whole response. This is machinery you would have to build yourself over a WebSocket.
Choosing, Concretely
📡 Use SSE for
- • LLM token streaming
- • Notification and activity feeds
- • Build and deploy logs
- • Live price or metric tickers
- • Long-running job progress
🔌 Use WebSockets for
- • Collaborative editors and whiteboards
- • Chat with typing indicators and presence
- • Multiplayer games
- • Anything sending binary frames
- • High-frequency client-to-server input
✅ Key takeaways
- ✓Pick the protocol from the traffic shape. One-directional streams do not need a duplex channel.
- ✓SSE is plain HTTP. No extra server, no upgrade handshake, no proxy surprises.
- ✓Disable proxy buffering. A buffered stream is not a stream.
- ✓Buffer to the delimiter on the client. Chunk boundaries land mid-payload.
- ✓Close the stream when the client leaves. An abandoned generation still costs money.
- ✓Check your platform's function timeout before committing to streaming from it.
For the duplex side of this, see architecting a real-time chat engine.
Frequently asked questions
Should I use SSE or WebSockets for streaming LLM tokens?
SSE. Token streaming is one-directional, and SSE runs over plain HTTP with automatic reconnection built in, so it works through proxies and needs no separate server.
What are the limitations of Server-Sent Events?
It is server-to-client only, text only, and older browsers limit concurrent connections per domain over HTTP/1.1. If the client needs to send a continuous stream too, use a WebSocket.
How do you handle a dropped connection while streaming?
SSE reconnects on its own and replays from the last event id if you send one. Track what the client has already received so a reconnect resumes rather than restarting the response.
Home · Projects · Blog · Services · Résumé · Contact