Architecting a Real-Time Chat Engine: WebSockets, Socket.IO Rooms, and Heartbeat Presence
Inside the real-time engine: handling socket reconnection, scaling broadcast channels across rooms, managing online status detection, and persistent chat storage.
By Uttam Thapa · · Backend
⚡ Executive Summary (TL;DR)
Polling is a workaround for not having a connection. This is the architecture behind a React, Node.js, Express, Socket.IO and MongoDB chat platform:
rooms to keep broadcasts scoped, asynchronous persistence so a database write never delays delivery, heartbeat-driven presence that survives a phone going
into a tunnel, and the adapter you need the moment there is more than one server process.
Figure 1: Messages, typing indicators and presence are three different event streams over one connection.
Polling, Long Polling, or a Real Connection
Knowing when to leave HTTP behind is the first architectural decision, and the cost difference is not subtle.
| Approach |
How it works |
What it costs |
| Short polling |
Request every two seconds regardless of activity |
Full HTTP headers per request, battery drain, poor concurrency |
| Long polling |
Server holds the request open until data exists |
Better, but constant connection churn and reconnect storms |
| WebSockets |
Persistent bidirectional TCP connection |
2–8 bytes of frame overhead; sub-50 ms propagation |
The decisive difference is direction. Polling means the client repeatedly asks whether anything happened; a socket means the server says so. Typing indicators,
presence and read receipts are all server-initiated, and none of them are viable over polling at any acceptable cost.
Rooms: Scoping the Broadcast
Without rooms, every message is delivered to every connected socket and filtered on the client — which is both a privacy problem and an O(n)
bandwidth bill per message. Socket.IO rooms move that filtering to the server.
io.on('connection', (socket) => {
socket.on('join_room', (roomId) => {
socket.join(roomId);
});
socket.on('send_message', (data) => {
// Deliver first: everyone in the room except the sender.
socket.to(data.roomId).emit('receive_message', data);
// Persist after. A slow write must never delay delivery.
saveMessageToDatabase(data);
});
});
🚨 Authorise the join, not just the connection
socket.join(roomId) takes whatever the client sends. Without a membership check on that handler, any authenticated user can join any room by
guessing an id — a private-conversation leak that no amount of frontend routing prevents. Verify room membership server-side inside the
join_room handler, every time.
Deliver First, Persist Second
Emitting before writing is deliberate. A MongoDB insert takes single-digit milliseconds when healthy and considerably longer when it is not; putting it in front
of the emit means every participant's chat latency tracks your database's worst percentile.
The trade is that a message can be delivered and then fail to persist. Handle it explicitly: assign the message a client-generated id, acknowledge the write back
to the sender, and surface a retry affordance if that acknowledgement never lands. Users tolerate "not sent — retry" far better than a chat that feels sluggish.
Presence Without Lies
Presence is deceptively hard because the interesting case is not a user logging out — it is a user walking into a tunnel. No logout event is ever
emitted; the connection simply stops answering.
🟢
Online
Heartbeat answered within the interval. The connection is confirmed alive, not merely opened once.
🟡
Away
Socket connected, but the tab reports hidden via the Page Visibility API.
⚪
Offline
Heartbeat missed past the timeout; Socket.IO fires disconnect and peers update within ~20 seconds.
Socket.IO's built-in ping/pong does this work for you, and the timeout is a product decision as much as a technical one. Shorter means presence indicators
correct themselves quickly but flicker on flaky mobile networks; longer means stable indicators that are occasionally wrong.
💡 The moment you run two server processes
Rooms are per-process. Scale to two instances behind a load balancer and users in the same room, connected to different instances, stop seeing each other's
messages — with no error anywhere. You need a Socket.IO adapter (Redis is the usual choice) to broadcast across instances, and sticky sessions at the balancer
so a reconnecting client returns to a process that knows it. Plan for this before the first horizontal scale, not during the incident.
✅ Key takeaways
- ✓Use rooms to scope broadcasts. Client-side filtering is a leak with a bandwidth bill attached.
- ✓Authorise every
join. Room ids arrive from the client and must be checked like any other input.
- ✓Emit before you persist. Then acknowledge the write so the sender learns if it failed.
- ✓Derive presence from heartbeats. Explicit logout events do not fire when a network disappears.
- ✓Add the adapter before the second instance. Cross-process broadcast failures are silent.
Choosing between a socket and a one-way stream for a given feature? WebSockets vs Server-Sent Events
covers when a full duplex connection is more than you need.
Frequently asked questions
When should you use WebSockets instead of HTTP polling?
As soon as the server needs to initiate messages. Typing indicators, presence and live delivery are all server-driven, and polling makes them either expensive or impossible.
How do you stop users joining chat rooms they should not see?
Authorise inside the join handler. The room id arrives from the client like any other input, so without a membership check any authenticated user can join a private room by guessing an id.
Why do messages stop working after adding a second server instance?
Socket.IO rooms are per-process, so users connected to different instances no longer share a room. You need a Redis adapter to broadcast across instances and sticky sessions at the load balancer.
Home · Projects · Blog · Services · Résumé · Contact