Skip to main content

Realtime Chat (WebSocket)

WSS /api/llm/chat/realtime

Streaming counterpart to POST /chat/completions: same X-API-Key auth, the same automatic system-prompt/MCP resolution, and the same usage billing, but the assistant's reply comes back token-by-token over a persistent connection instead of as one JSON body once generation finishes.

This page is maintained by hand rather than generated from the OpenAPI spec — OpenAPI has no way to describe a WebSocket operation, so it's absent from openapi.public.json even though the endpoint itself is fully public.

Connecting​

wss://api.r-mad.ai/api/llm/chat/realtime

Send X-API-Key as a header on the WebSocket handshake request, exactly as you would for any REST call (see Authentication) — the handshake is an ordinary HTTP request before it's upgraded, so the same header works:

import asyncio
import json
import websockets

async def main():
async with websockets.connect(
"wss://api.r-mad.ai/api/llm/chat/realtime",
additional_headers={"X-API-Key": "YOUR_API_KEY"},
) as ws:
await ws.send(json.dumps({
"type": "message",
"messages": [{"role": "user", "content": "Say hello in one sentence."}],
"return_markdown": False,
}))
async for raw in ws:
event = json.loads(raw)
if event["type"] == "delta":
print(event["text"], end="", flush=True)
elif event["type"] == "done":
print()
break
elif event["type"] == "error":
print(f"\nerror: {event['message']}")
break

asyncio.run(main())
// Node (e.g. the `ws` package) can set the handshake header directly.
// Browsers' native WebSocket API cannot set custom headers on a handshake
// at all — proxy the connection through your own backend if you need this
// from client-side browser code, the same way you would for any other
// credentialed API call you don't want shipped to a browser.
import WebSocket from "ws";

const ws = new WebSocket("wss://api.r-mad.ai/api/llm/chat/realtime", {
headers: {"X-API-Key": "YOUR_API_KEY"},
});

ws.on("open", () => {
ws.send(JSON.stringify({
type: "message",
messages: [{role: "user", content: "Say hello in one sentence."}],
return_markdown: false,
}));
});

ws.on("message", (raw) => {
const event = JSON.parse(raw);
if (event.type === "delta") process.stdout.write(event.text);
else if (event.type === "done") console.log();
else if (event.type === "error") console.error(`\nerror: ${event.message}`);
});

A missing or invalid key closes the connection immediately with WS close code 4401, before any messages are exchanged — same failure verify_partner_api_key would produce as a 401 for the REST endpoint.

Wire protocol​

JSON text frames, one exchange per conversation turn. The connection stays open across many turns — you keep sending message frames on the same socket rather than reconnecting each time.

Client → server

{
"type": "message",
"messages": [{"role": "user", "content": "..."}],
"return_markdown": false,
"system_prompt": "optional, layered on top of the org/key default for this turn only",
"user": "optional opaque end-user id, recorded against usage"
}

messages is the full conversation so far, OpenAI chat format — same contract as messages on the REST endpoint. The client owns conversation history; the server doesn't retain it between turns.

return_markdown is required on every turn, same contract as the REST endpoint's return_markdown — true asks the assistant to format its reply as Markdown, false for plain prose. A turn that omits it, or sends a non-boolean, gets an error frame for that turn (the connection stays open) rather than a 400 — a WebSocket has no per-message HTTP status to fail with.

Server → client, in order, per turn:

{"type": "delta", "text": "..."}

Repeated as the reply streams in.

{"type": "done", "usage": {"prompt_tokens": 12, "completion_tokens": 11, "total_tokens": 23}}

Ends the turn. usage is null only if an MCP tool loop hit its iteration cap without producing a priced final answer.

{"type": "error", "message": "..."}

This turn failed — the connection stays open, ready for another message frame.

Unlike the REST endpoint, a turn doesn't accept max_tokens, temperature, or the other sampling fields — every turn runs at a fixed 4096-token cap and temperature 0.2. top_p, stop, frequency_penalty, presence_penalty, and seed aren't available over this endpoint yet.

Rate limits and quota​

Enforced per turn, not just at connect — see Rate Limits for the underlying RPM/quota numbers, which are identical to the REST endpoint's and share the same per-key budget.

ConditionWhat happens
RPM exceeded (429)That turn gets an error frame; the connection stays open. Back off and send the next turn later.
Free-tier quota exhausted (402)Connection closes with WS code 4402.
Org no longer exists (404)Connection closes with WS code 4404.
API key revoked mid-sessionConnection closes with code 1011 and reason "API key is no longer valid.".

Next steps​

See Authentication for how X-API-Key is issued and validated, System Prompts for how the system prompt is resolved, and MCP Integration for tool-use behavior — all identical between this endpoint and POST /chat/completions.