Realtime Chat (WebSocket)
WSS /api/llm/chat/realtime
Streaming counterpart to POST /chat/completions: same X-API-Key auth, the same automatic system-prompt/MCP resolution, and the same usage billing, but the assistant's reply comes back token-by-token over a persistent connection instead of as one JSON body once generation finishes.
This page is maintained by hand rather than generated from the OpenAPI spec — OpenAPI has no way to describe a WebSocket operation, so it's absent from openapi.public.json even though the endpoint itself is fully public.
Connecting
wss://api.r-mad.ai/api/llm/chat/realtime
Send X-API-Key as a header on the WebSocket handshake request, exactly as you would for any REST call (see Authentication) — the handshake is an ordinary HTTP request before it's upgraded, so the same header works:
import asyncio
import json
import websockets
async def main():
async with websockets.connect(
"wss://api.r-mad.ai/api/llm/chat/realtime",
additional_headers={"X-API-Key": "YOUR_API_KEY"},
) as ws:
await ws.send(json.dumps({
"type": "message",
"messages": [{"role": "user", "content": "Say hello in one sentence."}],
"return_markdown": False,
}))
async for raw in ws:
event = json.loads(raw)
if event["type"] == "delta":
print(event["text"], end="", flush=True)
elif event["type"] == "done":
print()
break
elif event["type"] == "error":
print(f"\nerror: {event['message']}")
break
asyncio.run(main())
// Node (e.g. the `ws` package) can set the handshake header directly.
// Browsers' native WebSocket API cannot set custom headers on a handshake
// at all — proxy the connection through your own backend if you need this
// from client-side browser code, the same way you would for any other
// credentialed API call you don't want shipped to a browser.
import WebSocket from "ws";
const ws = new WebSocket("wss://api.r-mad.ai/api/llm/chat/realtime", {
headers: {"X-API-Key": "YOUR_API_KEY"},
});
ws.on("open", () => {
ws.send(JSON.stringify({
type: "message",
messages: [{role: "user", content: "Say hello in one sentence."}],
return_markdown: false,
}));
});
ws.on("message", (raw) => {
const event = JSON.parse(raw);
if (event.type === "delta") process.stdout.write(event.text);
else if (event.type === "done") console.log();
else if (event.type === "error") console.error(`\nerror: ${event.message}`);
});
A missing or invalid key closes the connection immediately with WS close code 4401, before any messages are exchanged — same failure verify_partner_api_key would produce as a 401 for the REST endpoint.
Wire protocol
JSON text frames, one exchange per conversation turn. The connection stays open across many turns — you keep sending message frames on the same socket rather than reconnecting each time.
Client → server
{
"type": "message",
"messages": [{"role": "user", "content": "..."}],
"return_markdown": false,
"system_prompt": "optional, layered on top of the org/key default for this turn only",
"user": "optional opaque end-user id, recorded against usage"
}
messages is the full conversation so far, OpenAI chat format — same contract as messages on the REST endpoint. The client owns conversation history; the server doesn't retain it between turns.
return_markdown is required on every turn, same contract as the REST endpoint's return_markdown — true asks the assistant to format its reply as Markdown, false for plain prose. A turn that omits it, or sends a non-boolean, gets an error frame for that turn (the connection stays open) rather than a 400 — a WebSocket has no per-message HTTP status to fail with.
Server → client, in order, per turn:
{"type": "delta", "text": "..."}
Repeated as the reply streams in.
{"type": "done", "usage": {"prompt_tokens": 12, "completion_tokens": 11, "total_tokens": 23}}
Ends the turn. usage is null only if an MCP tool loop hit its iteration cap without producing a priced final answer.
{"type": "error", "message": "..."}
This turn failed — the connection stays open, ready for another message frame.
Unlike the REST endpoint, a turn doesn't accept max_tokens, temperature, or the other sampling fields — every turn runs at a fixed 4096-token cap and temperature 0.2. top_p, stop, frequency_penalty, presence_penalty, and seed aren't available over this endpoint yet.
Rate limits and quota
Enforced per turn, not just at connect — see Rate Limits for the underlying RPM/quota numbers, which are identical to the REST endpoint's and share the same per-key budget.
| Condition | What happens |
|---|---|
RPM exceeded (429) | That turn gets an error frame; the connection stays open. Back off and send the next turn later. |
Free-tier quota exhausted (402) | Connection closes with WS code 4402. |
Org no longer exists (404) | Connection closes with WS code 4404. |
| API key revoked mid-session | Connection closes with code 1011 and reason "API key is no longer valid.". |
Next steps
See Authentication for how X-API-Key is issued and validated, System Prompts for how the system prompt is resolved, and MCP Integration for tool-use behavior — all identical between this endpoint and POST /chat/completions.