AudioHook Streaming Transcription (WebSocket)
WSS /api/audiohook/ws
Implements Genesys Cloud's AudioHook Protocol for a Bot Transcription Connector: Genesys Cloud streams both call legs of a live conversation to this endpoint as they happen, and gets transcript events back in real time. It's the same transcription engine as POST /transcribe, applied to a live call instead of an uploaded file — plus an on-demand intent-classification extension (see below).
This page is maintained by hand rather than generated from the OpenAPI spec, the same reason Realtime Chat is: OpenAPI has no way to describe a WebSocket operation, so this endpoint never appears in openapi.public.json even though it's fully public.
Setting up a connector
Create an AudioHook connector from the dashboard's Genesys page (AudioHook tab). This gives you three things, once:
- API Key — the connector's client ID. Sent as the
X-API-Keyheader on the WebSocket handshake, identifying which org and connector the connection belongs to. - Client Secret — shown once at creation time. Used to sign every handshake request (see Authentication below), never sent on the wire itself.
- WebSocket URL —
wss://api.r-mad.ai/api/audiohook/ws, fixed and org-independent; the org is resolved from the API key, not from the URL.
Paste the URL and credentials into Genesys Cloud's own Bot Transcription Connector configuration. Genesys Cloud's AudioHook client signs and sends every handshake itself from then on — you don't construct requests to this endpoint by hand, the same way you never hand-construct the Genesys Summarization Connector's JWT (see Authentication).
Authentication
Two things travel on the WebSocket upgrade request:
X-API-Key: <client id>— which connector is connecting.Signature/Signature-Input— an RFC 9421 HTTP Message Signature over the request, HMAC-SHA256'd with the connector's client secret. This is what Genesys's AudioHook client generates automatically from the client secret you pasted into its connector config.
A missing X-API-Key, an unknown connector, or a signature that doesn't verify all close the connection immediately with WS code 4401, before any protocol messages are exchanged.
Wire protocol
JSON text frames for control messages, binary frames for audio — the same envelope shape as Genesys's AudioHook reference implementation.
Client → server
{
"version": "2",
"id": "<session id>",
"seq": 1,
"type": "open",
"parameters": {
"language": "en-US",
"media": [
{"type": "audio", "format": "PCMU", "rate": 8000, "channels": ["external", "internal"]}
]
}
}
Only the first audio entry offering PCMU or L16 at 8000 Hz is negotiated — real AudioHook sessions offer exactly one in practice. channels names each leg (typically external = customer, internal = agent); binary audio frames afterwards carry all negotiated channels interleaved sample-by-sample in that same order.
Server → client, in response:
{"type": "opened", "parameters": {"media": [{"type": "audio", "format": "PCMU", "rate": 8000, "channels": ["external", "internal"]}]}}
If nothing offered is supported, or the org's transcription quota is already exhausted, the server sends disconnect instead (see Rate limits and quota below) and closes the connection.
After opened, stream call audio as binary WebSocket frames — raw PCM16 or PCMU, interleaved per the negotiated channels order, no envelope. Each channel is transcribed independently and transcripts come back tagged with the channel they came from:
{
"type": "event",
"parameters": {
"entities": [
{
"type": "transcript",
"data": {
"id": "...",
"channelId": 0,
"isFinal": true,
"offset": "PT4.32S",
"alternatives": [{"confidence": 1.0, "interpretations": [{"type": "display", "transcript": "I'd like to check my balance"}]}]
}
}
]
}
}
Only final transcripts are emitted (isFinal is always true) — there's no interim/partial-result stream on this endpoint.
Other control messages, both directions:
| Client sends | Server responds |
|---|---|
ping | pong |
update (e.g. mid-call language change) | updated |
close | closed, then the socket closes |
paused, resumed, and discarded from the client are accepted and logged but don't change server behavior — there's no barge-in or pause-on-hold handling in this version. The server never sends pause.
On-demand intent classification
Not part of the standard AudioHook protocol — an extension for clients (in practice, the dashboard's own AudioHook playground) that want a customer-intent read on the call so far, on demand rather than continuously:
Client → server
{"type": "analyze", "parameters": {}}
Server → client
{
"type": "intent_update",
"parameters": {
"intent": {
"intents": [{"intent": "balance_inquiry", "confidence": "high"}],
"reasoning": "Customer explicitly asked to check their account balance."
}
}
}
Classifies the full transcript accumulated so far (same model and profile as POST /transcribe's intent extraction). If another analyze arrives while one is already running, it's coalesced into a single follow-up call rather than run concurrently. There's no continuous/streaming intent output, and no BANT/NEAT-P extraction on live transcripts — both are file-upload-only, on POST /transcribe.
Errors and disconnects
| Condition | What happens |
|---|---|
Missing X-API-Key, unknown connector, or bad signature | WS closes immediately with code 4401, no messages exchanged. |
No supported audio format offered in open | disconnect message (reason: "error"), then the socket closes with code 1000. |
Transcription quota exhausted (at open, or mid-call — rechecked every 30s) | disconnect message (reason: "error", info: "Transcription quota exhausted for this plan"), then closes with code 1000. |
| Unhandled server-side error | Socket closes with code 1011. |
Rate limits and quota
There's no per-connection RPM limit — a session is a single long-lived connection, not a burst of requests. Usage is billed in audio-minutes against the same transcription-minutes pool as POST /transcribe (see Rate Limits), flushed incrementally roughly every 30 seconds so a long call doesn't wait until it ends to be metered.
Next steps
See Authentication for how connector credentials are issued, Rate Limits for the shared transcription quota, and POST /transcribe for the batch/file-upload equivalent — including BANT/NEAT-P extraction, which isn't available on live AudioHook transcripts.