Streaming API
Perceive8 exposes three streaming interfaces:
- Live WebSocket streaming (
/v1/stream) — real-time transcription, speaker labels, alerts, and conversation scoring over a single WebSocket, served by the streaming service. - Chunked upload sessions (
/v1/analysis/stream-upload) — upload a recording in chunks over HTTP, then finalise it into a standard analysis. - Pipeline progress events — SSE (
GET /v1/analysis/{id}/stream) or WebSocket (/ws/analysis/{id}/events) progress for the post-processing pipeline.
sequenceDiagram
participant C as Client
participant S as Server
C->>S: WebSocket connect
S-->>C: WebSocket open
C->>S: { action: "start", ... }
S-->>C: { event: "session_started" }
loop Audio Streaming
C->>S: PCM audio binary frames
S-->>C: { event: "transcript_partial" }
S-->>C: { event: "transcript_final" }
S-->>C: { event: "alert_triggered" } (if scenarios enabled)
S-->>C: { event: "analysis_chunk" } (if enable_analysis)
end
C->>S: { action: "stop" }
S-->>C: { event: "analysis_started" } (post-stream pipeline)
S-->>C: { event: "session_ended" }
C->>S: WebSocket close
The diagram shows the logical flow; the exact wire message names are documented below.
Live WebSocket Streaming
GET /v1/stream (WebSocket)
Real-time bidirectional streaming endpoint. Authenticate with ?token= (required for browser WebSockets) or the Authorization: Bearer header:
wss://api.perceive8.com/v1/stream?token=pk_live_xxxxxxxxxxxx
Missing or invalid credentials are rejected with HTTP 401 before the upgrade. All JSON messages share an envelope: clients send { "v": 1, "type": ..., "data": ... }; the server adds session_id, user_id, a monotonically increasing seq, and ts:
{
"v": 1,
"type": "transcript",
"session_id": "9b1c3d7e-2f4a-4c5b-8d6e-1a2b3c4d5e6f",
"user_id": "7a8b9c0d-1e2f-3a4b-5c6d-7e8f9a0b1c2d",
"seq": 12,
"ts": "2026-01-15T10:00:01Z",
"data": { "...": "..." }
}
Track the seq of the last event you receive — it is the resume cursor.
Client → server messages
| Message | Format | When | data payload |
|---|---|---|---|
session.init |
JSON text | First frame, within 10 s of connect | See below |
| (audio) | Binary | After session.ack |
Raw PCM bytes (see Send audio) |
ping |
JSON text | Any time | { "nonce": "abc" } — server replies pong with the same nonce |
audio.meta |
JSON text | Optional telemetry | { "client_offset_ms": 3200, "frame_count": 32 } |
transcript.inject |
JSON text | Trusted relays only (e.g. the meeting runtime) | { "speaker_id": "Aria", "text": "Hello there.", "t_start_ms": 12345, "t_end_ms": 14500 } — injects a final transcript segment with an explicit speaker label, for speech that never reached the audio stream (see below) |
session.close |
JSON text | End the session | { "reason": "done" } (reason optional) |
transcript.inject is how trusted relay clients (such as the meeting-runtime adapter) put text the assistant spoke through a separate synthesis path into the transcript, since that audio never reaches the ASR stream. speaker_id must be 1–64 chars and text 1–4000 chars; t_start_ms/t_end_ms are optional non-negative milliseconds since session start (t_end_ms defaults to the current session time; a missing or out-of-order t_start_ms collapses to t_end_ms, a zero-duration segment). The server emits it as a final transcript event like any ASR final. Invalid payloads — and injects arriving after session.close — are dropped with a warning; the session is unaffected. Injected segments are persisted and replayed like ASR finals but are excluded from the alert engine and live conversation summaries, and injected text is not PII-redacted server-side (ASR redaction happens provider-side; the injector is a trusted relay sending the assistant's own synthesized speech).
session.init payload:
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
mode |
"new" | "resume" |
✓ | — | Use "resume" to reattach to an interrupted session |
scenario_id |
UUID | ✓ | — | Scenario whose questions the alert engine evaluates. A scenario with no questions disables alerts. |
audio |
object | ✓ | — | Declared format, e.g. { "encoding": "pcm_s16le", "sample_rate": 16000, "channels": 1 }. The server always negotiates PCM s16le / 16 kHz / mono (see negotiated_audio in the ack). |
resume |
object | with mode: "resume" |
— | { "session_id", "resume_token", "last_seq" } |
quality_tier |
"lite" | "pro" | "max" |
— | "pro" |
Usage multiplier 1.0 / 1.5 / 3.0; selects ASR model and feature set |
retention |
"1w" | "1m" | "6m" | "2y" |
— | "1m" |
Audio storage retention class |
language |
string | — | — | Language code stored on the session record |
pii |
object | — | — | Accepted, but PII redaction is enforced server-side regardless (see below) |
client_meta |
object | — | — | Arbitrary metadata; client_meta.workspace_id attributes the session to a workspace |
PII redaction is always on: phone_number, email_address, credit_card_number, us_social_security_number, and location entities are substituted in transcripts.
session.ack
The first server message (seq: 1) confirms the session:
{
"v": 1,
"type": "session.ack",
"session_id": "9b1c3d7e-2f4a-4c5b-8d6e-1a2b3c4d5e6f",
"seq": 1,
"ts": "2026-01-15T10:00:00Z",
"data": {
"accepted": true,
"resume_token": "rt_9f2c...hex",
"scenario": { "id": "550e8400-e29b-41d4-a716-446655440000", "question_count": 7 },
"quality_tier": "pro",
"usage_multiplier": 1.5,
"enabled_features": ["transcription", "diarization", "pii_redaction", "sentiment", "entity", "topic"],
"negotiated_audio": { "encoding": "pcm_s16le", "sample_rate": 16000, "channels": 1 },
"limits": { "max_audio_ms_buffer": 30000, "min_llm_interval_ms": 4000 }
}
}
Store session_id (envelope field) and resume_token — you need both, plus your last received seq, to resume. resumed_from_seq is present only on resumed sessions.
Send audio
After session.ack, send audio as binary WebSocket frames:
- Encoding: 16-bit signed PCM, little-endian (
pcm_s16le); 16000 Hz; mono — the values the server negotiates - Frames over 256,000 bytes are dropped; 100 ms frames (3,200 bytes) work well
Server → client events
All events use the envelope shown above. data payloads:
type |
When sent | data fields |
|---|---|---|
session.ack |
Once, after a valid session.init |
Shown above |
transcript |
Each ASR partial and final turn | segment_id, revision, is_final, speaker_id?, t_start_ms, t_end_ms, text, words?, confidence? |
speaker_revision |
Diarization relabels earlier segments | revisions[] (segment_id, old_speaker_id, new_speaker_id, revision), reason |
alert |
A scenario question fires | alert_id, question_id, question_text, answer, severity, evidence_text, evidence_segment_ids, t_start_ms, t_end_ms, rationale, recommendation, dedup_hash |
conversation.summary |
When scores change (min 2 s apart) and once at close | total_score, sub_scores (compliance, talk_ratio, alert_load), suggested_next_line, severity_counts (high, med, low) |
flow_control |
ASR backpressure state changes | state ("ok" | "throttled"), reason ("asr_backpressure" | "asr_recovered"), suggested_delay_ms? |
session.status |
Lifecycle transitions | status ("active" | "closed" | "error"), detail? |
pong |
Reply to client ping |
nonce |
error |
Failures (see below) | code, message?, fatal, retryable |
Final transcript example (partials omit words and have is_final: false):
{
"v": 1,
"type": "transcript",
"session_id": "9b1c3d7e-2f4a-4c5b-8d6e-1a2b3c4d5e6f",
"seq": 12,
"ts": "2026-01-15T10:00:01Z",
"data": {
"segment_id": "3fa85f64-5717-4562-b3fc-2c963f66afa6",
"revision": 0,
"is_final": true,
"speaker_id": "A",
"t_start_ms": 1200,
"t_end_ms": 3600,
"text": "Hello, how are you today?",
"words": [
{ "w": "Hello", "t_start_ms": 1200, "t_end_ms": 1500, "conf": 0.99 }
],
"confidence": 0.97
}
}
Alert severity is taken from the scenario question's configured level, or derived from the LLM confidence: INFO (< 0.5), WARNING (≥ 0.5), HIGH (≥ 0.75), CRITICAL (≥ 0.9).
error codes include asr_not_configured (fatal), asr_upstream_unavailable and asr_upstream_error (retryable), and internal (fatal). A fatal error ends the session; watch for the following session.status.
Reconnect and resume
If the connection drops, reconnect within ~60 seconds and send session.init with "mode": "resume" and resume: { "session_id", "resume_token", "last_seq" }. The server validates the token, replays retained events with seq > last_seq (up to 500), and sets resumed_from_seq in the ack. If your cursor is older than the retained window, you instead get session.status with detail: "resume_gap" and the session continues live. Resume failures close the socket with code 4003.
Session limits and close behaviour
- Idle: the server ends the session after 30 s without audio frames; maximum session duration is 4 hours.
- Client-initiated end (
session.closeor socket close) yieldssession.status: "closed"; server-side termination (idle, max duration, upstream failure) yields"error". A finalconversation.summaryprecedessession.status. - Close codes:
4001invalid or latesession.init;4003resume failure. Auth failures are HTTP401before the upgrade.
Example (browser JavaScript)
const ws = new WebSocket("wss://api.perceive8.com/v1/stream?token=pk_live_xxxxxxxxxxxx");
ws.onopen = () => {
ws.send(JSON.stringify({
v: 1, type: "session.init",
data: {
mode: "new",
scenario_id: "550e8400-e29b-41d4-a716-446655440000",
audio: { encoding: "pcm_s16le", sample_rate: 16000, channels: 1 },
},
}));
};
ws.onmessage = (msg) => {
const env = JSON.parse(msg.data);
if (env.type === "transcript" && env.data.is_final) console.log(env.data.text);
if (env.type === "alert") console.warn(env.data.severity, env.data.question_text);
};
// After session.ack, stream mic audio as Int16LE mono 16 kHz binary frames:
processor.onaudioprocess = (e) => {
const f32 = e.inputBuffer.getChannelData(0);
const i16 = new Int16Array(f32.length);
for (let i = 0; i < f32.length; i++) i16[i] = Math.max(-1, Math.min(1, f32[i])) * 32767;
ws.send(i16.buffer);
};
// End: ws.send(JSON.stringify({ v: 1, type: "session.close", data: { reason: "done" } }));
sequenceDiagram
participant C as Client
participant S as Server
participant P as Pipeline
C->>S: 1. Open wss://api.perceive8.com/v1/stream?token=...
C->>S: 2. Send { action: "start", language: "en", sample_rate: 16000 }
S-->>C: 3. { event: "session_started", session_id: "..." }
loop 4. Stream PCM frames
C->>S: Binary PCM audio frames
S-->>C: transcript_partial
S-->>C: transcript_final
S-->>C: analysis_chunk (if enabled)
S-->>C: alert_triggered (if scenarios)
end
C->>S: 5. Send { action: "stop" }
S->>P: Start post-stream pipeline
S-->>C: 6. { event: "analysis_started", analysis_id: "..." }
S-->>C: 7. { event: "session_ended", analysis_id: "..." }
C->>S: 8. Close WebSocket
Note over C,P: 9. Optional: Poll /v1/analyses/{id} or subscribe to SSE progress
SSE Pipeline Progress
GET /v1/analysis/{analysis_id}/stream
Server-Sent Events stream of post-processing pipeline progress for one analysis. Auth: X-API-Key or Authorization: Bearer header — this endpoint does not accept ?token=, so browser EventSource (which cannot set headers) cannot call it directly; use the WebSocket variant below or a fetch-based SSE reader. Returns 404 if the analysis is not in the authenticated workspace.
| Parameter | In | Type | Description |
|---|---|---|---|
analysis_id |
path | UUID | Analysis to track |
| Event | When sent | Data |
|---|---|---|
ping |
Immediately on connect, then every 15 s (heartbeat) | {} |
pipeline_step |
Each pipeline step update | { "analysis_id", "step", "status", "progress_pct", "message", "timestamp" } |
done |
Terminal: progress_pct reaches 100 or status is failed; connection closes afterwards |
{ "status": "completed" | "failed", "analysis_id", "error"? } |
curl -N https://api.perceive8.com/v1/analysis/9b1c3d7e-2f4a-4c5b-8d6e-1a2b3c4d5e6f/stream \
-H "X-API-Key: pk_live_xxxxxxxxxxxx"
event: ping
data: {}
event: pipeline_step
data: {"analysis_id":"9b1c3d7e-...","step":"store","status":"completed","progress_pct":5,"message":"Stored audio file","timestamp":"2026-01-15T10:00:01.123456"}
event: pipeline_step
data: {"analysis_id":"9b1c3d7e-...","step":"complete","status":"completed","progress_pct":100,"message":"Analysis complete","timestamp":"2026-01-15T10:01:35.654321"}
event: done
data: {"status":"completed","analysis_id":"9b1c3d7e-2f4a-4c5b-8d6e-1a2b3c4d5e6f"}
Pipeline steps and their progress_pct values:
| Step | progress_pct |
Step | progress_pct |
|---|---|---|---|
store |
5 | match_speakers |
65 |
preprocess |
10 | merge |
75 |
quality_check |
15 | audio_intelligence |
80 |
enhance |
20 | video_analysis |
84 |
diarize |
40 | embed |
88 |
transcribe |
55 | persist |
95 |
complete |
100 |
WebSocket Pipeline Progress
GET /ws/analysis/{analysis_id}/events (WebSocket)
The same pipeline progress stream over a WebSocket, for clients that cannot set headers on SSE. Auth: ?token= query parameter (JWT or pk_live_ API key).
wss://api.perceive8.com/ws/analysis/9b1c3d7e-2f4a-4c5b-8d6e-1a2b3c4d5e6f/events?token=pk_live_xxxxxxxxxxxx
Connection failures close the socket: 4001 missing/invalid token, 4004 analysis not found, 4500 internal error. The client may send { "action": "cancel" } to close the stream.
| Server message | When sent | Payload |
|---|---|---|
{ "event": "ping" } |
Every 15 s (heartbeat) | — |
{ "event": "pipeline_step", "data": {...} } |
Each pipeline event | Same payload as the SSE pipeline_step data |
{ "event": "done", "data": {...} } |
Completion or failure | { "status": "completed" | "failed", "analysis_id", "error"? } |
Chunked Upload Sessions
For pre-recorded audio too large or unreliable to send in one request: create a session, upload chunks, then finalise into an analysis. All endpoints use standard workspace auth (X-API-Key or Authorization: Bearer) and are scoped to the authenticated workspace.
POST /v1/analysis/stream-upload
Create an upload session. The JSON body is optional.
Request body
{
"language": "en",
"filename": "call-recording.mp3"
}
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
language |
string | — | "en" |
Language code for the analysis |
filename |
string | — | — | Original filename; its extension is kept for the stored audio |
Response 200 — the session expires after 4 hours (default):
{
"session_id": "5f1a2b3c-4d5e-6f7a-8b9c-0d1e2f3a4b5c",
"expires_at": "2026-01-15T14:00:00+00:00"
}
curl -X POST https://api.perceive8.com/v1/analysis/stream-upload \
-H "X-API-Key: pk_live_xxxxxxxxxxxx" \
-H "Content-Type: application/json" \
-d '{"language": "en", "filename": "call-recording.mp3"}'
POST /v1/analysis/stream-upload/{session_id}/chunk
Upload one chunk. Send raw bytes with Content-Type: application/octet-stream, or multipart/form-data with an audio_file field.
| Parameter | In | Type | Description |
|---|---|---|---|
session_id |
path | UUID | Upload session |
The first chunk is validated against an audio allowlist (WAV, MP3, MP4/M4A, OGG, FLAC, AAC, AIFF, WebM) by MIME type and magic bytes. Chunks larger than 512 KB (default) are rejected.
Response 200
{ "chunk_index": 0, "total_bytes": 262144 }
Errors: 400 empty chunk or wrong content type · 403 session owned by another workspace · 404 unknown session · 409 session not open · 410 session expired · 413 chunk too large · 415 not a recognised audio file.
curl -X POST https://api.perceive8.com/v1/analysis/stream-upload/5f1a2b3c-.../chunk \
-H "X-API-Key: pk_live_xxxxxxxxxxxx" \
-H "Content-Type: application/octet-stream" \
--data-binary @chunk-001.bin
GET /v1/analysis/stream-upload/{session_id}/stream (SSE)
SSE stream for an upload session. Same path parameters and auth as the chunk endpoint.
| Event | When sent | Data |
|---|---|---|
(comment) : ping |
Heartbeat | — |
segment |
A transcript segment is published for the session | Segment payload (JSON string) |
done |
Completion signal, or the 5-minute stream cap is reached | Plain text: completed or timeout |
POST /v1/analysis/stream-upload/{session_id}/finalize
Combine the buffered chunks, store the audio, create the Analysis, and dispatch the async pipeline. The session transitions to processing.
Response 202
{ "analysis_id": "9b1c3d7e-2f4a-4c5b-8d6e-1a2b3c4d5e6f", "status": "processing" }
Errors: 400 no audio accumulated · 404 no buffer (no chunks uploaded) · 409 session not open (plus 403 / 410 as above).
Track the resulting pipeline with GET /v1/analysis/{analysis_id}/stream (SSE) or /ws/analysis/{analysis_id}/events (WebSocket) above.