Skip to main content
Home›Docs›API Reference›Analyses
API Reference

Analyses

Upload audio or video for transcription, speaker diarization, and audio intelligence; manage analyses and speaker voiceprints.

Analyses

An analysis is an uploaded audio or video file that Perceive8 processes into a transcript with speaker labels, signals (sentiment, topics, entities, key phrases), and alerts. You can create analyses synchronously or asynchronously, list and inspect them, re-run or resume processing, and manage per-user stars. All endpoints are scoped to a workspace: send the X-Workspace-Id header to target a specific workspace, otherwise your Personal workspace is used.

The analysis router is reachable under both /v1/analyses and /v1/analysis; the examples below use /v1/analyses.

Common enums and defaults

Value Options Default
language auto, en, he, es en
diarization_provider pyannote, replicate, assemblyai pyannote
transcription_providers openai_whisper, replicate, assemblyai ["openai_whisper"]
tier lite, pro, max pro
Analysis status pending, processing, failed, completed —

audio_intelligence_features is a JSON object of feature flags; omitted flags keep their defaults: sentiment_analysis (true), entity_detection (true), topic_detection (true), auto_chapters (true), content_moderation (true), pii_redaction (false), key_phrases (true), speaker_labels (true).

Create an analysis

POST /v1/analyses/analyze-upload-async

Upload an audio or video file (multipart/form-data) and process it in the background. Returns 202 immediately; poll GET /v1/analyses/{analysis_id} for progress. Recommended way to submit files.

Form fields

Field Type Required Description
audio_file file yes Audio (max 200 MB) or video (max 400 MB). MIME type and magic bytes are validated.
language string no See enum table above.
diarization_provider string no See enum table above.
transcription_providers string[] no One or more providers; each produces its own processing run.
playbook_id string (UUID) no Override the playbook used; defaults to your active playbook.
video_modules string no Video only: JSON array of modules, e.g. '["emotion","scene","pose"]'. Defaults to ["emotion","scene"].
audio_intelligence_features string no JSON object of feature flags (see above).
num_speakers int no Accepted but not currently forwarded to the pipeline.
curl -X POST https://api.perceive8.com/v1/analyses/analyze-upload-async \
  -H "X-API-Key: pk_live_xxxxxxxxxxxx" \
  -F "[email protected]" \
  -F "language=en" \
  -F 'audio_intelligence_features={"sentiment_analysis":true,"pii_redaction":false}'

Response 202 (audio upload; a video upload also returns "type": "video" and video_modules):

{ "analysis_id": "3fa85f64-5717-4562-b3fc-2c963f66afa6", "status": "pending" }

Errors: 400 empty file, 402 usage limit exceeded, 413 file too large, 415 unsupported or invalid media type.

POST /v1/analyses/analyze-async

Submit a file that is already in storage for background processing. Form-encoded; no file body.

Form fields

Field Type Required Description
storage_path string yes Path of a previously uploaded file, matching {user_id}/{uuid}/{filename}.
language string no See enum table above.
provider string no Legacy single transcription provider (default assemblyai). Ignored when transcription_providers is set.
transcription_providers string[] no See enum table above.
diarization_provider string no See enum table above.
playbook_id string (UUID) no Must belong to you; 404 otherwise.
tier string no lite, pro, or max (default pro).
audio_intelligence_features string no JSON object of feature flags.
num_speakers, diarize, redact_pii int / bool / bool no Accepted but not currently forwarded to the pipeline; use audio_intelligence_features instead.

Response 202:

{ "analysis_id": "3fa85f64-5717-4562-b3fc-2c963f66afa6", "status": "pending" }

Errors: 400 invalid storage_path or playbook_id, 402 usage limit exceeded, 404 playbook not found.

POST /v1/analyses

Upload audio (multipart/form-data, field audio_file) and process it synchronously — the request blocks until diarization and transcription finish. Accepts the same audio form fields as the async upload endpoint, plus diarization_model and transcription_models (string / string[]) to pin specific provider models.

Response 200:

{
  "id": "3fa85f64-5717-4562-b3fc-2c963f66afa6",
  "language": "en",
  "created_at": "2026-01-15T10:30:00.000000",
  "audio_intelligence_features": { "sentiment_analysis": true, "pii_redaction": false }
}

POST /v1/analyses/analyze

Like POST /v1/analyses, but runs the full pipeline synchronously (including embeddings and, when you have an active playbook, automatic report generation). Same form fields and response shape; a pipeline failure returns 502.

List and inspect analyses

GET /v1/analyses

List analyses in the active workspace, newest first. Query parameters: limit (int, default 100) and offset (int, default 0).

Response 200:

{
  "analyses": [
    {
      "id": "3fa85f64-5717-4562-b3fc-2c963f66afa6",
      "name": "Q4 discovery call",
      "language": "en",
      "status": "completed",
      "pipeline_checkpoint": null,
      "created_at": "2026-01-15T10:30:00.000000",
      "metadata": null
    }
  ],
  "total": 1
}

GET /v1/analyses/{analysis_id}

Get a single analysis with its processing runs. Includes video_analysis (job status, segment counts) when the analysis is a video, and live_summary when it originated from a live stream; both are null otherwise. Returns 404 when not found.

Response 200:

{
  "id": "3fa85f64-5717-4562-b3fc-2c963f66afa6",
  "name": "Q4 discovery call",
  "language": "en",
  "status": "completed",
  "pipeline_checkpoint": null,
  "error_message": null,
  "storage_path": "user-id/3fa85f64-5717-4562-b3fc-2c963f66afa6/call-recording.mp3",
  "video_modules": null,
  "video_analysis": null,
  "audio_intelligence_features": { "sentiment_analysis": true },
  "key_phrases": ["pricing", "renewal"],
  "metadata": null,
  "stream_session_id": null,
  "live_summary": null,
  "created_at": "2026-01-15T10:30:00.000000",
  "runs": [
    { "id": "8b1a9953-c461-4c5c-9b3f-3d2f4c5b6a01", "run_type": "transcription", "provider_name": "assemblyai", "model_name": null, "status": "completed", "processing_time_seconds": 42.5 }
  ]
}

PATCH /v1/analyses/{analysis_id}

Rename an analysis (or clear its name). Requires the member role on the analysis's use case. Body: { "name": "Q4 discovery call" } — name is capped at 255 characters (longer values are truncated); an empty or whitespace-only string clears it to null.

Response 200:

{ "name": "Q4 discovery call" }

Starred analyses

Stars are per user, within the active workspace.

POST /v1/analyses/{analysis_id}/star

Star an analysis. Idempotent — starring an already-starred analysis succeeds.

Response 201:

{ "starred": true, "analysis_id": "3fa85f64-5717-4562-b3fc-2c963f66afa6" }

DELETE /v1/analyses/{analysis_id}/star

Remove the star from an analysis. Returns 200 with { "starred": false, "analysis_id": "..." }.

GET /v1/analyses/starred

List your starred analyses, most recently starred first. Accepts the same limit/offset query parameters and returns the same shape as GET /v1/analyses.

GET /v1/analyses/starred/ids

Return only the IDs of your starred analyses (for UI state).

Response 200:

{ "starred_ids": ["3fa85f64-5717-4562-b3fc-2c963f66afa6"] }

Reprocess an analysis

POST /v1/analyses/{analysis_id}/reanalyze

Re-run the pipeline from scratch: deletes all existing processing runs and their segments, resets the status to pending, and queues a new run using the stored providers. Requires member role on the analysis's use case. Errors: 400 the analysis has no storage_path, 404 not found.

Response 202:

{ "analysis_id": "3fa85f64-5717-4562-b3fc-2c963f66afa6", "status": "pending" }

POST /v1/analyses/{analysis_id}/retry

Resume a stuck (processing) or failed analysis from its last pipeline checkpoint, preserving completed work. Analyses in any other status are rejected with 400. Requires member role.

Response 202:

{
  "analysis_id": "3fa85f64-5717-4562-b3fc-2c963f66afa6",
  "status": "pending",
  "pipeline_checkpoint": "transcription",
  "message": "Analysis re-queued. It will resume from the last checkpoint."
}

Processing runs

GET /v1/analyses/{analysis_id}/runs

List all processing runs for an analysis — one per pipeline execution (diarization, transcription), each with its provider, model, status, and duration.

Response 200:

{
  "runs": [
    { "id": "8b1a9953-c461-4c5c-9b3f-3d2f4c5b6a01", "run_type": "diarization", "provider_name": "pyannote", "model_name": null, "status": "completed", "processing_time_seconds": 18.2 }
  ]
}

GET /v1/analyses/{analysis_id}/runs/{run_id}

Get one run with its segments. The segment shape depends on run_type: diarization runs return speaker_label, start_time, end_time, confidence; other runs return transcript segments with start_time, end_time, text, confidence, word_timestamps. Errors: 404 analysis or run not found.

Response 200 (transcript run):

{
  "id": "8b1a9953-c461-4c5c-9b3f-3d2f4c5b6a01",
  "run_type": "transcription",
  "provider_name": "assemblyai",
  "model_name": null,
  "status": "completed",
  "processing_time_seconds": 42.5,
  "error_message": null,
  "segments": [
    { "start_time": 0.0, "end_time": 3.84, "text": "Thanks for joining the call today.", "confidence": 0.97, "word_timestamps": null }
  ]
}

Audio intelligence

GET /v1/analyses/{analysis_id}/audio-intelligence

Return the audio-intelligence segments for an analysis (sentiment, emotions, topics, entities, content moderation), ordered by start time. Each segment has: start_time, end_time, text, speaker_label, provider_name, prosody_emotions, vocal_burst, sentiment_label, sentiment_confidence, topics, entities, content_moderation, contextual_analysis.

Response 200:

{
  "analysis_id": "3fa85f64-5717-4562-b3fc-2c963f66afa6",
  "audio_intelligence_features": { "sentiment_analysis": true },
  "key_phrases": ["pricing", "renewal"],
  "segments": [
    {
      "start_time": 0.0,
      "end_time": 3.84,
      "text": "Thanks for joining the call today.",
      "speaker_label": "SPEAKER_00",
      "provider_name": "assemblyai",
      "prosody_emotions": null,
      "vocal_burst": null,
      "sentiment_label": "positive",
      "sentiment_confidence": 0.91,
      "topics": null,
      "entities": null,
      "content_moderation": null,
      "contextual_analysis": null
    }
  ],
  "total": 1
}

POST /v1/analyses/{analysis_id}/extract-topics

Extract topics and entities from the transcript with an LLM. Results are cached on the analysis's intelligence segments: if topics were already extracted, the cached values are returned without another LLM call. Requires member role.

Response 200 — source is "cached" when existing results were returned, "llm" when a new extraction ran:

{
  "analysis_id": "3fa85f64-5717-4562-b3fc-2c963f66afa6",
  "source": "llm",
  "topics": [{ "label": "pricing", "count": 1 }],
  "entities": [{ "label": "Acme Corp", "count": 1 }]
}

Errors: 404 not found, 422 no usable transcript, 502 LLM call failed, 503 LLM not configured.

Speakers (voiceprints)

Speakers are enrolled voiceprints used to identify people by name in transcripts. All speaker endpoints require a dashboard permission on one of the sales, support, or cs dashboards: viewer for reads, member for writes — except deletion, which requires the workspace admin role. Enrollment additionally requires voiceprint consent in your privacy settings and a paid plan: Free-plan workspaces cannot enroll speakers, and paid plans are capped at their speaker limit.

POST /v1/speakers

Enroll a speaker from a voice sample (multipart/form-data). The sample is preprocessed, embedded, and stored for matching. Form fields: name (string, required) and voice_sample (file, required).

curl -X POST https://api.perceive8.com/v1/speakers \
  -H "X-API-Key: pk_live_xxxxxxxxxxxx" \
  -F "name=Jane Doe" \
  -F "[email protected]"

Response 200:

{
  "id": "5d2f4c5b-6a01-4b2c-8e3f-9a1b2c3d4e5f",
  "name": "Jane Doe",
  "user_id": "auth0|abc123",
  "created_at": "2026-01-15T10:30:00.000000"
}

Errors: 403 voiceprint consent disabled, Free plan, or plan speaker limit reached; 502 voiceprint extraction failed; 503 embedding service unavailable.

GET /v1/speakers

List speakers in the workspace. Supports limit (default 100) and offset (default 0) query parameters.

Response 200:

{
  "speakers": [
    { "id": "5d2f4c5b-6a01-4b2c-8e3f-9a1b2c3d4e5f", "name": "Jane Doe", "created_at": "2026-01-15T10:30:00.000000" }
  ],
  "total": 1
}

GET /v1/speakers/{speaker_id}

Get a single speaker, including chromadb_id (the voiceprint store reference).

Response 200:

{
  "id": "5d2f4c5b-6a01-4b2c-8e3f-9a1b2c3d4e5f",
  "name": "Jane Doe",
  "chromadb_id": "5d2f4c5b-6a01-4b2c-8e3f-9a1b2c3d4e5f",
  "created_at": "2026-01-15T10:30:00.000000"
}

PATCH /v1/speakers/{speaker_id}

Rename a speaker. Requires member permission. Body: { "name": "Jane D." } — a blank name is rejected with 422.

Response 200:

{ "id": "5d2f4c5b-6a01-4b2c-8e3f-9a1b2c3d4e5f", "name": "Jane D.", "created_at": "2026-01-15T10:30:00.000000" }

DELETE /v1/speakers/{speaker_id}

Delete a speaker and its voiceprint. Existing transcript and diarization segments keep their text but lose the link to this speaker. Requires the workspace admin role.

Response 200:

{ "message": "Speaker deleted", "id": "5d2f4c5b-6a01-4b2c-8e3f-9a1b2c3d4e5f" }

See also

  • Streaming API — live WebSocket transcription and SSE pipeline progress.
  • Reports — playbook reports generated from completed analyses.