Skip to main content
Home›Docs›Architecture›Architecture overview
Architecture

Architecture overview

How Perceive8 captures, understands, and delivers conversation intelligence — and where your systems plug in.

Architecture overview

Perceive8 turns conversations — uploaded recordings, live audio streams, and online meetings — into transcripts, speaker labels, sentiment, topics, alerts, and reports. This page shows how the pieces fit together and where your systems connect. If you only read one page before integrating, read this one.

The big picture

   YOUR SOURCES                    PERCEIVE8                       YOUR STACK
 ┌──────────────────┐      ┌────────────────────────┐      ┌──────────────────────┐
 │ Audio/video files │────▶│  Capture               │      │ REST API             │──▶ your backend
 │ (upload)          │      │  (upload, WS, bot)     │      │ Webhooks             │──▶ your endpoints
 ├──────────────────┤      ├────────────────────────┤      │ SDKs (Python / Node) │──▶ your code
 │ Live audio stream │────▶│  Understand            │─────▶│ MCP server           │──▶ AI assistants
 │ (WebSocket)       │      │  (transcribe, diarize, │      │ CRM sync             │──▶ HubSpot
 ├──────────────────┤      │   analyze, embed)      │      └──────────────────────┘
 │ Meeting bot       │────▶├────────────────────────┤
 │ (Meet, Zoom,      │      │  Store & deliver       │
 │  Teams via        │      │  (transcripts, alerts, │
 │  calendar)        │      │   reports, embeddings) │
 └──────────────────┘      └────────────────────────┘

Three things to notice:

  • Everything is an API. Every capture method and every result is reachable programmatically. The dashboard is built on the same API you use.
  • Your data stays scoped to your workspace. Each API key, webhook, and integration operates inside one workspace, with per-user access control on top. See Access control model.
  • Perceive8 calls specialist AI providers (speech-to-text, diarization, language models) inside the pipeline. You get one integration instead of ten.

The data lifecycle

Every conversation goes through the same four stages, regardless of how it arrives.

1. Capture

Audio enters Perceive8 in one of three ways:

  • File upload — you POST an audio or video file for asynchronous analysis.
  • Live stream — you open a WebSocket and stream audio in real time (for example from a phone system, a browser mic, or a meeting bot).
  • Meeting bot — Perceive8 joins a Google Meet, Zoom, or Teams meeting as a participant and streams the audio for you. See Joining meetings.

2. Understand

The pipeline turns raw audio into structured intelligence:

  • Transcription with word-level timestamps and confidence scores
  • Speaker diarization — who spoke when, labeled across the recording
  • Audio intelligence — sentiment, topics, entities, emotions, and content moderation
  • Scenario analysis — your own questions and scoring criteria applied to the conversation (see Use cases and skills)
  • Embeddings — every transcript segment is vectorized so you can search and ask questions semantically

3. Store

Results are persisted as queryable artifacts: transcripts with segments, analyses, alerts, reports, and embeddings. Everything is retrievable later through the API — see Understanding your analysis.

4. Deliver

Results reach your systems through the delivery channels below. You choose per use case: pull (REST), push (webhooks), live (WebSocket events), or assistant (MCP).

The integration surface

These are all the doors into Perceive8. Each has a dedicated guide.

Integration point Use it when Guide
REST API You upload recordings and fetch results from a backend API quickstart
Streaming WebSocket You need transcripts and alerts while audio is still happening Streaming API
Webhooks You want Perceive8 to notify your servers when work completes Webhooks overview
Python / Node SDKs You prefer typed clients over raw HTTP Python SDK · Node.js SDK
MCP server You want an AI assistant (Claude and others) to query your analyses API quickstart
Calendar & meeting integrations You want automatic meeting capture from Google or Microsoft calendars Connect Google Workspace · Connect Microsoft 365
CRM & calling integrations You want call outcomes synced to HubSpot, or calls recorded via Aircall, RingCentral, or Twilio Available integrations · Connect HubSpot
Aria assistant You want a voice/chat assistant that answers from your conversations Aria assistant

Common integration patterns

Most integrations are one of these four recipes.

Pattern 1 — Async file analysis

Upload a recording, get notified when it's ready, then pull the results.

1. Your backend  →  POST /v1/analyses (audio file)        → returns an analysis id
2. Perceive8     →  runs the pipeline asynchronously
3. Perceive8     →  POST your webhook endpoint            → event: analysis.completed
4. Your backend  →  GET /v1/analyses/{id}                 → transcript, summary, topics

Register the webhook once; no polling loop required. If you can't receive webhooks, polling GET /v1/analyses/{id} until status is completed also works — the API quickstart shows both.

Pattern 2 — Live streaming

Open one WebSocket per live audio source and receive results as they happen.

1. Your app      →  wss://api.perceive8.com/v1/stream?token=...
2. Your app      →  streams PCM16 audio frames
3. Perceive8     →  pushes transcript events and scenario alerts in real time
4. Stream ends   →  a full async analysis is queued automatically for the deep dive

This is the pattern for voice agents, live call coaching, and in-meeting alerts. Details, message formats, and event types: Streaming API.

Pattern 3 — Meeting capture

No audio plumbing on your side at all — the bot handles it.

1. Connect Google or Microsoft calendar  (one-time OAuth)
2. Perceive8  →  joins scheduled meetings automatically 90s before start
3. Perceive8  →  transcribes, analyzes, and scores the meeting
4. Your stack →  receives results via webhook, or they sync to HubSpot

You can also trigger joins on demand via the API or dashboard. Details: Joining meetings.

Pattern 4 — AI assistant access via MCP

Expose your analyses to AI assistants (Claude and other MCP clients) without writing an integration: the Perceive8 MCP server provides tools like list_analyses, get_analysis, query (semantic search over transcripts), and generate_report, secured with OAuth 2.1 and granular scopes. Your assistant users connect once and can ask questions grounded in your actual conversation data.

Events and async delivery

Analysis is asynchronous — anything longer than a few seconds of audio finishes after the HTTP request that started it. You learn about outcomes through events:

Event Meaning Delivery
analysis.completed Pipeline finished; transcript and intelligence available Webhook
analysis.failed Pipeline failed; includes an error message Webhook
Transcript / alert events Live results while streaming WebSocket / SSE

Operational details live in the dedicated guides: Verifying webhook signatures and Retry behavior (five attempts with backoff).

Authentication and data access in 60 seconds

  • API keys (X-API-Key: pk_live_...) — for server-side integrations. Create them under Audio → Settings → Developer. Keys are scoped to a workspace.
  • User JWTs — a Supabase token sent in the Authorization header, for acting on behalf of a signed-in user (for example from your own frontend). Format details: Authentication.
  • WebSocket and SSE clients can't set headers, so they pass the credential as a ?token= query parameter.

Every request is authorized against the workspace, and user-scoped requests additionally respect that user's role. The full model is in Authentication and Access control model. For data handling, retention, and encryption guarantees, see the Security overview.

Where to go next