Realtime Voice Chat

Build bidirectional voice conversations with AI in React. Realtime audio streaming, interruption handling, and visual state, integrated via assistant-ui.

assistant-ui supports realtime bidirectional voice through the RealtimeVoiceAdapter interface. Users speak into their microphone, the agent answers with audio, and transcripts appear in the thread while the session runs.

idle

Three voice modes

assistant-ui ships three related but distinct voice capabilities. Choose the mode that matches the product surface:

ModeGuideAdapterDirection
Realtime duplexthis pageRealtimeVoiceAdapterAudio ↔ audio, live session
Push-to-talk dictationDictationDictationAdapterAudio → text into the composer
Read-aloudSpeechSpeechSynthesisAdapterText → audio for one message

Realtime voice is the only mode that owns both directions at once. Dictation fills the composer; speech reads a finished message aloud.

RealtimeVoiceAdapter

Implement RealtimeVoiceAdapter to connect any voice provider. The contract lives in @assistant-ui/core and is re-exported from @assistant-ui/react:

import type { RealtimeVoiceAdapter } from "@assistant-ui/react";

type RealtimeVoiceAdapter = {
  connect: (options: {
    abortSignal?: AbortSignal;
  }) => RealtimeVoiceAdapter.Session;
};

A session exposes connection state, mute controls, and event subscriptions:

namespace RealtimeVoiceAdapter {
  type Status =
    | { type: "starting" | "running" }
    | {
        type: "ended";
        reason: "finished" | "cancelled" | "error";
        error?: unknown;
      };

  type Mode = "listening" | "speaking";

  type TranscriptItem = {
    role: "user" | "assistant";
    text: string;
    isFinal?: boolean;
  };

  type Session = {
    status: Status;
    isMuted: boolean;

    disconnect: () => void;
    mute: () => void;
    unmute: () => void;
    sendText?: (text: string) => void | Promise<void>;

    onStatusChange: (callback: (status: Status) => void) => Unsubscribe;
    onTranscript: (
      callback: (transcript: TranscriptItem) => void,
    ) => Unsubscribe;
    onModeChange: (callback: (mode: Mode) => void) => Unsubscribe;
    onVolumeChange: (callback: (volume: number) => void) => Unsubscribe;
  };
}

Session status moves starting → running → ended. The ended status includes a reason: "finished", "cancelled", or "error" (with an optional error field).

Transcripts from onTranscript become messages in the thread:

  • User transcripts (role: "user", isFinal: true) become user messages.
  • Assistant transcripts (role: "assistant") stream into an assistant message. The message stays in a running status until isFinal: true.
  • Both carry metadata.modality: "voice", so your components can tell a spoken message from a typed one.

A finalized transcript is an ordinary message from then on. The local runtime writes it to the repository and, when one is configured, the history adapter as it finalizes. It persists across reloads only with a history adapter, and reaches the model on the next run like a typed turn. An external store runtime hands each finalized transcript to onVoiceTranscript; a host that does not implement it keeps transcripts for the session only, and disconnectVoice drops them. A transcript that finalizes while thread history is still loading stays in the session until the load settles and is committed after the loaded messages, whether the load began before or after the session connected; the local runtime writes it to the repository and history then, and an external store receives it through onVoiceTranscript once isLoading is false. A typed turn appears in the thread right away, and its append promise resolves when the commit lands, so it stays pending for the rest of the load. Switching threads before the load ends drops the waiting message rather than committing it to the thread the user moved to.

The local runtime and runtimes built on it, such as @assistant-ui/react-data-stream, commit transcripts as described above. On the external store runtimes that accept a voice adapter, where a finalized transcript goes depends on the runtime. Session only means the transcripts render while the session is connected, never reach the provider thread, a history adapter, or the model, and are dropped on disconnect.

RuntimeFinalized transcripts
@assistant-ui/ai-sdkAppended to the useChat messages, so they reach the model with the next request and a configured history adapter stores them like typed turns.
@assistant-ui/eveSession only. Every write to an eve session starts a turn, so a transcript cannot be recorded without the agent answering it again.
@assistant-ui/react-a2aAdded to the thread, and a configured history adapter stores them like typed turns. The agent does not receive them: A2A has no way to add a message to a context without asking the agent to act on it.
@assistant-ui/react-ag-uiAdded to the thread, so they reach the agent with the next run's messages and a configured history adapter stores them like typed turns.
@assistant-ui/react-google-adkSession only. The ADK API server adds content to an existing session only by running the agent, one user message per run.
@assistant-ui/react-langchainAdded to the thread and sent to the graph with the next run's input, which writes them to the thread state. A page reload before that run loses them.
@assistant-ui/react-langgraphAdded to the thread and sent to the graph with the next run's input, which writes them to the thread state. A page reload before that run loses them.
@assistant-ui/react-opencodeSession only. OpenCode cannot record an assistant message, so the session would keep only the user's half of the conversation.
@assistant-ui/react-piSession only. Pi records a message without prompting only as a custom message, which the model reads as the user's.

The Thread element renders voice turns as compact spoken rows grouped into a voice conversation block, and VoiceConversation from voice-conversation.aui renders the live call screen.

Typed text during a session

A session that implements sendText takes typed text while it is running: thread.append with a plain text user message at the end of the thread hands the text to sendText and, once it resolves, commits the message as an ordinary typed turn (no metadata.modality) through the same path as a finalized transcript, so the local runtime writes it to the repository and history and an external store receives it through onVoiceTranscript. The session must not echo the typed text back through onTranscript; the runtime records the turn exactly once. A rejected sendText, or a session that ends before the text is recorded, rejects append with MessageNotSentError (the provider error as its cause), commits nothing, and hands the draft back to the composer. If the thread runtime is replaced or unmounted while the text is in flight, append resolves instead and commits nothing, even when the send succeeded. The typed turn carries the composer metadata a text send would. VoiceSessionState.canSendText is true while a running session takes typed text, and the thread composer's canSend follows it; the send button and Enter key follow canSend alone during a session, so a reply being spoken does not lock them. Attachments, edits, and messages with non-text parts are still rejected during the session.

A session without sendText keeps the composer closed: canSend is false and thread.append rejects with "Cannot send a text message while a voice session is connected". Edits, reloads, and runs are rejected while connected on both runtimes whether or not the session takes typed text. Tool results and approvals are rejected on the local runtime; an external store forwards them to the host, which decides. Hanging up commits the sentence the assistant was speaking. Discarding the thread's runtime also ends the session: an external store switching to another thread id, or a thread list stopping or unmounting the thread, disconnects the provider and drops the sentence in progress, while a thread list remounting the thread for threads.reloadMainThread() hangs up and commits that sentence unless the thread's history is still loading. A queue on the local runtime holds its items while a session is connected and resumes after disconnect; a queue owned by an external store host is the host's to pause. The rule holds in both directions: connectVoice rejects while a run is in progress or paused on a pending tool action, and the kit disables the connect button until the run settles.

onModeChange reports "listening" (user's turn) or "speaking" (agent's turn). onVolumeChange reports a real-time level from 0 to 1 for visual feedback such as the voice orb.

createVoiceSession

createVoiceSession removes the manual callback-set boilerplate when you implement an adapter. Pass it the connect options and an async setup function that receives VoiceSessionHelpers and returns VoiceSessionControls:

import {
  createVoiceSession,
  type RealtimeVoiceAdapter,
  type VoiceSessionControls,
  type VoiceSessionHelpers,
} from "@assistant-ui/react";

export class MyVoiceAdapter implements RealtimeVoiceAdapter {
  connect(options: {
    abortSignal?: AbortSignal;
  }): RealtimeVoiceAdapter.Session {
    return createVoiceSession(options, async (helpers: VoiceSessionHelpers) => {
      const client = await MyVoiceClient.connect();

      client.on("open", () => helpers.setStatus({ type: "running" }));
      client.on("close", () => helpers.end("finished"));
      client.on("error", (err: unknown) => helpers.end("error", err));
      client.on("transcript", (item) => helpers.emitTranscript(item));
      client.on("mode", (mode) => helpers.emitMode(mode));
      client.on("volume", (v: number) => helpers.emitVolume(v));

      const controls: VoiceSessionControls = {
        disconnect: () => client.close(),
        mute: () => client.setMuted(true),
        unmute: () => client.setMuted(false),
        sendText: (text) => client.sendText(text),
      };
      return controls;
    });
  }
}

VoiceSessionHelpers provides:

HelperRole
setStatus(status)Update session status (for example to { type: "running" })
end(reason, error?)End the session and clean up subscribers
emitTranscript(item)Push a transcript into the thread
emitMode(mode)Report "listening" or "speaking"
emitVolume(volume)Report a level from 0 to 1
isDisposed()Skip events after teardown

createVoiceSession wires status tracking, mute state, abort-signal disconnect, and all on* subscriptions for you. Return sendText from the controls only when the provider takes typed text; the session exposes it once the controls resolve, and the runtime records the typed turn itself.

Configuration

Pass a RealtimeVoiceAdapter implementation on the runtime:

const runtime = useChatRuntime({
  adapters: {
    voice: new MyVoiceAdapter({ /* provider options */ }),
  },
});

When a voice adapter is provided, capabilities.voice is set to true automatically.

React hooks

These hooks are exported from @assistant-ui/react and read the active voice session on the current thread.

useVoiceState

Returns the current VoiceSessionState, or undefined when no session is active:

import { useVoiceState } from "@assistant-ui/react";

const voiceState = useVoiceState();
// voiceState?.status.type: "starting" | "running" | "ended"
// voiceState?.isMuted: boolean
// voiceState?.mode: "listening" | "speaking"
// voiceState?.canSendText: boolean

VoiceSessionState is:

type VoiceSessionState = {
  readonly status: RealtimeVoiceAdapter.Status;
  readonly isMuted: boolean;
  readonly mode: RealtimeVoiceAdapter.Mode;
  readonly canSendText: boolean;
};

useVoiceVolume

Subscribes to real-time audio level independently of session state:

import { useVoiceVolume } from "@assistant-ui/react";

const volume = useVoiceVolume(); // number from 0 to 1

useVoiceControls

Returns methods that drive the session:

import { useVoiceControls } from "@assistant-ui/react";

const { connect, disconnect, mute, unmute } = useVoiceControls();

Headless control bar

Build your own controls from the hooks when you need a custom layout:

import { useVoiceState, useVoiceControls } from "@assistant-ui/react";
import { PhoneIcon, PhoneOffIcon, MicIcon, MicOffIcon } from "lucide-react";

function VoiceControls() {
  const voiceState = useVoiceState();
  const { connect, disconnect, mute, unmute } = useVoiceControls();

  const isRunning = voiceState?.status.type === "running";
  const isStarting = voiceState?.status.type === "starting";
  const isMuted = voiceState?.isMuted ?? false;

  if (!isRunning && !isStarting) {
    return (
      <button onClick={() => connect()}>
        <PhoneIcon /> Connect
      </button>
    );
  }

  return (
    <>
      <button onClick={() => (isMuted ? unmute() : mute())} disabled={!isRunning}>
        {isMuted ? <MicOffIcon /> : <MicIcon />}
        {isMuted ? "Unmute" : "Mute"}
      </button>
      <button onClick={() => disconnect()}>
        <PhoneOffIcon /> Disconnect
      </button>
    </>
  );
}

Voice UI component

The fastest path is the styled voice registry component. It ships a VoiceControl bar, mute and disconnect buttons, status indicator, and animated VoiceOrb built on the hooks above.

Install

npx shadcn@latest add @assistant-ui/voice

The @assistant-ui namespace resolves the Radix or Base UI flavor from your project's style through the style-aware registry entry in components.json. Without that entry, add by direct URL instead:

npx shadcn@latest add https://r.assistant-ui.com/base/voice.json

This adds /components/assistant-ui/elements/voice.tsx to your project. Adjust styling as needed.

Use with a runtime

app/page.tsx
import { Thread } from "@/components/assistant-ui/elements/thread.aui";
import { VoiceControl } from "@/components/assistant-ui/elements/voice.aui";
import { AuiIf } from "@assistant-ui/react";

export default function Chat() {
  return (
    <div className="flex h-full flex-col">
      <AuiIf condition={(s) => s.thread.capabilities.voice}>
        <VoiceControl />
      </AuiIf>
      <div className="min-h-0 flex-1">
        <Thread />
      </div>
    </div>
  );
}

See the Orb page for anatomy, variants, and state samples.

Example: LiveKit

LiveKit provides realtime voice over WebRTC rooms with transcription support. The browser adapter joins a room; a separate agent worker (STT, LLM, TTS) joins the same room. Without an agent in the room, the client connects but has nothing to talk to.

The with-livekit example shows the full wiring:

  1. Adapter (examples/with-livekit/lib/livekit-voice-adapter.ts): LiveKitVoiceAdapter implements RealtimeVoiceAdapter. connect calls createVoiceSession, connects a LiveKit Room, enables the local microphone, attaches remote audio tracks, maps room events to helpers (setStatus, end, emitMode, emitVolume, emitTranscript), and forwards typed text with localParticipant.sendText on the lk.chat topic, which the agent reads as text input.
  2. Token route (examples/with-livekit/app/api/livekit-token/route.ts): mints a ten-minute token for an isolated room so secrets stay off the client.
  3. Agent worker (examples/with-livekit/agent/): a Python LiveKit Agents process that joins the room and runs the voice pipeline.

Minimal runtime wiring:

import { LiveKitVoiceAdapter } from "@/lib/livekit-voice-adapter";

const runtime = useChatRuntime({
  adapters: {
    voice: new LiveKitVoiceAdapter({
      url: process.env.NEXT_PUBLIC_LIVEKIT_URL!,
      token: async () => {
        const res = await fetch("/api/livekit-token", { method: "POST" });
        if (!res.ok) throw new Error("Failed to create a LiveKit session");
        const { token } = await res.json();
        return token;
      },
    }),
  },
});

Clone the example for the adapter source, token endpoint, and agent worker rather than re-implementing the full LiveKit event map from scratch.

The token route accepts a request only when Sec-Fetch-Site is same-origin or none, and falls back to comparing the Origin header against the request URL for clients that omit Fetch Metadata. Behind a reverse proxy, that fallback needs the public scheme and host preserved in the request URL.

Warning

A request-context check is not authentication. Before deploying, require your application session in the token route and apply a durable rate limit, so anonymous callers cannot consume voice capacity.

You can also scaffold it with the CLI:

npx assistant-ui create my-app -e with-livekit
  • Dictation: push-to-talk speech-to-text into the composer (DictationAdapter, ComposerPrimitive.Dictate).
  • Speech: read-aloud for assistant messages (SpeechSynthesisAdapter, ActionBarPrimitive.Speak).
  • Orb: registry component reference for VoiceControl and VoiceOrb.
  • Voice API reference: generated types for sessions, speech, and dictation.