927tanmay/react-ai-voice-avatar

Open-source alternative to real-time avatar APIs like HeyGen Interactive Avatar and Tavus. A talking, lip-synced 3D avatar that runs on your user's GPU, driven by your own LLM. No video stream, no per-minute billing. Or voice mode alone, ChatGPT-style, as a headless React hook. 📦 npm: react-ai-voice-avatar

TypeScript

12

206 commits

updated Oct 1, 2026

See the code

See what people are saying

README

React AI Voice Avatar (react-ai-voice-avatar) 🚀🗣️🧬

NPM Version TypeScript Tested with Playwright Live Demo License: MIT

The open-source alternative to real-time avatar APIs, as a React component.

HeyGen Interactive Avatar, Tavus and Soul Machines sell a talking avatar that holds a conversation, rendered on their servers and billed per streaming minute. This does the same job as an npm install: the avatar renders on your user's GPU, so there is no video stream, no per-minute cost, and no third party in the middle of your conversations.

You bring the model. Point it at OpenAI, Anthropic, your own fine-tune or your existing chat endpoint, and keep your keys on your own backend. The package owns the parts that are tedious to build and easy to get wrong: microphone capture, knowing when someone has finished speaking, interrupting the avatar mid-sentence when they talk over it, streaming speech synthesis, and driving 52 ARKit facial blendshapes at 60 FPS so the mouth matches the words.

It can also run with no backend at all. Speech recognition, generation and voice all have in-browser implementations, which makes for a convincing demo and a genuinely offline kiosk. Most production apps will use their own model and keep only speech and lip-sync on the device.

Don't need a face? The same engine is a headless React hook for voice mode, the way ChatGPT and Gemini do it: react-ai-voice-avatar/headless, with no three.js in your bundle. See it ➔ · How ➔

🌐 Try the live demo ➔ · Voice only ➔

Speech recognition, voice synthesis and lip-sync, all running in your browser tab.

The first visit downloads about 590 MB of models, roughly a minute and a half on a 50 Mbit/s connection. Your browser keeps them, so a return visit is talking in about 2 seconds with nothing downloaded. Replies on the demo come from a hosted model on Groq's free tier, shared by everyone trying it, and start about 3 seconds after you stop talking; on a busy day that tier can run out, and examples/groq-voice runs the same thing on your own free key.

Talking to the avatar: it answers with captions and gestures, is interrupted mid-answer, stops and answers the new question


🌟 Two Entry Points (How to use it)

react-ai-voice-avatar provides two distinct ways to integrate into your app depending on your design needs. Both share the exact same underlying conversational state machine, Voice Activity Detection (VAD), and turn-taking logic.

🎧 Entry 1: The Headless Hook ("Voice Mode for your App")

Voice mode for your app: talk to it the way you talk to ChatGPT or Gemini, and interrupt it mid-sentence. The hook owns the microphone, knowing when someone has finished speaking, interruption, and streaming the reply into speech. You own the UI, and every frame it hands you a loudness level to animate.

Voice mode with no avatar: an orb that pulses with the voice, live captions, and a reply interrupted mid-answer by a new question

Try it ➔: a page built on this hook alone. examples/voice-only is the same page as an app to copy.

Import it from react-ai-voice-avatar/headless and Three.js never enters your module graph. Measured on the same Next.js App Router build, one route rendering the 3D avatar and one rendering only the hook:

RouteFirst Load JS
3D avatar389 kB
Headless hook117 kB
npm install react-ai-voice-avatar
import { useRef, useState } from 'react';
// The /headless subpath is what keeps Three.js out of your bundle.
// Importing the hook from the package root pulls the 3D stack in with it.
import { useAiVoiceAvatar } from 'react-ai-voice-avatar/headless';

export function VoiceMode() {
  const orb = useRef<HTMLDivElement>(null);
  const [micOn, setMicOn] = useState(false);

  const voice = useAiVoiceAvatar({
    // Your model. Return a string, or a stream so speech starts sooner.
    onSubmit: text => fetch('/api/chat', { method: 'POST', body: text }).then(r => r.body),

    // 0 to 1 every frame, from whoever is talking. Written straight to the
    // DOM: routing it through state would re-render sixty times a second.
    onAudioLevelChange: level => {
      if (orb.current) orb.current.style.transform = `scale(${1 + level * 0.3})`;
    },
  });

  // The mic stays open across turns, so it is tracked apart from `status`,
  // which moves through listening, thinking and speaking.
  const toggle = () => {
    if (micOn) {
      voice.stopListening();
      voice.interrupt();
    } else {
      voice.startListening(); // From a click: this is where the mic prompt appears.
    }
    setMicOn(!micOn);
  };

  return (
    <>
      <div ref={orb} className="orb" data-status={voice.status} />
      <button onClick={toggle} disabled={!voice.isReady}>
        {voice.isReady ? (micOn ? 'Stop' : 'Talk') : 'Loading…'}
      </button>
    </>
  );
}

For captions, onTranscriptUpdate(text, 'user') gives what the user said and onSpeechStart(text) each sentence of the reply as it starts playing. speak(text) says something without a model turn, such as a greeting, and sendText(text) takes typed input. Everything it returns is in the hook reference.

🚀 Voice mode in production

Each stage runs in the browser or on your backend, and the choice is mostly about the download:

SetupOptionsDownloaded by each visitorAudio leaves the device
Cloud speechonTranscribe + onSubmit + onSynthesizeNothing until the mic opens, then ~4 MB for the voice detectorYes, to your providers
Local speech (the demo)onSubmit~590 MB once, kept for later visits (~320 MB on iPhone)No
Fully localnone~1.3 GB onceNothing leaves at all

Cloud speech is the setup to reach for on phones and on pages people visit once: it is ready as soon as the page is. Local speech keeps what people say on their device and costs nothing per minute, and suits an app they come back to, since the models are stored after the first visit. Fully local is for kiosks and offline use. Use loadModels to hold any download until someone actually engages.

☁️ Cloud Adapters

onTranscribe and onSynthesize replace the local speech models with your own providers. Each replaces its model entirely: supply one and that model is never downloaded. Drop it later and the local model loads, so a provider outage can fall back to the browser mid-session.

To have that fallback ready rather than downloading it at the moment you need it, pass preloadLocalSpeech: the local hearing and voice download behind the conversation, and onLocalSpeechReady fires when they could take over. examples/groq-voice starts on Groq and offers the switch then.

const voice = useAiVoiceAvatar({
  // Instead of Whisper: 16 kHz mono samples of one utterance, to any
  // speech-to-text service. Encode them as WAV if the service wants a file.
  onTranscribe: async (samples: Float32Array) => {
    return await mySpeechToText(samples);
  },

  // Instead of Kokoro: return an encoded MP3/WAV (ArrayBuffer), which the
  // hook decodes, or raw 24 kHz PCM as a Float32Array. Called a sentence at
  // a time, so the first plays while the rest are fetched.
  onSynthesize: async (text: string) => {
    const res = await fetch('/api/tts', { method: 'POST', body: text });
    return await res.arrayBuffer();
  },

  onSubmit: async (text) => fetch('/api/chat', { method: 'POST', body: text }).then(r => r.body),
});

🧑‍💼 Entry 2: The Full 3D Avatar

If you want the full visual presence with 60FPS ARKit lip-syncing, use the drop-in 3D component. Under the hood, this is just a wrapper around the useAiVoiceAvatar hook that procedurally maps the audio to a 3D model!

📦 View Package on the Official NPM Registry ➔

# If using the 3D Avatar, you must also install the Three.js ecosystem
npm install react-ai-voice-avatar three @react-three/fiber @react-three/drei

[!NOTE] React 18 Users: Installing the latest @react-three/drei defaults to version 10, which demands React 19. If your project runs on React 18, install compatible Three.js React bindings explicitly:

npm install @react-three/drei@^9 @react-three/fiber@^8

[!IMPORTANT] React 19.3 and ERESOLVE: @react-three/fiber@9.7 still declares its React peer as >=19 <19.3, so a default npm install alongside React 19.3 or newer fails with ERESOLVE unable to resolve dependency tree. This is a Three.js binding constraint, not a limit of this package: our own peer range accepts React 19.3.

That range is over-cautious. We build and run a Next.js App Router app against React 19.3.0 with @react-three/fiber@9.7.0 and the avatar renders, loads its models and speaks with no errors. So install past it rather than downgrading:

npm install react-ai-voice-avatar three @react-three/fiber @react-three/drei --legacy-peer-deps

If you would rather keep strict peer resolution, pinning React works too:

npm install react@~19.2.0 react-dom@~19.2.0

The headless entry point (react-ai-voice-avatar/headless) pulls in no Three.js at all, so it never hits this and works on any React 18 or 19 version.

⚡ DX & Performance (Lazy Code-Splitting)

To prevent the massive ML assets (WebGPU workers, 3D engines) from bloating your initial page load, use the built-in lazy wrapper. It will automatically code-split the 3D dependencies and render a sleek holographic Skeleton UI while the assets download in the background!

import { AiVoiceAvatarLazy } from 'react-ai-voice-avatar';

// Use it exactly like the normal component!
<AiVoiceAvatarLazy avatarPreset="ananya" />

Show the avatar before anyone commits to a download

Lazy loading defers the 3D engine. The speech and language models are the larger cost, and by default they start downloading as soon as the avatar mounts. On a landing page most visitors only look, so render the avatar idle and load the models when someone actually engages:

const [engaged, setEngaged] = useState(false);

<AiVoiceAvatar loadModels={engaged} hideStatusPill={!engaged} />
<button onClick={() => setEngaged(true)}>Talk to it</button>

The live demo's landing page works this way.

📐 Architectural Best Practices

  • Standard Import (AiVoiceAvatar): Recommended for full-screen applications where the avatar is the primary product (e.g., Kiosks, Digital Tutors). The browser aggressively downloads the 3D canvas and ML models immediately so the avatar is ready instantly.
  • Lazy Import (AiVoiceAvatarLazy): Recommended for widgets, modals, or sub-routes (e.g., a "Support Desk" chat bubble in a SaaS dashboard). Defers downloading the 1.5MB 3D engine and WebWorkers until the user actually opens the widget.

⚙️ Server Configuration (Optional Performance Boost)

The react-ai-voice-avatar engine is truly zero-config. You do not need to configure Vite optimizeDeps, Next.js Webpack overrides, or manually host Web Worker files—everything is dynamically bundled and executed automatically!

However, because our ONNX WebGPU engine leverages modern multi-threaded SharedArrayBuffer memory pipelines for maximum inference speed, your hosting server can optionally emit standard Cross-Origin Isolation HTTP headers (COOP/COEP) to unlock peak performance. If these headers are not present, the engine automatically falls back to single-threaded WebAssembly without crashing.

🌐 Enabling Multi-threading on Production (Vercel, Netlify & Cloudflare)

To unlock multi-threaded performance, specify these isolation headers in your routing manifests:

  • Vercel (vercel.json): Add "headers": [{ "source": "/(.*)", "headers": [{ "key": "Cross-Origin-Opener-Policy", "value": "same-origin" }, { "key": "Cross-Origin-Embedder-Policy", "value": "require-corp" }] }].
  • Netlify / Cloudflare Pages (_headers or netlify.toml): Add /*\n Cross-Origin-Opener-Policy: same-origin\n Cross-Origin-Embedder-Policy: require-corp to public/_headers.

⚡ Enabling Multi-threading in Local Dev (Vite & Next.js)

Vite (vite.config.ts):

export default defineConfig({
  plugins: [react()],
  server: {
    headers: {
      'Cross-Origin-Opener-Policy': 'same-origin',
      'Cross-Origin-Embedder-Policy': 'require-corp',
    },
  },
});

Next.js (next.config.mjs):

export default {
  async headers() {
    return [
      {
        source: '/(.*)',
        headers: [
          { key: 'Cross-Origin-Opener-Policy', value: 'same-origin' },
          { key: 'Cross-Origin-Embedder-Policy', value: 'require-corp' },
        ],
      },
    ];
  },
};

[!CAUTION] Strict CSP Policies: If your enterprise enforces strict Content Security Policies that block blob: workers (worker-src 'self'), you can bypass our zero-config Blob loaders by passing the workerBaseUrl prop to the avatar and hosting the pre-compiled .worker.js files from our dist/assets/ directory yourself.

⚡ 3D Avatar Quickstart

import React, { useRef, useState } from 'react';
import { Canvas } from '@react-three/fiber';
import { OrbitControls } from '@react-three/drei';
import { AiVoiceAvatar, type AiVoiceAvatarHandle } from 'react-ai-voice-avatar';

export function App() {
  const avatarRef = useRef<AiVoiceAvatarHandle>(null);
  const [text, setText] = useState('');

  return (
    <div style={{ width: '100vw', height: '100vh', position: 'relative' }}>
      <Canvas camera={{ position: [0, 0.15, 2.2], fov: 32 }}>
        <color attach="background" args={['#101116']} />
        
        {/* Subtle studio lighting */}
        <pointLight position={[-3, 2, -2]} intensity={25} color="#E67E22" distance={6} />
        <pointLight position={[3, 1, -2]} intensity={20} color="#2980B9" distance={6} />
        
        <OrbitControls target={[0, 0.05, 0]} />
        
        {/* Connected Brain: Zero download, instant initialization! */}
        <AiVoiceAvatar
          ref={avatarRef}
          avatarPreset="ananya"
          lightingPreset="studio"
          ttsEngine="kokoro"
          ttsVoice="af_heart"
          // Connect your backend here (receives user speech transcript):
          onSubmit={async (text) => {
            const res = await fetch('/api/chat', { 
              method: 'POST', 
              body: JSON.stringify({ prompt: text }) 
            });
            return res.body; // Avatar natively reads streams!
          }}
        />
      </Canvas>

      {/* Fallback Text Input for noisy environments */}
      <form
        onSubmit={(e) => {
          e.preventDefault();
          if (text.trim() && avatarRef.current) {
            avatarRef.current.sendText(text);
            setText('');
          }
        }}
        style={{ position: 'absolute', bottom: '20px', left: '50%', transform: 'translateX(-50%)', display: 'flex', gap: '8px', zIndex: 100 }}
      >
        <input 
          value={text} 
          onChange={e => setText(e.target.value)} 
          placeholder="Type a message..." 
          style={{ padding: '8px 16px', borderRadius: '20px', border: 'none', background: 'rgba(255,255,255,0.9)', width: '300px' }}
        />
        <button type="submit" style={{ padding: '8px 16px', borderRadius: '20px', border: 'none', background: '#3b82f6', color: 'white', cursor: 'pointer' }}>
          Send
        </button>
      </form>
    </div>
  );
}

🔌 Connect Your Backend (Recipes)

The onSubmit prop natively accepts a string, an AsyncIterable<string>, or a ReadableStream. To connect your actual backend, simply drop in one of these copy-paste recipes to parse your streaming format!

[!CAUTION] API Keys Belong on the Server! Never put your OpenAI or Anthropic API keys directly in the frontend browser code. Always route through your own backend endpoint (/api/chat).

Recipe 0: Plain Text Stream (Fastest & Simplest)

If your backend uses Vercel AI SDK's streamText(...).toTextStreamResponse() or otherwise streams plain raw text, you can pass the stream natively without any parsing!

onSubmit={async (text) => {
  const res = await fetch('/api/chat', { method: 'POST', body: JSON.stringify({ prompt: text }) });
  return res.body; // Natively supported!
}}

Recipe 1: Vercel AI SDK (≤v4 Data Stream)

Older versions of the Vercel AI SDK stream data using a specific protocol (e.g., 0:"Hello"). This recipe parses those chunks into clean text with a carry-over buffer for safe network boundaries.

onSubmit={async function* (text) {
  const res = await fetch('/api/chat', { method: 'POST', body: JSON.stringify({ prompt: text }) });
  if (!res.body) return;
  const reader = res.body.getReader();
  const decoder = new TextDecoder();
  let buffer = '';
  
  while (true) {
    const { done, value } = await reader.read();
    if (done) break;
    buffer += decoder.decode(value, { stream: true });
    const lines = buffer.split('\n');
    buffer = lines.pop() ?? ''; // keep the trailing partial chunk
    for (const line of lines) {
      if (line.startsWith('0:')) {
        try { yield JSON.parse(line.substring(2)); } catch { /* ignore keep-alive / non-JSON frames */ }
      }
    }
  }
}}

Recipe 2: OpenAI-Compatible SSE Endpoint (and AI SDK v5)

Standard Server-Sent Events (SSE) stream data: {...} blocks. This handles safe parsing across broken network chunk boundaries.

onSubmit={async function* (text) {
  const res = await fetch('/api/chat', { method: 'POST', body: JSON.stringify({ prompt: text }) });
  if (!res.body) return;
  const reader = res.body.getReader();
  const decoder = new TextDecoder();
  let buffer = '';
  
  while (true) {
    const { done, value } = await reader.read();
    if (done) break;
    buffer += decoder.decode(value, { stream: true });
    const lines = buffer.split('\n');
    buffer = lines.pop() ?? ''; // keep the trailing partial chunk
    for (const line of lines) {
      if (line.startsWith('data: ') && line !== 'data: [DONE]') {
        try {
          const parsed = JSON.parse(line.substring(6));
          // AI SDK v5 emits {type:'text-delta', delta:'...'}; OpenAI emits choices[0].delta.content
          if (parsed.type === 'text-delta' && parsed.delta) {
            yield parsed.delta;
          } else if (parsed.choices?.[0]?.delta?.content) {
            yield parsed.choices[0].delta.content;
          }
        } catch { /* ignore keep-alive / non-JSON frames */ }
      }
    }
  }
}}

🗣️ Languages

English and Hindi both run entirely on the device. Set ttsLanguage and the engine picks a matching voice, a matching speech-recognition model, and a matching phoneme path.

<AiVoiceAvatar ttsLanguage="hi-IN" />

Hindi needed real work rather than a config flag. Kokoro ships four Hindi voices inside the same checkpoint as the English ones, but the JavaScript wrapper does not list them, and the bundled eSpeak build carries English data only and rejects hi outright. So this package includes its own Devanagari-to-phoneme converter. Devanagari is close to phonemic, which makes that tractable; the hard part is schwa deletion, the rule that makes कमल read as "kamal" rather than "kamala", and getting it wrong produces speech that still sounds like speech while sounding like someone spelling Hindi out.

Because the converter emits real phonemes, Hindi gets phoneme-driven lip-sync with distinct mouth shapes for the retroflex consonants, not the amplitude-only mouth flapping most engines fall back to outside English. Code-switching works too: English words inside a Hindi sentence are routed to the English phonemiser, so "मुझे coffee चाहिए" is pronounced correctly throughout.

[!NOTE] Local Hindi is a demo, not a product. Models small enough to run in a browser are far weaker in Hindi than in English. It is good enough to show the pipeline working end to end and not good enough to ship. For production Hindi, route hearing and thinking to an API through onTranscribe and onSubmit; the voice and lip-sync stay local and are genuinely good.

A session is one language at a time. Someone who switches language mid-conversation will be mistranscribed, because the recognition model is told which language to expect and browser Whisper cannot detect it.


🎤 Taking turns

The user can talk over the avatar and cut it off mid-sentence. This is on by default, because waiting for a reply to finish is the thing that makes a voice agent feel like a walkie-talkie.

<AiVoiceAvatar
  allowInterruption={false}   // default: true
  onUserInterrupt={() => analytics.track('barge_in')}
/>

Turn it off for a kiosk or a noisy room, where the avatar hearing its own voice through the speakers and stopping itself is worse than waiting. It is ignored in push-to-talk, which owns the floor explicitly.

How a turn is decided

Voice detection scores every 96ms frame for how much it sounds like speech. A cough, a door and a chair all clear a loudness bar as easily as a word does, so loudness alone cannot separate them — sustain can. A sound has to keep scoring as speech for minSpeechMs before the engine treats it as a turn.

Below that bar, a sound only ducks the avatar's voice, reversibly. If it turns out to be a cough the reply resumes from where it paused, at the right place in the sentence and with the mouth still in sync. Nothing about the conversation changed, because nothing was decided on a noise.

Tuning for your room

The defaults suit a quiet room and headphones. A shop floor is a different problem, and only you know which you have.

<AiVoiceAvatar
  speechDetection={{ positiveSpeechThreshold: 0.6, minSpeechMs: 700 }}
/>
FieldDefaultRaise it whenLower it when
positiveSpeechThreshold0.5Passing noise is mistaken for talkingQuiet speakers go unheard
negativeSpeechThreshold0.35Turns end too slowly in a noisy roomTurns end while someone is still talking
minSpeechMs500Short noises still start turnsSingle-word answers are ignored
redemptionMs1400People are cut off while thinking mid-sentenceReplies feel slow to start
preSpeechPadMs800The first word is still being clippedRarely — this is the audio kept from before the trigger, and it is what stops the first syllable going missing

Two costs worth knowing before you change anything. A speaker who sits below positiveSpeechThreshold produces no events at all — raising it trades quiet voices for quiet rooms. And a filler like "hmm" held long enough to pass minSpeechMs still reaches transcription; a denylist catches the common ones after the fact, but sustain cannot tell a long "hmm" from a short word.


🚨 Handling failures

Pass onError and your app learns when something breaks, rather than finding out from a console message it cannot see.

<AiVoiceAvatar
  onError={(e) => {
    if (e.severity === 'fatal') showFallbackUI(e.stage);
    logToSentry(e);
  }}
/>

Check severity before reacting. Most failures here are survivable because the engine falls back: WebGPU to WASM, Kokoro to a smaller voice model. Those arrive as degraded and the avatar still works, so treating them as fatal would hide a working experience behind an error screen. A refused microphone is fatal for listening while typed input still works, which is a judgement only your app can make.

stage is one of microphone, speech-recognition, language-model, speech-synthesis, audio-output, worker, model-storage or conversation. There is also a detail string carrying the engine's internal stage name for bug reports; it is not stable across versions, so do not branch on it.

model-storage is always degraded: a model loaded and works, but could not be kept, so the next visit downloads it again. The message says why — usually not enough free space, which on a phone is the common case.

Models are kept between visits

Downloaded models are stored in the browser's Origin Private File System, so a returning visitor skips the download: on an Apple M4 with WebGPU, a return visit is ready to talk in about 2 seconds, with nothing fetched. The browser's Cache API — what the model loader uses by default — refuses any single file of 256 MiB or more, and both the voice (~310 MB) and the local language model (~750 MB) are larger than that, so before this they were downloaded again on every visit. Where OPFS is unavailable the Cache API is used as before.

Browsers can clear this storage when disk space runs low. If your app is one people return to, call navigator.storage.persist() from a user gesture to ask the browser to keep it; the engine does not, because Firefox can answer that call with a permission prompt, and a library should not put one in front of your visitors unasked.


🎨 Bring Your Own 3D Avatar (Custom GLB)

You are not locked into our built-in avatars (ananya and aarav)! You can use any custom .glb humanoid model by passing its URL or local path to the modelSrc prop:

<AiVoiceAvatar
  modelSrc="/models/my-custom-avatar.glb"
  // ...
/>

📋 Custom Avatar Requirements

To ensure the lip-sync and procedural facial dynamics engines work correctly, your custom model must meet the following standard requirements:

  1. Format: .glb (GLTF Binary).
  2. Facial Blendshapes (Morph Targets): The model's head/face mesh must contain the standard 52 Apple ARKit blendshapes (e.g., jawOpen, eyeBlinkLeft, mouthSmileRight). Our engine automatically traverses your model to find these targets.
  3. Bone Naming: For the interactive mouse-tracking and head-tilting physics to function, the armature should use standard bone names (e.g., a neck/head bone named Head, head, Neck, or neck).

Where the built-in avatars come from

ananya and aarav are converted from the Microsoft Rocketbox library, which is MIT licensed like this project, so you can redistribute them without restriction. The library has 115 avatars and any of them can be converted with scripts/convert-rocketbox.py. See assets/avatars/LICENSE.md.

Models exported from Ready Player Me also work, since they carry the same ARKit blendshapes and bone names. Note that Ready Player Me shut down in January 2026, so you can no longer create new avatars there, and existing exports are licensed CC BY-NC-SA rather than MIT.


🏗️ What runs where

Four stages, and you choose where each one happens. The defaults are all local, which is why the demo needs no keys, but the interesting production setups are mixed.

flowchart TD
  mic(["🎙️ You speak"]) --> vad["Voice detector<br/><i>in the browser</i>"]
  vad --> hear["Hearing<br/><i>Whisper in the browser</i><br/>or your onTranscribe"]
  hear --> think["Thinking<br/><i>a model in the browser</i><br/>or your onSubmit"]
  think --> speak["Speaking<br/><i>Kokoro in the browser</i><br/>or your onSynthesize"]
  speak --> face["🔊 Voice, with a lip-synced<br/>3D face or your own UI"]
  vad -. "talk over it and it stops" .-> speak

Each stage runs in the browser unless you pass the adapter beside it, which hands that stage to your own service.

StageOn the deviceYour backend instead
Hearing — speech to textWhisper via ONNXonTranscribe
Thinking — the replyQwen or Gemma via WebGPUonSubmit
Speaking — text to audioKokoro-82MonSynthesize
Face — lip-sync and animationAlways hereNot applicable

The face never leaves the device, which is the whole point: that is what avatar APIs charge per minute for, and it is the one stage that cannot be outsourced without a video stream.

The common production shape is onSubmit alone. Hearing and speaking stay local, so no audio ever leaves the browser, while generation goes to whatever model you already run. That keeps the download to about 600 MB (Whisper base and the Kokoro voice), keeps your keys on your server, and still gives you a conversation nobody else can read.

Add onTranscribe and onSynthesize as well and nothing downloads at all, apart from the ~4 MB voice detector once the microphone opens. Each adapter replaces its model entirely, so the page is ready as soon as it loads. That is the shape for phones and for pages people visit once.

Fully local is real, not a demo trick, and it is the right answer for a kiosk, a regulated environment, or anywhere without reliable connectivity. Be aware of the cost: in English the first visit downloads roughly 1.3 GB before anyone can speak (the language model alone is 750 MB), about 2.1 GB in Hindi, and the quality ceiling is whatever a model that size can do.

⏱️ How fast it answers

Measured with scripts/measure-latency.mjs on an Apple M4 with WebGPU in Chromium: a spoken question through a fake microphone, the median of five turns after a warm-up turn, timed by the hook's own callbacks.

StageTime
Waiting for silence, to be sure you've finished1,400 ms (redemptionMs, adjustable)
Hearing: Whisper base transcribes the utterance330 ms
Thinking, hosted: GPT-OSS 20B on Groq, the demo's route, one round trip670 ms
Thinking and the first sentence of voice, all in the browser (Qwen 0.5B, Kokoro)740 ms
Speaking: Kokoro's first sentence ready and playing, with an instant reply460 ms
Starting up on a return visit, models already storedabout 2 s, nothing downloaded

So from the moment you stop talking to the first word of the reply:

SetupFirst word after you stop
All in the browserabout 2.5 s
Hosted replies, local hearing and voice (the demo)about 2.9 s
With an instant reply, the floor for speech aloneabout 2.2 s

The hosted figure adds the route's round trip to the measured speech stages, since that route answers only the live site. Most of the wait is the pause for silence, which is what stops the avatar cutting people off while they think. Lower redemptionMs in speechDetection for a snappier feel, at the cost of answering half-finished sentences. Without WebGPU, recognition and the voice run on the CPU and are several times slower.

🔀 Hosted now, local when warm

You do not have to choose once. preloadLocalLlm fetches the in-browser model behind the conversation while your backend answers, and onLocalLlmReady tells you when it can take over:

const [localReady, setLocalReady] = useState(false);

<AiVoiceAvatar
  // Dropping onSubmit is what hands the conversation over.
  onSubmit={localReady ? undefined : askMyBackend}
  preloadLocalLlm
  onLocalLlmReady={() => setLocalReady(true)}
/>

Nobody waits for a gigabyte to say the first word, and the turns your backend answered are handed to the local model, so it does not restart the conversation from nothing. Useful for a kiosk that must keep working when the wifi drops, a demo on someone else's quota, or a phone, which will never accept the download but can hold a conversation the moment it loads.

The live demo runs exactly this: replies come from a hosted model, and on a desktop with WebGPU the in-browser model downloads during the conversation and takes the floor when it lands. The page says which one is answering at any moment.

Explore the canonical patterns in the examples/ directory:

Example PatternFolderHighlights & Architecture
Live Interactive DemosandboxDeploy on Vercel ➔ — Our full-featured interactive testbed featuring live character switching (ananya, aarav), voice persona switching (af_heart, am_michael), real-time diagnostic probe metrics, and Leva 3D lighting controls.
Quickstartexamples/quickstartMinimal, zero-configuration plug-and-play AI voice avatar deployment with built-in studio lighting & sizing.
Local Kioskexamples/local-kiosk100% offline on-device retail & restaurant ordering kiosk with embedded menu reasoning. Demonstrates the On-Device Brain; operates without internet access once model weights are locally cached.
Connected Appexamples/hybrid-cloudIllustrates the Connected Brain (onSubmit). Bypasses gigabyte-scale local LLM downloads by routing reasoning to OpenAI, Claude, or corporate APIs while keeping ASR, TTS, and 3D lip blending 100% on-device!
Voice Onlyexamples/voice-onlyVoice mode with no avatar, like ChatGPT or Gemini voice: the react-ai-voice-avatar/headless hook, an audio-reactive orb, live captions and interruption. No three.js in the bundle, and no backend needed to try it.
Groq Voiceexamples/groq-voiceVoice mode on Groq's free tier with your own key: hearing, replies and voice start on Groq with nothing to download, the in-browser hearing and voice download behind the conversation (preloadLocalSpeech), and the page offers to switch when they are ready. Shows Groq's live limits.
Headless Custom UIexamples/headless-custom-uiStill renders the 3D avatar; for no avatar at all, see Voice Only above. Demonstrates hiding built-in DOM overlays (hideStatusPill={true}, showCaptions={false}), streaming transcripts into a custom enterprise UI, and controlling voice outputs imperatively via ref.current?.speak(text).

📖 API Reference

<AiVoiceAvatar /> Props

PropTypeDefaultDescription
avatarPreset'ananya' | 'aarav' | 'default' | 'kiosk''ananya'Built-in 3D character models featuring both female ('ananya') and male ('aarav') voice concierges out of the box with full ARKit facial blendshapes!
avatarSize'sm' | 'md' | 'lg' | number'md' (0.48)Intuitive model sizing presets or custom decimal scaling multiplier applied directly to the 3D humanoid mesh.
modelSrcstringundefinedAbsolute local path or remote URL to a custom GLTF/GLB humanoid armature avatar model.
lightingPreset'studio' | 'cyberpunk_violet' | 'cool_azure' | 'warm_amber' | 'clean_white' | 'none''studio'Pre-built cinematic studio lighting atmospheres directly applied to your 3D viewport without manual Three.js configuration!
systemPromptstring"You are Ananya..."Conversational persona directives and context injected into active LLMs.
llmModelstringby languageHugging Face id for the local WebGPU reasoning model, used only when onSubmit is absent. Defaults to Qwen2.5-0.5B for English and Gemma 3 1B for Hindi, which Qwen that size cannot speak. Set it to pin one model for every language.
asrModelstringby languageHugging Face id for the local Whisper model. Defaults to Whisper base for English and Whisper small for Hindi, which base transcribes badly. Pass "Xenova/whisper-tiny" for a faster download and worse accuracy.
ttsEngine'kokoro' | 'mms''kokoro'High-fidelity neural voice synthesis engine executing inside dedicated Web Workers.
ttsVoicestringby languageKokoro voice id. af_* and am_* American, bf_* and bm_* British, hf_* and hm_* Hindi (af_heart, am_michael, bf_emma, hf_alpha, hm_omega). Defaults to one matching ttsLanguage. A voice whose language disagrees is corrected with a warning.
ttsLanguage'en-US' | 'en-GB' | 'hi-IN''en-US'Conversation language. Selects the voice, the recognition model and the phoneme path. Also sets asrLanguage unless you set that yourself. For any other language, pass onSynthesize and use a cloud voice provider.
asrLanguagestringttsLanguageLanguage to transcribe. Follows ttsLanguage by default, since a conversation is almost always held in one language.
showCaptionsbooleantrueRenders a sleek glassmorphic subtitle overlay displaying spoken interaction dialog.
hideStatusPillbooleanfalseWhen true, suppresses the default bottom-left microphone interactive control pill.
listenMode'continuous' | 'push-to-talk''continuous'continuous keeps the mic hot after the avatar finishes speaking naturally, but explicitly clicking Stop forces it off until tapped again. push-to-talk strictly requires manually tapping to start listening for every single turn.
loadModelsbooleantrueSet false to render the avatar without downloading any models, then flip it true when the visitor engages. For landing pages and widgets most visitors never talk to. Status stays 'loading' until it is true and the models are up, so show your own call to action meanwhile.
onModelLoaded() => voidundefinedFires once the 3D mesh is parsed and in the scene. Parsing a multi-megabyte GLB leaves the canvas empty for a few seconds; use this to hold a placeholder over it.
allowInterruptionbooleantrueLets the user talk over the avatar and cut it off mid-sentence. Turn off for a kiosk or noisy room, where the avatar hearing itself through the speakers is worse than waiting. Ignored in push-to-talk. See Taking turns.
speechDetection{ positiveSpeechThreshold?, negativeSpeechThreshold?, minSpeechMs?, redemptionMs?, preSpeechPadMs? }see Taking turnsTunes how the microphone decides someone is talking. Every field optional. The defaults suit a quiet room; a shop floor needs a higher threshold and a longer minSpeechMs.
onUserInterrupt() => voidundefinedFires when the user talks over the avatar and takes the floor. Only fires if the avatar actually had audio playing.
gesturesboolean | numbertrueHand and arm gestures while the avatar speaks: one hand or both brought up in front of the chest for each phrase, with small beats on stressed syllables, and the arms back at rest when it stops. false keeps the arms still; a number sets the size, 0 to 1.5. Needs a skeleton with LeftArm, LeftForeArm and LeftHand and the right-hand equivalents; custom avatars without them simply do not gesture.
onAudioLevelChange(level: number, source: 'mic' | 'tts' | 'idle') => voidundefinedReal-time audio amplitude (0-1) callbacks for the active stream. Essential for building highly responsive, audio-reactive 3D Visualizers and HUDs! Fires with 0 and 'idle' between turns, so a meter falls to rest rather than freezing.
onSubmit(text: string) => Promise<string | AsyncIterable<string> | ReadableStream>undefinedConnected Brain API: Bypasses local LLMs; routes transcribed user microphone strings to your cloud or custom LLM API endpoint.
preloadLocalLlmbooleanfalseOnly meaningful alongside onSubmit, which otherwise skips the local language model download entirely. Set it to fetch that model in the background while your hosted one answers, so the conversation survives a rate limit, an expired quota or a lost network. See Hosted now, local when warm.
onLocalLlmReady() => voidundefinedFires once the model requested by preloadLocalLlm has loaded. Drop onSubmit here to hand the conversation over; the turns your backend answered are carried across, so the local model knows what was already said.
onTranscribe(audio: Float32Array) => Promise<string>undefinedReplaces local speech recognition with your own service, and Whisper is then not downloaded. Receives one utterance as 16 kHz mono samples.
onSynthesize(text: string) => Promise<Float32Array | ArrayBuffer>undefinedReplaces local voice synthesis with your own service, and Kokoro is then not downloaded. Return an encoded MP3/WAV buffer, or raw 24 kHz PCM. Called a sentence at a time.
onError(e: AiVoiceAvatarError) => voidundefinedFires when a stage fails. Carries stage, message and a severity of degraded or fatal. See Handling failures.
onTranscriptUpdate(text: string, speaker: 'user' | 'avatar') => voidundefinedCallback delivering real-time microphone transcriptions and assistant spoken utterance strings.
onStatusChange(status: string) => voidundefinedEmits live state transitions (loading, idle, listening, thinking, speaking).
debugbooleanfalseWhen true, renders an interactive floating GUI (Leva) to inspect and tune individual 3D blendshapes.
vadAssetPathstringundefinedOptional URL or local path override for self-hosting @ricky0123/vad-web ONNX asset binaries in airgapped deployments.
onnxWasmPathstringundefinedOptional URL override for self-hosting onnxruntime-web WASM distribution files.
workerBaseUrlstringundefinedCSP Escape Hatch: if blob: workers are blocked by your server, fetch pre-compiled Web Workers from this URL directory.
enableLocalAssetProbebooleanfalseWhen true, HEAD-checks /ananya.glb in your own public directory before falling back to the CDN. Off by default: with no local copy the probe 404s, and that 404 lands in every visitor's console.
statusPillStyleReact.CSSPropertiesundefinedOptional custom CSS styling & absolute positioning overrides for the interactive Status Pill overlay.
accentColorstringundefinedCustom CSS color string (e.g., #38BDF8) for the active status indicator rings and highlights.

Imperative Ref API (AiVoiceAvatarHandle)

Attach a React ref (useRef<AiVoiceAvatarHandle>(null)) to access imperative real-time controls:

interface AiVoiceAvatarHandle {
  /** Command the 3D avatar to speak an arbitrary string with synchronized acoustic lip blending */
  speak: (text: string) => void;
  /** Manually engage microphone recording and Voice Activity Detection (VAD) */
  startListening: () => void;
  /** Pause active microphone listening */
  stopListening: () => void;
  /** Instantly interrupt and halt active voice speech synthesis and clear the audio queue */
  interrupt: () => void;
  /** Manually submit text to the onSubmit handler, simulating a spoken utterance (useful for text-only fallback) */
  sendText: (text: string) => void;
  /** Wipe multi-turn conversation memory history and caption overlay states */
  clearHistory: () => void;
  /** Retrieve live Web Audio API AnalyserNode powering real-time spectral lip sync */
  getAnalyser: () => AnalyserNode | undefined;
}

💬 Text-Only Input (sendText)

If your users cannot use a microphone (e.g., noisy environments, privacy concerns, or lack of permissions), you can easily wire up a standard text input field to bypass the speech recognition pipeline entirely!

Simply attach a ref and call sendText() to pass a string directly to your onSubmit handler (or local LLM):

const avatarRef = useRef<AiVoiceAvatarHandle>(null);

// In your UI, attach this to a standard <form> submission:
const handleTextSubmit = (userInput: string) => {
  avatarRef.current?.sendText(userInput);
}

When you use sendText, the avatar immediately enters the thinking state and processes the interaction exactly as if the user had spoken it aloud.

useAiVoiceAvatar() (headless)

import { useAiVoiceAvatar } from 'react-ai-voice-avatar/headless';

Takes the same options as the component's props, except the ones about the 3D scene and its overlays (avatarPreset, avatarSize, modelSrc, lightingPreset, gestures, showCaptions, hideStatusPill, onModelLoaded, debug and the styling props). onStatusChange is replaced by the returned status. A few options matter mostly without an avatar:

OptionTypeDescription
onAudioLevelChange(level: number, source: 'mic' | 'tts' | 'idle') => voidLoudness from 0 to 1, every frame, from whichever side is talking. What an orb or waveform animates from. Write it to the DOM through a ref rather than into state.
onSpeechStart(text: string) => voidEach sentence of the reply as it starts playing. Captions that keep pace with the voice.
onTranscriptUpdate(text: string, speaker: 'user' | 'avatar') => voidWhat the user said, once transcribed, and the reply in full.
onInferenceStart / onInferenceEnd() => voidAround each turn, from the moment the user stops talking to the end of the reply.
onTtsEngineChange(engine: 'kokoro' | 'mms' | 'custom') => voidFires whenever activeTtsEngine changes. Also a prop on the component.
preloadLocalSpeechbooleanDownload the in-browser hearing and voice behind onTranscribe and onSynthesize, without waiting for them. Drop the adapters once onLocalSpeechReady fires and the local models take over. Also a prop on the component.
onLocalSpeechReady() => voidFires once, when the models requested by preloadLocalSpeech are loaded.
loadingProgress(pct: number, label: string) => voidDownload progress per model: 'asr', 'kokoro' (or 'tts' for MMS) and 'llm'. They download in parallel, so keep one figure per label.

It returns:

ValueTypeDescription
status'loading' | 'idle' | 'listening' | 'thinking' | 'speaking'Where the conversation is. 'loading' until the models are up, and until loadModels is true.
isLoading, isIdle, isListening, isThinking, isSpeakingbooleanShorthands for status.
isReadybooleanThe models are loaded. startListening does nothing before this.
isLocalSpeechReadybooleanThe in-browser hearing and voice are loaded, whether or not they are in use.
activeTtsEngine'kokoro' | 'mms' | 'custom'The voice actually speaking. iPhones and iPads get 'mms', a single plainer voice that ignores ttsVoice, because Kokoro runs Safari out of memory; so does any device where Kokoro fails to load. 'custom' is your onSynthesize. Worth showing if your users will compare devices.
startListening() => Promise<void>Opens the microphone and starts listening. Call it from a click: the first call is where the browser asks for permission. In 'continuous' mode the microphone then stays open across turns.
stopListening() => voidCloses the microphone.
interrupt() => voidStops the reply mid-sentence and clears what was queued.
speak(text: string) => voidSays the text without a model turn: a greeting, a notification. Plays a sentence at a time.
sendText(text: string) => voidTyped input, handled as if it had been spoken.
clearHistory() => voidForgets the conversation so far.
micErrorstring | nullWhy the microphone could not open, such as a denied permission. null otherwise.
analyserAnalyserNode | undefinedThe reply's audio, for a frequency visualiser. onAudioLevelChange is simpler if you only need loudness.

It also returns several refs (currentSpeechTextRef, audioContextRef and others) that the 3D component uses for lip-sync. A voice UI can ignore them.


🌐 Performance & Asset Caching

  1. Native WebGPU & WASM Degradation:
    • Modern Chromium browsers (Chrome, Edge, Opera, Arc) on desktop and mobile platforms benefit from hardware-accelerated WebGPU neural execution.
    • On systems without WebGPU, inference automatically falls back to multi-threaded WebAssembly (WASM) quantization without app crashes.
  2. Persistent Local Caching:
    • AI models (Whisper, Kokoro, the local language model) are downloaded once and kept in the browser's origin private file system, so later visits skip the download. See Models are kept between visits.

🤝 Contributing & Open Issues Roadmap

We actively welcome community contributions. CONTRIBUTING.md has the local development guide, and ROADMAP.md has what is worth doing next, why, and what is already known about each item — including the measurements behind the open questions.

The largest pieces currently open:

  1. 🖐️ Open-palm gestures. Hands now gesture while the avatar speaks, but always with the palms facing inward. Turning them up — the open, offering gesture people use when explaining — needs the forearm's roll, which has to be worked out from the finger bones rather than assumed, for the same reason as everything else in armRig.ts.
  2. 🎤 Turn-taking in real rooms. Turns are tuned against one speaker in a quiet room. A short "yes" can be dropped, and a quiet speaker can go unheard. Recordings from more voices and noisier rooms would let the defaults be set against something other than one person.
  3. 🎭 Expanding regional 3D avatar personas. Ananya and Aarav ship out of the box. Royalty-free character GLBs (~3MB) rigged with the standard 52 Apple ARKit facial blendshapes are welcome. Avatar meshes are served from a CDN rather than bundled, so adding one adds nothing to the npm install (2.3 MB tarball, 6.0 MB unpacked, asserted in CI by scripts/verify-pack.mjs).
  4. 🙌 Gestures. The avatar stands still while it speaks. Hand and arm movement tied to speech is the most visible thing still missing.
  5. 📱 React Native / Expo support. Exploring bindings to run ONNX inference and Three.js on mobile runtimes.

Hindi speech and VAD sensitivity tuning (speechDetection) were previously listed here and have both shipped.


🧭 Browser Compatibility Matrix

This library heavily relies on modern Web APIs (WebGPU, WebGL, Web Audio, and Web Workers). It gracefully degrades when certain APIs are unavailable.

BrowserOS3D Rendering (WebGL)Voice Synthesis (WebGPU/WASM)Voice Recognition (Web Audio)Status
Chrome / EdgeWindows, macOS, Android✅ Native✅ WebGPU (Ultra Fast)✅ Native🟢 Tier 1 (Recommended)
Safari / iOSmacOS, iOS✅ Native🔄 Lightweight Models by Default✅ Native🟡 Supported
FirefoxWindows, macOS✅ Native⚠️ WASM Fallback✅ Native🟡 Tier 2 (Slower TTS)

[!NOTE]

  • WebGPU is currently enabled by default in Chrome/Edge. On browsers without WebGPU, the library automatically falls back to WASM execution.
  • iOS/Safari Preemptive Fallback: Safari and iOS impose strict memory limits that cause 80MB+ models (like Kokoro) to crash the tab. The engine automatically preempts this by forcing the lightweight MMS TTS model (~30MB) on iOS devices, and tracks crash breadcrumbs to prevent OOM reload loops.
  • Strict CSP Environments: Safari and Firefox may block blob: worker execution depending on your Content-Security-Policy headers. If this occurs, host the .worker.js files statically and pass their base path via the workerBaseUrl prop.

💻 Hardware Requirements

Running Neural Networks in the browser requires capable hardware.

Deployment ModeMin RAMGPU RequirementRecommended Devices
Connected Brain (ASR + TTS only)4GBNone (WASM Fallback ok)iPhone 11+, Mid-range Android (2021+), Any Laptop
Full Local AI (ASR + 500M LLM + TTS)8GBWebGPU Support PreferrediPhone 13 Pro+, High-end Android (Snapdragon 8 Gen 1+), M1/M2 Macs, Modern PCs

[!TIP] Mobile Memory Limits: Mobile browsers rigidly enforce memory limits per tab (often terminating tabs exceeding ~1GB). If your mobile app crashes "after some time", ensure you are utilizing the Connected Brain mode (onSubmit API) which offloads the heavy LLM memory footprint to your server while keeping ultra-fast lip-sync and TTS local.


📜 License

MIT © React AI Voice Avatar Contributors.

ai-avatar
conversational-ai
headless
heygen-alternative
hindi
kokoro
lipsync
onnx
r3f
react
react-hook
speech-to-speech
threejs
tts
virtual-assistant
voice-ai
voice-assistant
voice-mode
webgpu
whisper

927tanmay/react-ai-voice-avatar

Open-source alternative to real-time avatar APIs like HeyGen Interactive Avatar and Tavus. A talking, lip-synced 3D avatar that runs on your user's GPU, driven by your own LLM. No video stream, no per-minute billing. Or voice mode alone, ChatGPT-style, as a headless React hook. 📦 npm: react-ai-voice-avatar

TypeScript

12

206 commits

updated Oct 1, 2026

See the code

See what people are saying

README

React AI Voice Avatar (react-ai-voice-avatar) 🚀🗣️🧬

NPM Version TypeScript Tested with Playwright Live Demo License: MIT

The open-source alternative to real-time avatar APIs, as a React component.

HeyGen Interactive Avatar, Tavus and Soul Machines sell a talking avatar that holds a conversation, rendered on their servers and billed per streaming minute. This does the same job as an npm install: the avatar renders on your user's GPU, so there is no video stream, no per-minute cost, and no third party in the middle of your conversations.

You bring the model. Point it at OpenAI, Anthropic, your own fine-tune or your existing chat endpoint, and keep your keys on your own backend. The package owns the parts that are tedious to build and easy to get wrong: microphone capture, knowing when someone has finished speaking, interrupting the avatar mid-sentence when they talk over it, streaming speech synthesis, and driving 52 ARKit facial blendshapes at 60 FPS so the mouth matches the words.

It can also run with no backend at all. Speech recognition, generation and voice all have in-browser implementations, which makes for a convincing demo and a genuinely offline kiosk. Most production apps will use their own model and keep only speech and lip-sync on the device.

Don't need a face? The same engine is a headless React hook for voice mode, the way ChatGPT and Gemini do it: react-ai-voice-avatar/headless, with no three.js in your bundle. See it ➔ · How ➔

🌐 Try the live demo ➔ · Voice only ➔

Speech recognition, voice synthesis and lip-sync, all running in your browser tab.

The first visit downloads about 590 MB of models, roughly a minute and a half on a 50 Mbit/s connection. Your browser keeps them, so a return visit is talking in about 2 seconds with nothing downloaded. Replies on the demo come from a hosted model on Groq's free tier, shared by everyone trying it, and start about 3 seconds after you stop talking; on a busy day that tier can run out, and examples/groq-voice runs the same thing on your own free key.

Talking to the avatar: it answers with captions and gestures, is interrupted mid-answer, stops and answers the new question


🌟 Two Entry Points (How to use it)

react-ai-voice-avatar provides two distinct ways to integrate into your app depending on your design needs. Both share the exact same underlying conversational state machine, Voice Activity Detection (VAD), and turn-taking logic.

🎧 Entry 1: The Headless Hook ("Voice Mode for your App")

Voice mode for your app: talk to it the way you talk to ChatGPT or Gemini, and interrupt it mid-sentence. The hook owns the microphone, knowing when someone has finished speaking, interruption, and streaming the reply into speech. You own the UI, and every frame it hands you a loudness level to animate.

Voice mode with no avatar: an orb that pulses with the voice, live captions, and a reply interrupted mid-answer by a new question

Try it ➔: a page built on this hook alone. examples/voice-only is the same page as an app to copy.

Import it from react-ai-voice-avatar/headless and Three.js never enters your module graph. Measured on the same Next.js App Router build, one route rendering the 3D avatar and one rendering only the hook:

RouteFirst Load JS
3D avatar389 kB
Headless hook117 kB
npm install react-ai-voice-avatar
import { useRef, useState } from 'react';
// The /headless subpath is what keeps Three.js out of your bundle.
// Importing the hook from the package root pulls the 3D stack in with it.
import { useAiVoiceAvatar } from 'react-ai-voice-avatar/headless';

export function VoiceMode() {
  const orb = useRef<HTMLDivElement>(null);
  const [micOn, setMicOn] = useState(false);

  const voice = useAiVoiceAvatar({
    // Your model. Return a string, or a stream so speech starts sooner.
    onSubmit: text => fetch('/api/chat', { method: 'POST', body: text }).then(r => r.body),

    // 0 to 1 every frame, from whoever is talking. Written straight to the
    // DOM: routing it through state would re-render sixty times a second.
    onAudioLevelChange: level => {
      if (orb.current) orb.current.style.transform = `scale(${1 + level * 0.3})`;
    },
  });

  // The mic stays open across turns, so it is tracked apart from `status`,
  // which moves through listening, thinking and speaking.
  const toggle = () => {
    if (micOn) {
      voice.stopListening();
      voice.interrupt();
    } else {
      voice.startListening(); // From a click: this is where the mic prompt appears.
    }
    setMicOn(!micOn);
  };

  return (
    <>
      <div ref={orb} className="orb" data-status={voice.status} />
      <button onClick={toggle} disabled={!voice.isReady}>
        {voice.isReady ? (micOn ? 'Stop' : 'Talk') : 'Loading…'}
      </button>
    </>
  );
}

For captions, onTranscriptUpdate(text, 'user') gives what the user said and onSpeechStart(text) each sentence of the reply as it starts playing. speak(text) says something without a model turn, such as a greeting, and sendText(text) takes typed input. Everything it returns is in the hook reference.

🚀 Voice mode in production

Each stage runs in the browser or on your backend, and the choice is mostly about the download:

SetupOptionsDownloaded by each visitorAudio leaves the device
Cloud speechonTranscribe + onSubmit + onSynthesizeNothing until the mic opens, then ~4 MB for the voice detectorYes, to your providers
Local speech (the demo)onSubmit~590 MB once, kept for later visits (~320 MB on iPhone)No
Fully localnone~1.3 GB onceNothing leaves at all

Cloud speech is the setup to reach for on phones and on pages people visit once: it is ready as soon as the page is. Local speech keeps what people say on their device and costs nothing per minute, and suits an app they come back to, since the models are stored after the first visit. Fully local is for kiosks and offline use. Use loadModels to hold any download until someone actually engages.

☁️ Cloud Adapters

onTranscribe and onSynthesize replace the local speech models with your own providers. Each replaces its model entirely: supply one and that model is never downloaded. Drop it later and the local model loads, so a provider outage can fall back to the browser mid-session.

To have that fallback ready rather than downloading it at the moment you need it, pass preloadLocalSpeech: the local hearing and voice download behind the conversation, and onLocalSpeechReady fires when they could take over. examples/groq-voice starts on Groq and offers the switch then.

const voice = useAiVoiceAvatar({
  // Instead of Whisper: 16 kHz mono samples of one utterance, to any
  // speech-to-text service. Encode them as WAV if the service wants a file.
  onTranscribe: async (samples: Float32Array) => {
    return await mySpeechToText(samples);
  },

  // Instead of Kokoro: return an encoded MP3/WAV (ArrayBuffer), which the
  // hook decodes, or raw 24 kHz PCM as a Float32Array. Called a sentence at
  // a time, so the first plays while the rest are fetched.
  onSynthesize: async (text: string) => {
    const res = await fetch('/api/tts', { method: 'POST', body: text });
    return await res.arrayBuffer();
  },

  onSubmit: async (text) => fetch('/api/chat', { method: 'POST', body: text }).then(r => r.body),
});

🧑‍💼 Entry 2: The Full 3D Avatar

If you want the full visual presence with 60FPS ARKit lip-syncing, use the drop-in 3D component. Under the hood, this is just a wrapper around the useAiVoiceAvatar hook that procedurally maps the audio to a 3D model!

📦 View Package on the Official NPM Registry ➔

# If using the 3D Avatar, you must also install the Three.js ecosystem
npm install react-ai-voice-avatar three @react-three/fiber @react-three/drei

[!NOTE] React 18 Users: Installing the latest @react-three/drei defaults to version 10, which demands React 19. If your project runs on React 18, install compatible Three.js React bindings explicitly:

npm install @react-three/drei@^9 @react-three/fiber@^8

[!IMPORTANT] React 19.3 and ERESOLVE: @react-three/fiber@9.7 still declares its React peer as >=19 <19.3, so a default npm install alongside React 19.3 or newer fails with ERESOLVE unable to resolve dependency tree. This is a Three.js binding constraint, not a limit of this package: our own peer range accepts React 19.3.

That range is over-cautious. We build and run a Next.js App Router app against React 19.3.0 with @react-three/fiber@9.7.0 and the avatar renders, loads its models and speaks with no errors. So install past it rather than downgrading:

npm install react-ai-voice-avatar three @react-three/fiber @react-three/drei --legacy-peer-deps

If you would rather keep strict peer resolution, pinning React works too:

npm install react@~19.2.0 react-dom@~19.2.0

The headless entry point (react-ai-voice-avatar/headless) pulls in no Three.js at all, so it never hits this and works on any React 18 or 19 version.

⚡ DX & Performance (Lazy Code-Splitting)

To prevent the massive ML assets (WebGPU workers, 3D engines) from bloating your initial page load, use the built-in lazy wrapper. It will automatically code-split the 3D dependencies and render a sleek holographic Skeleton UI while the assets download in the background!

import { AiVoiceAvatarLazy } from 'react-ai-voice-avatar';

// Use it exactly like the normal component!
<AiVoiceAvatarLazy avatarPreset="ananya" />

Show the avatar before anyone commits to a download

Lazy loading defers the 3D engine. The speech and language models are the larger cost, and by default they start downloading as soon as the avatar mounts. On a landing page most visitors only look, so render the avatar idle and load the models when someone actually engages:

const [engaged, setEngaged] = useState(false);

<AiVoiceAvatar loadModels={engaged} hideStatusPill={!engaged} />
<button onClick={() => setEngaged(true)}>Talk to it</button>

The live demo's landing page works this way.

📐 Architectural Best Practices

  • Standard Import (AiVoiceAvatar): Recommended for full-screen applications where the avatar is the primary product (e.g., Kiosks, Digital Tutors). The browser aggressively downloads the 3D canvas and ML models immediately so the avatar is ready instantly.
  • Lazy Import (AiVoiceAvatarLazy): Recommended for widgets, modals, or sub-routes (e.g., a "Support Desk" chat bubble in a SaaS dashboard). Defers downloading the 1.5MB 3D engine and WebWorkers until the user actually opens the widget.

⚙️ Server Configuration (Optional Performance Boost)

The react-ai-voice-avatar engine is truly zero-config. You do not need to configure Vite optimizeDeps, Next.js Webpack overrides, or manually host Web Worker files—everything is dynamically bundled and executed automatically!

However, because our ONNX WebGPU engine leverages modern multi-threaded SharedArrayBuffer memory pipelines for maximum inference speed, your hosting server can optionally emit standard Cross-Origin Isolation HTTP headers (COOP/COEP) to unlock peak performance. If these headers are not present, the engine automatically falls back to single-threaded WebAssembly without crashing.

🌐 Enabling Multi-threading on Production (Vercel, Netlify & Cloudflare)

To unlock multi-threaded performance, specify these isolation headers in your routing manifests:

  • Vercel (vercel.json): Add "headers": [{ "source": "/(.*)", "headers": [{ "key": "Cross-Origin-Opener-Policy", "value": "same-origin" }, { "key": "Cross-Origin-Embedder-Policy", "value": "require-corp" }] }].
  • Netlify / Cloudflare Pages (_headers or netlify.toml): Add /*\n Cross-Origin-Opener-Policy: same-origin\n Cross-Origin-Embedder-Policy: require-corp to public/_headers.

⚡ Enabling Multi-threading in Local Dev (Vite & Next.js)

Vite (vite.config.ts):

export default defineConfig({
  plugins: [react()],
  server: {
    headers: {
      'Cross-Origin-Opener-Policy': 'same-origin',
      'Cross-Origin-Embedder-Policy': 'require-corp',
    },
  },
});

Next.js (next.config.mjs):

export default {
  async headers() {
    return [
      {
        source: '/(.*)',
        headers: [
          { key: 'Cross-Origin-Opener-Policy', value: 'same-origin' },
          { key: 'Cross-Origin-Embedder-Policy', value: 'require-corp' },
        ],
      },
    ];
  },
};

[!CAUTION] Strict CSP Policies: If your enterprise enforces strict Content Security Policies that block blob: workers (worker-src 'self'), you can bypass our zero-config Blob loaders by passing the workerBaseUrl prop to the avatar and hosting the pre-compiled .worker.js files from our dist/assets/ directory yourself.

⚡ 3D Avatar Quickstart

import React, { useRef, useState } from 'react';
import { Canvas } from '@react-three/fiber';
import { OrbitControls } from '@react-three/drei';
import { AiVoiceAvatar, type AiVoiceAvatarHandle } from 'react-ai-voice-avatar';

export function App() {
  const avatarRef = useRef<AiVoiceAvatarHandle>(null);
  const [text, setText] = useState('');

  return (
    <div style={{ width: '100vw', height: '100vh', position: 'relative' }}>
      <Canvas camera={{ position: [0, 0.15, 2.2], fov: 32 }}>
        <color attach="background" args={['#101116']} />
        
        {/* Subtle studio lighting */}
        <pointLight position={[-3, 2, -2]} intensity={25} color="#E67E22" distance={6} />
        <pointLight position={[3, 1, -2]} intensity={20} color="#2980B9" distance={6} />
        
        <OrbitControls target={[0, 0.05, 0]} />
        
        {/* Connected Brain: Zero download, instant initialization! */}
        <AiVoiceAvatar
          ref={avatarRef}
          avatarPreset="ananya"
          lightingPreset="studio"
          ttsEngine="kokoro"
          ttsVoice="af_heart"
          // Connect your backend here (receives user speech transcript):
          onSubmit={async (text) => {
            const res = await fetch('/api/chat', { 
              method: 'POST', 
              body: JSON.stringify({ prompt: text }) 
            });
            return res.body; // Avatar natively reads streams!
          }}
        />
      </Canvas>

      {/* Fallback Text Input for noisy environments */}
      <form
        onSubmit={(e) => {
          e.preventDefault();
          if (text.trim() && avatarRef.current) {
            avatarRef.current.sendText(text);
            setText('');
          }
        }}
        style={{ position: 'absolute', bottom: '20px', left: '50%', transform: 'translateX(-50%)', display: 'flex', gap: '8px', zIndex: 100 }}
      >
        <input 
          value={text} 
          onChange={e => setText(e.target.value)} 
          placeholder="Type a message..." 
          style={{ padding: '8px 16px', borderRadius: '20px', border: 'none', background: 'rgba(255,255,255,0.9)', width: '300px' }}
        />
        <button type="submit" style={{ padding: '8px 16px', borderRadius: '20px', border: 'none', background: '#3b82f6', color: 'white', cursor: 'pointer' }}>
          Send
        </button>
      </form>
    </div>
  );
}

🔌 Connect Your Backend (Recipes)

The onSubmit prop natively accepts a string, an AsyncIterable<string>, or a ReadableStream. To connect your actual backend, simply drop in one of these copy-paste recipes to parse your streaming format!

[!CAUTION] API Keys Belong on the Server! Never put your OpenAI or Anthropic API keys directly in the frontend browser code. Always route through your own backend endpoint (/api/chat).

Recipe 0: Plain Text Stream (Fastest & Simplest)

If your backend uses Vercel AI SDK's streamText(...).toTextStreamResponse() or otherwise streams plain raw text, you can pass the stream natively without any parsing!

onSubmit={async (text) => {
  const res = await fetch('/api/chat', { method: 'POST', body: JSON.stringify({ prompt: text }) });
  return res.body; // Natively supported!
}}

Recipe 1: Vercel AI SDK (≤v4 Data Stream)

Older versions of the Vercel AI SDK stream data using a specific protocol (e.g., 0:"Hello"). This recipe parses those chunks into clean text with a carry-over buffer for safe network boundaries.

onSubmit={async function* (text) {
  const res = await fetch('/api/chat', { method: 'POST', body: JSON.stringify({ prompt: text }) });
  if (!res.body) return;
  const reader = res.body.getReader();
  const decoder = new TextDecoder();
  let buffer = '';
  
  while (true) {
    const { done, value } = await reader.read();
    if (done) break;
    buffer += decoder.decode(value, { stream: true });
    const lines = buffer.split('\n');
    buffer = lines.pop() ?? ''; // keep the trailing partial chunk
    for (const line of lines) {
      if (line.startsWith('0:')) {
        try { yield JSON.parse(line.substring(2)); } catch { /* ignore keep-alive / non-JSON frames */ }
      }
    }
  }
}}

Recipe 2: OpenAI-Compatible SSE Endpoint (and AI SDK v5)

Standard Server-Sent Events (SSE) stream data: {...} blocks. This handles safe parsing across broken network chunk boundaries.

onSubmit={async function* (text) {
  const res = await fetch('/api/chat', { method: 'POST', body: JSON.stringify({ prompt: text }) });
  if (!res.body) return;
  const reader = res.body.getReader();
  const decoder = new TextDecoder();
  let buffer = '';
  
  while (true) {
    const { done, value } = await reader.read();
    if (done) break;
    buffer += decoder.decode(value, { stream: true });
    const lines = buffer.split('\n');
    buffer = lines.pop() ?? ''; // keep the trailing partial chunk
    for (const line of lines) {
      if (line.startsWith('data: ') && line !== 'data: [DONE]') {
        try {
          const parsed = JSON.parse(line.substring(6));
          // AI SDK v5 emits {type:'text-delta', delta:'...'}; OpenAI emits choices[0].delta.content
          if (parsed.type === 'text-delta' && parsed.delta) {
            yield parsed.delta;
          } else if (parsed.choices?.[0]?.delta?.content) {
            yield parsed.choices[0].delta.content;
          }
        } catch { /* ignore keep-alive / non-JSON frames */ }
      }
    }
  }
}}

🗣️ Languages

English and Hindi both run entirely on the device. Set ttsLanguage and the engine picks a matching voice, a matching speech-recognition model, and a matching phoneme path.

<AiVoiceAvatar ttsLanguage="hi-IN" />

Hindi needed real work rather than a config flag. Kokoro ships four Hindi voices inside the same checkpoint as the English ones, but the JavaScript wrapper does not list them, and the bundled eSpeak build carries English data only and rejects hi outright. So this package includes its own Devanagari-to-phoneme converter. Devanagari is close to phonemic, which makes that tractable; the hard part is schwa deletion, the rule that makes कमल read as "kamal" rather than "kamala", and getting it wrong produces speech that still sounds like speech while sounding like someone spelling Hindi out.

Because the converter emits real phonemes, Hindi gets phoneme-driven lip-sync with distinct mouth shapes for the retroflex consonants, not the amplitude-only mouth flapping most engines fall back to outside English. Code-switching works too: English words inside a Hindi sentence are routed to the English phonemiser, so "मुझे coffee चाहिए" is pronounced correctly throughout.

[!NOTE] Local Hindi is a demo, not a product. Models small enough to run in a browser are far weaker in Hindi than in English. It is good enough to show the pipeline working end to end and not good enough to ship. For production Hindi, route hearing and thinking to an API through onTranscribe and onSubmit; the voice and lip-sync stay local and are genuinely good.

A session is one language at a time. Someone who switches language mid-conversation will be mistranscribed, because the recognition model is told which language to expect and browser Whisper cannot detect it.


🎤 Taking turns

The user can talk over the avatar and cut it off mid-sentence. This is on by default, because waiting for a reply to finish is the thing that makes a voice agent feel like a walkie-talkie.

<AiVoiceAvatar
  allowInterruption={false}   // default: true
  onUserInterrupt={() => analytics.track('barge_in')}
/>

Turn it off for a kiosk or a noisy room, where the avatar hearing its own voice through the speakers and stopping itself is worse than waiting. It is ignored in push-to-talk, which owns the floor explicitly.

How a turn is decided

Voice detection scores every 96ms frame for how much it sounds like speech. A cough, a door and a chair all clear a loudness bar as easily as a word does, so loudness alone cannot separate them — sustain can. A sound has to keep scoring as speech for minSpeechMs before the engine treats it as a turn.

Below that bar, a sound only ducks the avatar's voice, reversibly. If it turns out to be a cough the reply resumes from where it paused, at the right place in the sentence and with the mouth still in sync. Nothing about the conversation changed, because nothing was decided on a noise.

Tuning for your room

The defaults suit a quiet room and headphones. A shop floor is a different problem, and only you know which you have.

<AiVoiceAvatar
  speechDetection={{ positiveSpeechThreshold: 0.6, minSpeechMs: 700 }}
/>
FieldDefaultRaise it whenLower it when
positiveSpeechThreshold0.5Passing noise is mistaken for talkingQuiet speakers go unheard
negativeSpeechThreshold0.35Turns end too slowly in a noisy roomTurns end while someone is still talking
minSpeechMs500Short noises still start turnsSingle-word answers are ignored
redemptionMs1400People are cut off while thinking mid-sentenceReplies feel slow to start
preSpeechPadMs800The first word is still being clippedRarely — this is the audio kept from before the trigger, and it is what stops the first syllable going missing

Two costs worth knowing before you change anything. A speaker who sits below positiveSpeechThreshold produces no events at all — raising it trades quiet voices for quiet rooms. And a filler like "hmm" held long enough to pass minSpeechMs still reaches transcription; a denylist catches the common ones after the fact, but sustain cannot tell a long "hmm" from a short word.


🚨 Handling failures

Pass onError and your app learns when something breaks, rather than finding out from a console message it cannot see.

<AiVoiceAvatar
  onError={(e) => {
    if (e.severity === 'fatal') showFallbackUI(e.stage);
    logToSentry(e);
  }}
/>

Check severity before reacting. Most failures here are survivable because the engine falls back: WebGPU to WASM, Kokoro to a smaller voice model. Those arrive as degraded and the avatar still works, so treating them as fatal would hide a working experience behind an error screen. A refused microphone is fatal for listening while typed input still works, which is a judgement only your app can make.

stage is one of microphone, speech-recognition, language-model, speech-synthesis, audio-output, worker, model-storage or conversation. There is also a detail string carrying the engine's internal stage name for bug reports; it is not stable across versions, so do not branch on it.

model-storage is always degraded: a model loaded and works, but could not be kept, so the next visit downloads it again. The message says why — usually not enough free space, which on a phone is the common case.

Models are kept between visits

Downloaded models are stored in the browser's Origin Private File System, so a returning visitor skips the download: on an Apple M4 with WebGPU, a return visit is ready to talk in about 2 seconds, with nothing fetched. The browser's Cache API — what the model loader uses by default — refuses any single file of 256 MiB or more, and both the voice (~310 MB) and the local language model (~750 MB) are larger than that, so before this they were downloaded again on every visit. Where OPFS is unavailable the Cache API is used as before.

Browsers can clear this storage when disk space runs low. If your app is one people return to, call navigator.storage.persist() from a user gesture to ask the browser to keep it; the engine does not, because Firefox can answer that call with a permission prompt, and a library should not put one in front of your visitors unasked.


🎨 Bring Your Own 3D Avatar (Custom GLB)

You are not locked into our built-in avatars (ananya and aarav)! You can use any custom .glb humanoid model by passing its URL or local path to the modelSrc prop:

<AiVoiceAvatar
  modelSrc="/models/my-custom-avatar.glb"
  // ...
/>

📋 Custom Avatar Requirements

To ensure the lip-sync and procedural facial dynamics engines work correctly, your custom model must meet the following standard requirements:

  1. Format: .glb (GLTF Binary).
  2. Facial Blendshapes (Morph Targets): The model's head/face mesh must contain the standard 52 Apple ARKit blendshapes (e.g., jawOpen, eyeBlinkLeft, mouthSmileRight). Our engine automatically traverses your model to find these targets.
  3. Bone Naming: For the interactive mouse-tracking and head-tilting physics to function, the armature should use standard bone names (e.g., a neck/head bone named Head, head, Neck, or neck).

Where the built-in avatars come from

ananya and aarav are converted from the Microsoft Rocketbox library, which is MIT licensed like this project, so you can redistribute them without restriction. The library has 115 avatars and any of them can be converted with scripts/convert-rocketbox.py. See assets/avatars/LICENSE.md.

Models exported from Ready Player Me also work, since they carry the same ARKit blendshapes and bone names. Note that Ready Player Me shut down in January 2026, so you can no longer create new avatars there, and existing exports are licensed CC BY-NC-SA rather than MIT.


🏗️ What runs where

Four stages, and you choose where each one happens. The defaults are all local, which is why the demo needs no keys, but the interesting production setups are mixed.

flowchart TD
  mic(["🎙️ You speak"]) --> vad["Voice detector<br/><i>in the browser</i>"]
  vad --> hear["Hearing<br/><i>Whisper in the browser</i><br/>or your onTranscribe"]
  hear --> think["Thinking<br/><i>a model in the browser</i><br/>or your onSubmit"]
  think --> speak["Speaking<br/><i>Kokoro in the browser</i><br/>or your onSynthesize"]
  speak --> face["🔊 Voice, with a lip-synced<br/>3D face or your own UI"]
  vad -. "talk over it and it stops" .-> speak

Each stage runs in the browser unless you pass the adapter beside it, which hands that stage to your own service.

StageOn the deviceYour backend instead
Hearing — speech to textWhisper via ONNXonTranscribe
Thinking — the replyQwen or Gemma via WebGPUonSubmit
Speaking — text to audioKokoro-82MonSynthesize
Face — lip-sync and animationAlways hereNot applicable

The face never leaves the device, which is the whole point: that is what avatar APIs charge per minute for, and it is the one stage that cannot be outsourced without a video stream.

The common production shape is onSubmit alone. Hearing and speaking stay local, so no audio ever leaves the browser, while generation goes to whatever model you already run. That keeps the download to about 600 MB (Whisper base and the Kokoro voice), keeps your keys on your server, and still gives you a conversation nobody else can read.

Add onTranscribe and onSynthesize as well and nothing downloads at all, apart from the ~4 MB voice detector once the microphone opens. Each adapter replaces its model entirely, so the page is ready as soon as it loads. That is the shape for phones and for pages people visit once.

Fully local is real, not a demo trick, and it is the right answer for a kiosk, a regulated environment, or anywhere without reliable connectivity. Be aware of the cost: in English the first visit downloads roughly 1.3 GB before anyone can speak (the language model alone is 750 MB), about 2.1 GB in Hindi, and the quality ceiling is whatever a model that size can do.

⏱️ How fast it answers

Measured with scripts/measure-latency.mjs on an Apple M4 with WebGPU in Chromium: a spoken question through a fake microphone, the median of five turns after a warm-up turn, timed by the hook's own callbacks.

StageTime
Waiting for silence, to be sure you've finished1,400 ms (redemptionMs, adjustable)
Hearing: Whisper base transcribes the utterance330 ms
Thinking, hosted: GPT-OSS 20B on Groq, the demo's route, one round trip670 ms
Thinking and the first sentence of voice, all in the browser (Qwen 0.5B, Kokoro)740 ms
Speaking: Kokoro's first sentence ready and playing, with an instant reply460 ms
Starting up on a return visit, models already storedabout 2 s, nothing downloaded

So from the moment you stop talking to the first word of the reply:

SetupFirst word after you stop
All in the browserabout 2.5 s
Hosted replies, local hearing and voice (the demo)about 2.9 s
With an instant reply, the floor for speech aloneabout 2.2 s

The hosted figure adds the route's round trip to the measured speech stages, since that route answers only the live site. Most of the wait is the pause for silence, which is what stops the avatar cutting people off while they think. Lower redemptionMs in speechDetection for a snappier feel, at the cost of answering half-finished sentences. Without WebGPU, recognition and the voice run on the CPU and are several times slower.

🔀 Hosted now, local when warm

You do not have to choose once. preloadLocalLlm fetches the in-browser model behind the conversation while your backend answers, and onLocalLlmReady tells you when it can take over:

const [localReady, setLocalReady] = useState(false);

<AiVoiceAvatar
  // Dropping onSubmit is what hands the conversation over.
  onSubmit={localReady ? undefined : askMyBackend}
  preloadLocalLlm
  onLocalLlmReady={() => setLocalReady(true)}
/>

Nobody waits for a gigabyte to say the first word, and the turns your backend answered are handed to the local model, so it does not restart the conversation from nothing. Useful for a kiosk that must keep working when the wifi drops, a demo on someone else's quota, or a phone, which will never accept the download but can hold a conversation the moment it loads.

The live demo runs exactly this: replies come from a hosted model, and on a desktop with WebGPU the in-browser model downloads during the conversation and takes the floor when it lands. The page says which one is answering at any moment.

Explore the canonical patterns in the examples/ directory:

Example PatternFolderHighlights & Architecture
Live Interactive DemosandboxDeploy on Vercel ➔ — Our full-featured interactive testbed featuring live character switching (ananya, aarav), voice persona switching (af_heart, am_michael), real-time diagnostic probe metrics, and Leva 3D lighting controls.
Quickstartexamples/quickstartMinimal, zero-configuration plug-and-play AI voice avatar deployment with built-in studio lighting & sizing.
Local Kioskexamples/local-kiosk100% offline on-device retail & restaurant ordering kiosk with embedded menu reasoning. Demonstrates the On-Device Brain; operates without internet access once model weights are locally cached.
Connected Appexamples/hybrid-cloudIllustrates the Connected Brain (onSubmit). Bypasses gigabyte-scale local LLM downloads by routing reasoning to OpenAI, Claude, or corporate APIs while keeping ASR, TTS, and 3D lip blending 100% on-device!
Voice Onlyexamples/voice-onlyVoice mode with no avatar, like ChatGPT or Gemini voice: the react-ai-voice-avatar/headless hook, an audio-reactive orb, live captions and interruption. No three.js in the bundle, and no backend needed to try it.
Groq Voiceexamples/groq-voiceVoice mode on Groq's free tier with your own key: hearing, replies and voice start on Groq with nothing to download, the in-browser hearing and voice download behind the conversation (preloadLocalSpeech), and the page offers to switch when they are ready. Shows Groq's live limits.
Headless Custom UIexamples/headless-custom-uiStill renders the 3D avatar; for no avatar at all, see Voice Only above. Demonstrates hiding built-in DOM overlays (hideStatusPill={true}, showCaptions={false}), streaming transcripts into a custom enterprise UI, and controlling voice outputs imperatively via ref.current?.speak(text).

📖 API Reference

<AiVoiceAvatar /> Props

PropTypeDefaultDescription
avatarPreset'ananya' | 'aarav' | 'default' | 'kiosk''ananya'Built-in 3D character models featuring both female ('ananya') and male ('aarav') voice concierges out of the box with full ARKit facial blendshapes!
avatarSize'sm' | 'md' | 'lg' | number'md' (0.48)Intuitive model sizing presets or custom decimal scaling multiplier applied directly to the 3D humanoid mesh.
modelSrcstringundefinedAbsolute local path or remote URL to a custom GLTF/GLB humanoid armature avatar model.
lightingPreset'studio' | 'cyberpunk_violet' | 'cool_azure' | 'warm_amber' | 'clean_white' | 'none''studio'Pre-built cinematic studio lighting atmospheres directly applied to your 3D viewport without manual Three.js configuration!
systemPromptstring"You are Ananya..."Conversational persona directives and context injected into active LLMs.
llmModelstringby languageHugging Face id for the local WebGPU reasoning model, used only when onSubmit is absent. Defaults to Qwen2.5-0.5B for English and Gemma 3 1B for Hindi, which Qwen that size cannot speak. Set it to pin one model for every language.
asrModelstringby languageHugging Face id for the local Whisper model. Defaults to Whisper base for English and Whisper small for Hindi, which base transcribes badly. Pass "Xenova/whisper-tiny" for a faster download and worse accuracy.
ttsEngine'kokoro' | 'mms''kokoro'High-fidelity neural voice synthesis engine executing inside dedicated Web Workers.
ttsVoicestringby languageKokoro voice id. af_* and am_* American, bf_* and bm_* British, hf_* and hm_* Hindi (af_heart, am_michael, bf_emma, hf_alpha, hm_omega). Defaults to one matching ttsLanguage. A voice whose language disagrees is corrected with a warning.
ttsLanguage'en-US' | 'en-GB' | 'hi-IN''en-US'Conversation language. Selects the voice, the recognition model and the phoneme path. Also sets asrLanguage unless you set that yourself. For any other language, pass onSynthesize and use a cloud voice provider.
asrLanguagestringttsLanguageLanguage to transcribe. Follows ttsLanguage by default, since a conversation is almost always held in one language.
showCaptionsbooleantrueRenders a sleek glassmorphic subtitle overlay displaying spoken interaction dialog.
hideStatusPillbooleanfalseWhen true, suppresses the default bottom-left microphone interactive control pill.
listenMode'continuous' | 'push-to-talk''continuous'continuous keeps the mic hot after the avatar finishes speaking naturally, but explicitly clicking Stop forces it off until tapped again. push-to-talk strictly requires manually tapping to start listening for every single turn.
loadModelsbooleantrueSet false to render the avatar without downloading any models, then flip it true when the visitor engages. For landing pages and widgets most visitors never talk to. Status stays 'loading' until it is true and the models are up, so show your own call to action meanwhile.
onModelLoaded() => voidundefinedFires once the 3D mesh is parsed and in the scene. Parsing a multi-megabyte GLB leaves the canvas empty for a few seconds; use this to hold a placeholder over it.
allowInterruptionbooleantrueLets the user talk over the avatar and cut it off mid-sentence. Turn off for a kiosk or noisy room, where the avatar hearing itself through the speakers is worse than waiting. Ignored in push-to-talk. See Taking turns.
speechDetection{ positiveSpeechThreshold?, negativeSpeechThreshold?, minSpeechMs?, redemptionMs?, preSpeechPadMs? }see Taking turnsTunes how the microphone decides someone is talking. Every field optional. The defaults suit a quiet room; a shop floor needs a higher threshold and a longer minSpeechMs.
onUserInterrupt() => voidundefinedFires when the user talks over the avatar and takes the floor. Only fires if the avatar actually had audio playing.
gesturesboolean | numbertrueHand and arm gestures while the avatar speaks: one hand or both brought up in front of the chest for each phrase, with small beats on stressed syllables, and the arms back at rest when it stops. false keeps the arms still; a number sets the size, 0 to 1.5. Needs a skeleton with LeftArm, LeftForeArm and LeftHand and the right-hand equivalents; custom avatars without them simply do not gesture.
onAudioLevelChange(level: number, source: 'mic' | 'tts' | 'idle') => voidundefinedReal-time audio amplitude (0-1) callbacks for the active stream. Essential for building highly responsive, audio-reactive 3D Visualizers and HUDs! Fires with 0 and 'idle' between turns, so a meter falls to rest rather than freezing.
onSubmit(text: string) => Promise<string | AsyncIterable<string> | ReadableStream>undefinedConnected Brain API: Bypasses local LLMs; routes transcribed user microphone strings to your cloud or custom LLM API endpoint.
preloadLocalLlmbooleanfalseOnly meaningful alongside onSubmit, which otherwise skips the local language model download entirely. Set it to fetch that model in the background while your hosted one answers, so the conversation survives a rate limit, an expired quota or a lost network. See Hosted now, local when warm.
onLocalLlmReady() => voidundefinedFires once the model requested by preloadLocalLlm has loaded. Drop onSubmit here to hand the conversation over; the turns your backend answered are carried across, so the local model knows what was already said.
onTranscribe(audio: Float32Array) => Promise<string>undefinedReplaces local speech recognition with your own service, and Whisper is then not downloaded. Receives one utterance as 16 kHz mono samples.
onSynthesize(text: string) => Promise<Float32Array | ArrayBuffer>undefinedReplaces local voice synthesis with your own service, and Kokoro is then not downloaded. Return an encoded MP3/WAV buffer, or raw 24 kHz PCM. Called a sentence at a time.
onError(e: AiVoiceAvatarError) => voidundefinedFires when a stage fails. Carries stage, message and a severity of degraded or fatal. See Handling failures.
onTranscriptUpdate(text: string, speaker: 'user' | 'avatar') => voidundefinedCallback delivering real-time microphone transcriptions and assistant spoken utterance strings.
onStatusChange(status: string) => voidundefinedEmits live state transitions (loading, idle, listening, thinking, speaking).
debugbooleanfalseWhen true, renders an interactive floating GUI (Leva) to inspect and tune individual 3D blendshapes.
vadAssetPathstringundefinedOptional URL or local path override for self-hosting @ricky0123/vad-web ONNX asset binaries in airgapped deployments.
onnxWasmPathstringundefinedOptional URL override for self-hosting onnxruntime-web WASM distribution files.
workerBaseUrlstringundefinedCSP Escape Hatch: if blob: workers are blocked by your server, fetch pre-compiled Web Workers from this URL directory.
enableLocalAssetProbebooleanfalseWhen true, HEAD-checks /ananya.glb in your own public directory before falling back to the CDN. Off by default: with no local copy the probe 404s, and that 404 lands in every visitor's console.
statusPillStyleReact.CSSPropertiesundefinedOptional custom CSS styling & absolute positioning overrides for the interactive Status Pill overlay.
accentColorstringundefinedCustom CSS color string (e.g., #38BDF8) for the active status indicator rings and highlights.

Imperative Ref API (AiVoiceAvatarHandle)

Attach a React ref (useRef<AiVoiceAvatarHandle>(null)) to access imperative real-time controls:

interface AiVoiceAvatarHandle {
  /** Command the 3D avatar to speak an arbitrary string with synchronized acoustic lip blending */
  speak: (text: string) => void;
  /** Manually engage microphone recording and Voice Activity Detection (VAD) */
  startListening: () => void;
  /** Pause active microphone listening */
  stopListening: () => void;
  /** Instantly interrupt and halt active voice speech synthesis and clear the audio queue */
  interrupt: () => void;
  /** Manually submit text to the onSubmit handler, simulating a spoken utterance (useful for text-only fallback) */
  sendText: (text: string) => void;
  /** Wipe multi-turn conversation memory history and caption overlay states */
  clearHistory: () => void;
  /** Retrieve live Web Audio API AnalyserNode powering real-time spectral lip sync */
  getAnalyser: () => AnalyserNode | undefined;
}

💬 Text-Only Input (sendText)

If your users cannot use a microphone (e.g., noisy environments, privacy concerns, or lack of permissions), you can easily wire up a standard text input field to bypass the speech recognition pipeline entirely!

Simply attach a ref and call sendText() to pass a string directly to your onSubmit handler (or local LLM):

const avatarRef = useRef<AiVoiceAvatarHandle>(null);

// In your UI, attach this to a standard <form> submission:
const handleTextSubmit = (userInput: string) => {
  avatarRef.current?.sendText(userInput);
}

When you use sendText, the avatar immediately enters the thinking state and processes the interaction exactly as if the user had spoken it aloud.

useAiVoiceAvatar() (headless)

import { useAiVoiceAvatar } from 'react-ai-voice-avatar/headless';

Takes the same options as the component's props, except the ones about the 3D scene and its overlays (avatarPreset, avatarSize, modelSrc, lightingPreset, gestures, showCaptions, hideStatusPill, onModelLoaded, debug and the styling props). onStatusChange is replaced by the returned status. A few options matter mostly without an avatar:

OptionTypeDescription
onAudioLevelChange(level: number, source: 'mic' | 'tts' | 'idle') => voidLoudness from 0 to 1, every frame, from whichever side is talking. What an orb or waveform animates from. Write it to the DOM through a ref rather than into state.
onSpeechStart(text: string) => voidEach sentence of the reply as it starts playing. Captions that keep pace with the voice.
onTranscriptUpdate(text: string, speaker: 'user' | 'avatar') => voidWhat the user said, once transcribed, and the reply in full.
onInferenceStart / onInferenceEnd() => voidAround each turn, from the moment the user stops talking to the end of the reply.
onTtsEngineChange(engine: 'kokoro' | 'mms' | 'custom') => voidFires whenever activeTtsEngine changes. Also a prop on the component.
preloadLocalSpeechbooleanDownload the in-browser hearing and voice behind onTranscribe and onSynthesize, without waiting for them. Drop the adapters once onLocalSpeechReady fires and the local models take over. Also a prop on the component.
onLocalSpeechReady() => voidFires once, when the models requested by preloadLocalSpeech are loaded.
loadingProgress(pct: number, label: string) => voidDownload progress per model: 'asr', 'kokoro' (or 'tts' for MMS) and 'llm'. They download in parallel, so keep one figure per label.

It returns:

ValueTypeDescription
status'loading' | 'idle' | 'listening' | 'thinking' | 'speaking'Where the conversation is. 'loading' until the models are up, and until loadModels is true.
isLoading, isIdle, isListening, isThinking, isSpeakingbooleanShorthands for status.
isReadybooleanThe models are loaded. startListening does nothing before this.
isLocalSpeechReadybooleanThe in-browser hearing and voice are loaded, whether or not they are in use.
activeTtsEngine'kokoro' | 'mms' | 'custom'The voice actually speaking. iPhones and iPads get 'mms', a single plainer voice that ignores ttsVoice, because Kokoro runs Safari out of memory; so does any device where Kokoro fails to load. 'custom' is your onSynthesize. Worth showing if your users will compare devices.
startListening() => Promise<void>Opens the microphone and starts listening. Call it from a click: the first call is where the browser asks for permission. In 'continuous' mode the microphone then stays open across turns.
stopListening() => voidCloses the microphone.
interrupt() => voidStops the reply mid-sentence and clears what was queued.
speak(text: string) => voidSays the text without a model turn: a greeting, a notification. Plays a sentence at a time.
sendText(text: string) => voidTyped input, handled as if it had been spoken.
clearHistory() => voidForgets the conversation so far.
micErrorstring | nullWhy the microphone could not open, such as a denied permission. null otherwise.
analyserAnalyserNode | undefinedThe reply's audio, for a frequency visualiser. onAudioLevelChange is simpler if you only need loudness.

It also returns several refs (currentSpeechTextRef, audioContextRef and others) that the 3D component uses for lip-sync. A voice UI can ignore them.


🌐 Performance & Asset Caching

  1. Native WebGPU & WASM Degradation:
    • Modern Chromium browsers (Chrome, Edge, Opera, Arc) on desktop and mobile platforms benefit from hardware-accelerated WebGPU neural execution.
    • On systems without WebGPU, inference automatically falls back to multi-threaded WebAssembly (WASM) quantization without app crashes.
  2. Persistent Local Caching:
    • AI models (Whisper, Kokoro, the local language model) are downloaded once and kept in the browser's origin private file system, so later visits skip the download. See Models are kept between visits.

🤝 Contributing & Open Issues Roadmap

We actively welcome community contributions. CONTRIBUTING.md has the local development guide, and ROADMAP.md has what is worth doing next, why, and what is already known about each item — including the measurements behind the open questions.

The largest pieces currently open:

  1. 🖐️ Open-palm gestures. Hands now gesture while the avatar speaks, but always with the palms facing inward. Turning them up — the open, offering gesture people use when explaining — needs the forearm's roll, which has to be worked out from the finger bones rather than assumed, for the same reason as everything else in armRig.ts.
  2. 🎤 Turn-taking in real rooms. Turns are tuned against one speaker in a quiet room. A short "yes" can be dropped, and a quiet speaker can go unheard. Recordings from more voices and noisier rooms would let the defaults be set against something other than one person.
  3. 🎭 Expanding regional 3D avatar personas. Ananya and Aarav ship out of the box. Royalty-free character GLBs (~3MB) rigged with the standard 52 Apple ARKit facial blendshapes are welcome. Avatar meshes are served from a CDN rather than bundled, so adding one adds nothing to the npm install (2.3 MB tarball, 6.0 MB unpacked, asserted in CI by scripts/verify-pack.mjs).
  4. 🙌 Gestures. The avatar stands still while it speaks. Hand and arm movement tied to speech is the most visible thing still missing.
  5. 📱 React Native / Expo support. Exploring bindings to run ONNX inference and Three.js on mobile runtimes.

Hindi speech and VAD sensitivity tuning (speechDetection) were previously listed here and have both shipped.


🧭 Browser Compatibility Matrix

This library heavily relies on modern Web APIs (WebGPU, WebGL, Web Audio, and Web Workers). It gracefully degrades when certain APIs are unavailable.

BrowserOS3D Rendering (WebGL)Voice Synthesis (WebGPU/WASM)Voice Recognition (Web Audio)Status
Chrome / EdgeWindows, macOS, Android✅ Native✅ WebGPU (Ultra Fast)✅ Native🟢 Tier 1 (Recommended)
Safari / iOSmacOS, iOS✅ Native🔄 Lightweight Models by Default✅ Native🟡 Supported
FirefoxWindows, macOS✅ Native⚠️ WASM Fallback✅ Native🟡 Tier 2 (Slower TTS)

[!NOTE]

  • WebGPU is currently enabled by default in Chrome/Edge. On browsers without WebGPU, the library automatically falls back to WASM execution.
  • iOS/Safari Preemptive Fallback: Safari and iOS impose strict memory limits that cause 80MB+ models (like Kokoro) to crash the tab. The engine automatically preempts this by forcing the lightweight MMS TTS model (~30MB) on iOS devices, and tracks crash breadcrumbs to prevent OOM reload loops.
  • Strict CSP Environments: Safari and Firefox may block blob: worker execution depending on your Content-Security-Policy headers. If this occurs, host the .worker.js files statically and pass their base path via the workerBaseUrl prop.

💻 Hardware Requirements

Running Neural Networks in the browser requires capable hardware.

Deployment ModeMin RAMGPU RequirementRecommended Devices
Connected Brain (ASR + TTS only)4GBNone (WASM Fallback ok)iPhone 11+, Mid-range Android (2021+), Any Laptop
Full Local AI (ASR + 500M LLM + TTS)8GBWebGPU Support PreferrediPhone 13 Pro+, High-end Android (Snapdragon 8 Gen 1+), M1/M2 Macs, Modern PCs

[!TIP] Mobile Memory Limits: Mobile browsers rigidly enforce memory limits per tab (often terminating tabs exceeding ~1GB). If your mobile app crashes "after some time", ensure you are utilizing the Connected Brain mode (onSubmit API) which offloads the heavy LLM memory footprint to your server while keeping ultra-fast lip-sync and TTS local.


📜 License

MIT © React AI Voice Avatar Contributors.

ai-avatar
conversational-ai
headless
heygen-alternative
hindi
kokoro
lipsync
onnx
r3f
react
react-hook
speech-to-speech
threejs
tts
virtual-assistant
voice-ai
voice-assistant
voice-mode
webgpu
whisper

Languages

TypeScript

74.9%

JavaScript

19.9%

Python

3.5%