Open-source alternative to real-time avatar APIs like HeyGen Interactive Avatar and Tavus. A talking, lip-synced 3D avatar that runs on your user's GPU, driven by your own LLM. No video stream, no per-minute billing. Or voice mode alone, ChatGPT-style, as a headless React hook. 📦 npm: react-ai-voice-avatar
See the codereact-ai-voice-avatar) 🚀🗣️🧬The open-source alternative to real-time avatar APIs, as a React component.
HeyGen Interactive Avatar, Tavus and Soul Machines sell a talking avatar that holds a conversation, rendered on their servers and billed per streaming minute. This does the same job as an npm install: the avatar renders on your user's GPU, so there is no video stream, no per-minute cost, and no third party in the middle of your conversations.
You bring the model. Point it at OpenAI, Anthropic, your own fine-tune or your existing chat endpoint, and keep your keys on your own backend. The package owns the parts that are tedious to build and easy to get wrong: microphone capture, knowing when someone has finished speaking, interrupting the avatar mid-sentence when they talk over it, streaming speech synthesis, and driving 52 ARKit facial blendshapes at 60 FPS so the mouth matches the words.
It can also run with no backend at all. Speech recognition, generation and voice all have in-browser implementations, which makes for a convincing demo and a genuinely offline kiosk. Most production apps will use their own model and keep only speech and lip-sync on the device.
Don't need a face? The same engine is a headless React hook for voice mode, the way ChatGPT and Gemini do it: react-ai-voice-avatar/headless, with no three.js in your bundle. See it ➔ · How ➔
Speech recognition, voice synthesis and lip-sync, all running in your browser tab.
The first visit downloads about 590 MB of models, roughly a minute and a half on a
50 Mbit/s connection. Your browser keeps them, so a return visit is talking in
about 2 seconds with nothing downloaded. Replies on the demo come from a
hosted model on Groq's free tier, shared by everyone trying it, and start about
3 seconds after you stop talking; on a busy day that tier can run out, and
examples/groq-voice
runs the same thing on your own free key.
![]()
react-ai-voice-avatar provides two distinct ways to integrate into your app depending on your design needs. Both share the exact same underlying conversational state machine, Voice Activity Detection (VAD), and turn-taking logic.
Voice mode for your app: talk to it the way you talk to ChatGPT or Gemini, and interrupt it mid-sentence. The hook owns the microphone, knowing when someone has finished speaking, interruption, and streaming the reply into speech. You own the UI, and every frame it hands you a loudness level to animate.
![]()
Try it ➔: a page built on this hook alone. examples/voice-only is the same page as an app to copy.
Import it from react-ai-voice-avatar/headless and Three.js never enters your module graph. Measured on the same Next.js App Router build, one route rendering the 3D avatar and one rendering only the hook:
| Route | First Load JS |
|---|---|
| 3D avatar | 389 kB |
| Headless hook | 117 kB |
npm install react-ai-voice-avatar
import { useRef, useState } from 'react';
// The /headless subpath is what keeps Three.js out of your bundle.
// Importing the hook from the package root pulls the 3D stack in with it.
import { useAiVoiceAvatar } from 'react-ai-voice-avatar/headless';
export function VoiceMode() {
const orb = useRef<HTMLDivElement>(null);
const [micOn, setMicOn] = useState(false);
const voice = useAiVoiceAvatar({
// Your model. Return a string, or a stream so speech starts sooner.
onSubmit: text => fetch('/api/chat', { method: 'POST', body: text }).then(r => r.body),
// 0 to 1 every frame, from whoever is talking. Written straight to the
// DOM: routing it through state would re-render sixty times a second.
onAudioLevelChange: level => {
if (orb.current) orb.current.style.transform = `scale(${1 + level * 0.3})`;
},
});
// The mic stays open across turns, so it is tracked apart from `status`,
// which moves through listening, thinking and speaking.
const toggle = () => {
if (micOn) {
voice.stopListening();
voice.interrupt();
} else {
voice.startListening(); // From a click: this is where the mic prompt appears.
}
setMicOn(!micOn);
};
return (
<>
<div ref={orb} className="orb" data-status={voice.status} />
<button onClick={toggle} disabled={!voice.isReady}>
{voice.isReady ? (micOn ? 'Stop' : 'Talk') : 'Loading…'}
</button>
</>
);
}
For captions, onTranscriptUpdate(text, 'user') gives what the user said and onSpeechStart(text) each sentence of the reply as it starts playing. speak(text) says something without a model turn, such as a greeting, and sendText(text) takes typed input. Everything it returns is in the hook reference.
Each stage runs in the browser or on your backend, and the choice is mostly about the download:
| Setup | Options | Downloaded by each visitor | Audio leaves the device |
|---|---|---|---|
| Cloud speech | onTranscribe + onSubmit + onSynthesize | Nothing until the mic opens, then ~4 MB for the voice detector | Yes, to your providers |
| Local speech (the demo) | onSubmit | ~590 MB once, kept for later visits (~320 MB on iPhone) | No |
| Fully local | none | ~1.3 GB once | Nothing leaves at all |
Cloud speech is the setup to reach for on phones and on pages people visit once: it is ready as soon as the page is. Local speech keeps what people say on their device and costs nothing per minute, and suits an app they come back to, since the models are stored after the first visit. Fully local is for kiosks and offline use. Use loadModels to hold any download until someone actually engages.
onTranscribe and onSynthesize replace the local speech models with your own providers. Each replaces its model entirely: supply one and that model is never downloaded. Drop it later and the local model loads, so a provider outage can fall back to the browser mid-session.
To have that fallback ready rather than downloading it at the moment you need it, pass preloadLocalSpeech: the local hearing and voice download behind the conversation, and onLocalSpeechReady fires when they could take over. examples/groq-voice starts on Groq and offers the switch then.
const voice = useAiVoiceAvatar({
// Instead of Whisper: 16 kHz mono samples of one utterance, to any
// speech-to-text service. Encode them as WAV if the service wants a file.
onTranscribe: async (samples: Float32Array) => {
return await mySpeechToText(samples);
},
// Instead of Kokoro: return an encoded MP3/WAV (ArrayBuffer), which the
// hook decodes, or raw 24 kHz PCM as a Float32Array. Called a sentence at
// a time, so the first plays while the rest are fetched.
onSynthesize: async (text: string) => {
const res = await fetch('/api/tts', { method: 'POST', body: text });
return await res.arrayBuffer();
},
onSubmit: async (text) => fetch('/api/chat', { method: 'POST', body: text }).then(r => r.body),
});
If you want the full visual presence with 60FPS ARKit lip-syncing, use the drop-in 3D component. Under the hood, this is just a wrapper around the useAiVoiceAvatar hook that procedurally maps the audio to a 3D model!
📦 View Package on the Official NPM Registry ➔
# If using the 3D Avatar, you must also install the Three.js ecosystem
npm install react-ai-voice-avatar three @react-three/fiber @react-three/drei
[!NOTE] React 18 Users: Installing the latest
@react-three/dreidefaults to version 10, which demands React 19. If your project runs on React 18, install compatible Three.js React bindings explicitly:npm install @react-three/drei@^9 @react-three/fiber@^8
[!IMPORTANT] React 19.3 and
ERESOLVE:@react-three/fiber@9.7still declares its React peer as>=19 <19.3, so a defaultnpm installalongside React 19.3 or newer fails withERESOLVE unable to resolve dependency tree. This is a Three.js binding constraint, not a limit of this package: our own peer range accepts React 19.3.That range is over-cautious. We build and run a Next.js App Router app against React 19.3.0 with
@react-three/fiber@9.7.0and the avatar renders, loads its models and speaks with no errors. So install past it rather than downgrading:npm install react-ai-voice-avatar three @react-three/fiber @react-three/drei --legacy-peer-depsIf you would rather keep strict peer resolution, pinning React works too:
npm install react@~19.2.0 react-dom@~19.2.0The headless entry point (
react-ai-voice-avatar/headless) pulls in no Three.js at all, so it never hits this and works on any React 18 or 19 version.
To prevent the massive ML assets (WebGPU workers, 3D engines) from bloating your initial page load, use the built-in lazy wrapper. It will automatically code-split the 3D dependencies and render a sleek holographic Skeleton UI while the assets download in the background!
import { AiVoiceAvatarLazy } from 'react-ai-voice-avatar';
// Use it exactly like the normal component!
<AiVoiceAvatarLazy avatarPreset="ananya" />
Lazy loading defers the 3D engine. The speech and language models are the larger cost, and by default they start downloading as soon as the avatar mounts. On a landing page most visitors only look, so render the avatar idle and load the models when someone actually engages:
const [engaged, setEngaged] = useState(false);
<AiVoiceAvatar loadModels={engaged} hideStatusPill={!engaged} />
<button onClick={() => setEngaged(true)}>Talk to it</button>
The live demo's landing page works this way.
AiVoiceAvatar): Recommended for full-screen applications where the avatar is the primary product (e.g., Kiosks, Digital Tutors). The browser aggressively downloads the 3D canvas and ML models immediately so the avatar is ready instantly.AiVoiceAvatarLazy): Recommended for widgets, modals, or sub-routes (e.g., a "Support Desk" chat bubble in a SaaS dashboard). Defers downloading the 1.5MB 3D engine and WebWorkers until the user actually opens the widget.The react-ai-voice-avatar engine is truly zero-config. You do not need to configure Vite optimizeDeps, Next.js Webpack overrides, or manually host Web Worker files—everything is dynamically bundled and executed automatically!
However, because our ONNX WebGPU engine leverages modern multi-threaded SharedArrayBuffer memory pipelines for maximum inference speed, your hosting server can optionally emit standard Cross-Origin Isolation HTTP headers (COOP/COEP) to unlock peak performance. If these headers are not present, the engine automatically falls back to single-threaded WebAssembly without crashing.
To unlock multi-threaded performance, specify these isolation headers in your routing manifests:
vercel.json): Add "headers": [{ "source": "/(.*)", "headers": [{ "key": "Cross-Origin-Opener-Policy", "value": "same-origin" }, { "key": "Cross-Origin-Embedder-Policy", "value": "require-corp" }] }]._headers or netlify.toml): Add /*\n Cross-Origin-Opener-Policy: same-origin\n Cross-Origin-Embedder-Policy: require-corp to public/_headers.Vite (vite.config.ts):
export default defineConfig({
plugins: [react()],
server: {
headers: {
'Cross-Origin-Opener-Policy': 'same-origin',
'Cross-Origin-Embedder-Policy': 'require-corp',
},
},
});
Next.js (next.config.mjs):
export default {
async headers() {
return [
{
source: '/(.*)',
headers: [
{ key: 'Cross-Origin-Opener-Policy', value: 'same-origin' },
{ key: 'Cross-Origin-Embedder-Policy', value: 'require-corp' },
],
},
];
},
};
[!CAUTION] Strict CSP Policies: If your enterprise enforces strict Content Security Policies that block
blob:workers (worker-src 'self'), you can bypass our zero-config Blob loaders by passing theworkerBaseUrlprop to the avatar and hosting the pre-compiled.worker.jsfiles from ourdist/assets/directory yourself.
import React, { useRef, useState } from 'react';
import { Canvas } from '@react-three/fiber';
import { OrbitControls } from '@react-three/drei';
import { AiVoiceAvatar, type AiVoiceAvatarHandle } from 'react-ai-voice-avatar';
export function App() {
const avatarRef = useRef<AiVoiceAvatarHandle>(null);
const [text, setText] = useState('');
return (
<div style={{ width: '100vw', height: '100vh', position: 'relative' }}>
<Canvas camera={{ position: [0, 0.15, 2.2], fov: 32 }}>
<color attach="background" args={['#101116']} />
{/* Subtle studio lighting */}
<pointLight position={[-3, 2, -2]} intensity={25} color="#E67E22" distance={6} />
<pointLight position={[3, 1, -2]} intensity={20} color="#2980B9" distance={6} />
<OrbitControls target={[0, 0.05, 0]} />
{/* Connected Brain: Zero download, instant initialization! */}
<AiVoiceAvatar
ref={avatarRef}
avatarPreset="ananya"
lightingPreset="studio"
ttsEngine="kokoro"
ttsVoice="af_heart"
// Connect your backend here (receives user speech transcript):
onSubmit={async (text) => {
const res = await fetch('/api/chat', {
method: 'POST',
body: JSON.stringify({ prompt: text })
});
return res.body; // Avatar natively reads streams!
}}
/>
</Canvas>
{/* Fallback Text Input for noisy environments */}
<form
onSubmit={(e) => {
e.preventDefault();
if (text.trim() && avatarRef.current) {
avatarRef.current.sendText(text);
setText('');
}
}}
style={{ position: 'absolute', bottom: '20px', left: '50%', transform: 'translateX(-50%)', display: 'flex', gap: '8px', zIndex: 100 }}
>
<input
value={text}
onChange={e => setText(e.target.value)}
placeholder="Type a message..."
style={{ padding: '8px 16px', borderRadius: '20px', border: 'none', background: 'rgba(255,255,255,0.9)', width: '300px' }}
/>
<button type="submit" style={{ padding: '8px 16px', borderRadius: '20px', border: 'none', background: '#3b82f6', color: 'white', cursor: 'pointer' }}>
Send
</button>
</form>
</div>
);
}
The onSubmit prop natively accepts a string, an AsyncIterable<string>, or a ReadableStream. To connect your actual backend, simply drop in one of these copy-paste recipes to parse your streaming format!
[!CAUTION] API Keys Belong on the Server! Never put your OpenAI or Anthropic API keys directly in the frontend browser code. Always route through your own backend endpoint (
/api/chat).
If your backend uses Vercel AI SDK's streamText(...).toTextStreamResponse() or otherwise streams plain raw text, you can pass the stream natively without any parsing!
onSubmit={async (text) => {
const res = await fetch('/api/chat', { method: 'POST', body: JSON.stringify({ prompt: text }) });
return res.body; // Natively supported!
}}
Older versions of the Vercel AI SDK stream data using a specific protocol (e.g., 0:"Hello"). This recipe parses those chunks into clean text with a carry-over buffer for safe network boundaries.
onSubmit={async function* (text) {
const res = await fetch('/api/chat', { method: 'POST', body: JSON.stringify({ prompt: text }) });
if (!res.body) return;
const reader = res.body.getReader();
const decoder = new TextDecoder();
let buffer = '';
while (true) {
const { done, value } = await reader.read();
if (done) break;
buffer += decoder.decode(value, { stream: true });
const lines = buffer.split('\n');
buffer = lines.pop() ?? ''; // keep the trailing partial chunk
for (const line of lines) {
if (line.startsWith('0:')) {
try { yield JSON.parse(line.substring(2)); } catch { /* ignore keep-alive / non-JSON frames */ }
}
}
}
}}
Standard Server-Sent Events (SSE) stream data: {...} blocks. This handles safe parsing across broken network chunk boundaries.
onSubmit={async function* (text) {
const res = await fetch('/api/chat', { method: 'POST', body: JSON.stringify({ prompt: text }) });
if (!res.body) return;
const reader = res.body.getReader();
const decoder = new TextDecoder();
let buffer = '';
while (true) {
const { done, value } = await reader.read();
if (done) break;
buffer += decoder.decode(value, { stream: true });
const lines = buffer.split('\n');
buffer = lines.pop() ?? ''; // keep the trailing partial chunk
for (const line of lines) {
if (line.startsWith('data: ') && line !== 'data: [DONE]') {
try {
const parsed = JSON.parse(line.substring(6));
// AI SDK v5 emits {type:'text-delta', delta:'...'}; OpenAI emits choices[0].delta.content
if (parsed.type === 'text-delta' && parsed.delta) {
yield parsed.delta;
} else if (parsed.choices?.[0]?.delta?.content) {
yield parsed.choices[0].delta.content;
}
} catch { /* ignore keep-alive / non-JSON frames */ }
}
}
}
}}
English and Hindi both run entirely on the device. Set ttsLanguage and the
engine picks a matching voice, a matching speech-recognition model, and a
matching phoneme path.
<AiVoiceAvatar ttsLanguage="hi-IN" />
Hindi needed real work rather than a config flag. Kokoro ships four Hindi voices
inside the same checkpoint as the English ones, but the JavaScript wrapper does
not list them, and the bundled eSpeak build carries English data only and rejects
hi outright. So this package includes its own Devanagari-to-phoneme converter.
Devanagari is close to phonemic, which makes that tractable; the hard part is
schwa deletion, the rule that makes कमल read as "kamal" rather than "kamala", and
getting it wrong produces speech that still sounds like speech while sounding
like someone spelling Hindi out.
Because the converter emits real phonemes, Hindi gets phoneme-driven lip-sync with distinct mouth shapes for the retroflex consonants, not the amplitude-only mouth flapping most engines fall back to outside English. Code-switching works too: English words inside a Hindi sentence are routed to the English phonemiser, so "मुझे coffee चाहिए" is pronounced correctly throughout.
[!NOTE] Local Hindi is a demo, not a product. Models small enough to run in a browser are far weaker in Hindi than in English. It is good enough to show the pipeline working end to end and not good enough to ship. For production Hindi, route hearing and thinking to an API through
onTranscribeandonSubmit; the voice and lip-sync stay local and are genuinely good.
A session is one language at a time. Someone who switches language mid-conversation will be mistranscribed, because the recognition model is told which language to expect and browser Whisper cannot detect it.
The user can talk over the avatar and cut it off mid-sentence. This is on by default, because waiting for a reply to finish is the thing that makes a voice agent feel like a walkie-talkie.
<AiVoiceAvatar
allowInterruption={false} // default: true
onUserInterrupt={() => analytics.track('barge_in')}
/>
Turn it off for a kiosk or a noisy room, where the avatar hearing its own voice
through the speakers and stopping itself is worse than waiting. It is ignored in
push-to-talk, which owns the floor explicitly.
Voice detection scores every 96ms frame for how much it sounds like speech. A
cough, a door and a chair all clear a loudness bar as easily as a word does, so
loudness alone cannot separate them — sustain can. A sound has to keep
scoring as speech for minSpeechMs before the engine treats it as a turn.
Below that bar, a sound only ducks the avatar's voice, reversibly. If it turns out to be a cough the reply resumes from where it paused, at the right place in the sentence and with the mouth still in sync. Nothing about the conversation changed, because nothing was decided on a noise.
The defaults suit a quiet room and headphones. A shop floor is a different problem, and only you know which you have.
<AiVoiceAvatar
speechDetection={{ positiveSpeechThreshold: 0.6, minSpeechMs: 700 }}
/>
| Field | Default | Raise it when | Lower it when |
|---|---|---|---|
positiveSpeechThreshold | 0.5 | Passing noise is mistaken for talking | Quiet speakers go unheard |
negativeSpeechThreshold | 0.35 | Turns end too slowly in a noisy room | Turns end while someone is still talking |
minSpeechMs | 500 | Short noises still start turns | Single-word answers are ignored |
redemptionMs | 1400 | People are cut off while thinking mid-sentence | Replies feel slow to start |
preSpeechPadMs | 800 | The first word is still being clipped | Rarely — this is the audio kept from before the trigger, and it is what stops the first syllable going missing |
Two costs worth knowing before you change anything. A speaker who sits below
positiveSpeechThreshold produces no events at all — raising it trades
quiet voices for quiet rooms. And a filler like "hmm" held long enough to pass
minSpeechMs still reaches transcription; a denylist catches the common ones
after the fact, but sustain cannot tell a long "hmm" from a short word.
Pass onError and your app learns when something breaks, rather than finding out
from a console message it cannot see.
<AiVoiceAvatar
onError={(e) => {
if (e.severity === 'fatal') showFallbackUI(e.stage);
logToSentry(e);
}}
/>
Check severity before reacting. Most failures here are survivable because the
engine falls back: WebGPU to WASM, Kokoro to a smaller voice model. Those arrive
as degraded and the avatar still works, so treating them as fatal would hide a
working experience behind an error screen. A refused microphone is fatal for
listening while typed input still works, which is a judgement only your app can
make.
stage is one of microphone, speech-recognition, language-model,
speech-synthesis, audio-output, worker, model-storage or conversation.
There is also a detail string carrying the engine's internal stage name for bug
reports; it is not stable across versions, so do not branch on it.
model-storage is always degraded: a model loaded and works, but could not be
kept, so the next visit downloads it again. The message says why — usually not
enough free space, which on a phone is the common case.
Downloaded models are stored in the browser's Origin Private File System, so a returning visitor skips the download: on an Apple M4 with WebGPU, a return visit is ready to talk in about 2 seconds, with nothing fetched. The browser's Cache API — what the model loader uses by default — refuses any single file of 256 MiB or more, and both the voice (~310 MB) and the local language model (~750 MB) are larger than that, so before this they were downloaded again on every visit. Where OPFS is unavailable the Cache API is used as before.
Browsers can clear this storage when disk space runs low. If your app is one
people return to, call navigator.storage.persist() from a user gesture to ask
the browser to keep it; the engine does not, because Firefox can answer that
call with a permission prompt, and a library should not put one in front of
your visitors unasked.
You are not locked into our built-in avatars (ananya and aarav)! You can use any custom .glb humanoid model by passing its URL or local path to the modelSrc prop:
<AiVoiceAvatar
modelSrc="/models/my-custom-avatar.glb"
// ...
/>
To ensure the lip-sync and procedural facial dynamics engines work correctly, your custom model must meet the following standard requirements:
.glb (GLTF Binary).jawOpen, eyeBlinkLeft, mouthSmileRight). Our engine automatically traverses your model to find these targets.Head, head, Neck, or neck).ananya and aarav are converted from the Microsoft Rocketbox library, which is MIT licensed like this project, so you can redistribute them without restriction. The library has 115 avatars and any of them can be converted with scripts/convert-rocketbox.py. See assets/avatars/LICENSE.md.
Models exported from Ready Player Me also work, since they carry the same ARKit blendshapes and bone names. Note that Ready Player Me shut down in January 2026, so you can no longer create new avatars there, and existing exports are licensed CC BY-NC-SA rather than MIT.
Four stages, and you choose where each one happens. The defaults are all local, which is why the demo needs no keys, but the interesting production setups are mixed.
flowchart TD
mic(["🎙️ You speak"]) --> vad["Voice detector<br/><i>in the browser</i>"]
vad --> hear["Hearing<br/><i>Whisper in the browser</i><br/>or your onTranscribe"]
hear --> think["Thinking<br/><i>a model in the browser</i><br/>or your onSubmit"]
think --> speak["Speaking<br/><i>Kokoro in the browser</i><br/>or your onSynthesize"]
speak --> face["🔊 Voice, with a lip-synced<br/>3D face or your own UI"]
vad -. "talk over it and it stops" .-> speak
Each stage runs in the browser unless you pass the adapter beside it, which hands that stage to your own service.
| Stage | On the device | Your backend instead |
|---|---|---|
| Hearing — speech to text | Whisper via ONNX | onTranscribe |
| Thinking — the reply | Qwen or Gemma via WebGPU | onSubmit |
| Speaking — text to audio | Kokoro-82M | onSynthesize |
| Face — lip-sync and animation | Always here | Not applicable |
The face never leaves the device, which is the whole point: that is what avatar APIs charge per minute for, and it is the one stage that cannot be outsourced without a video stream.
The common production shape is onSubmit alone. Hearing and speaking stay local,
so no audio ever leaves the browser, while generation goes to whatever model you
already run. That keeps the download to about 600 MB (Whisper base and the Kokoro
voice), keeps your keys on your server, and still gives you a conversation nobody
else can read.
Add onTranscribe and onSynthesize as well and nothing downloads at all, apart
from the ~4 MB voice detector once the microphone opens. Each adapter replaces
its model entirely, so the page is ready as soon as it loads. That is the shape
for phones and for pages people visit once.
Fully local is real, not a demo trick, and it is the right answer for a kiosk, a regulated environment, or anywhere without reliable connectivity. Be aware of the cost: in English the first visit downloads roughly 1.3 GB before anyone can speak (the language model alone is 750 MB), about 2.1 GB in Hindi, and the quality ceiling is whatever a model that size can do.
Measured with scripts/measure-latency.mjs on an
Apple M4 with WebGPU in Chromium: a spoken question through a fake microphone,
the median of five turns after a warm-up turn, timed by the hook's own
callbacks.
| Stage | Time |
|---|---|
| Waiting for silence, to be sure you've finished | 1,400 ms (redemptionMs, adjustable) |
| Hearing: Whisper base transcribes the utterance | 330 ms |
| Thinking, hosted: GPT-OSS 20B on Groq, the demo's route, one round trip | 670 ms |
| Thinking and the first sentence of voice, all in the browser (Qwen 0.5B, Kokoro) | 740 ms |
| Speaking: Kokoro's first sentence ready and playing, with an instant reply | 460 ms |
| Starting up on a return visit, models already stored | about 2 s, nothing downloaded |
So from the moment you stop talking to the first word of the reply:
| Setup | First word after you stop |
|---|---|
| All in the browser | about 2.5 s |
| Hosted replies, local hearing and voice (the demo) | about 2.9 s |
| With an instant reply, the floor for speech alone | about 2.2 s |
The hosted figure adds the route's round trip to the measured speech stages,
since that route answers only the live site. Most of the wait is the pause for
silence, which is what stops the avatar cutting people off while they think.
Lower redemptionMs in speechDetection for a snappier feel, at the cost of
answering half-finished sentences. Without WebGPU, recognition and the voice
run on the CPU and are several times slower.
You do not have to choose once. preloadLocalLlm fetches the in-browser model
behind the conversation while your backend answers, and onLocalLlmReady tells
you when it can take over:
const [localReady, setLocalReady] = useState(false);
<AiVoiceAvatar
// Dropping onSubmit is what hands the conversation over.
onSubmit={localReady ? undefined : askMyBackend}
preloadLocalLlm
onLocalLlmReady={() => setLocalReady(true)}
/>
Nobody waits for a gigabyte to say the first word, and the turns your backend answered are handed to the local model, so it does not restart the conversation from nothing. Useful for a kiosk that must keep working when the wifi drops, a demo on someone else's quota, or a phone, which will never accept the download but can hold a conversation the moment it loads.
The live demo runs exactly this: replies come from a hosted model, and on a desktop with WebGPU the in-browser model downloads during the conversation and takes the floor when it lands. The page says which one is answering at any moment.
Explore the canonical patterns in the examples/ directory:
| Example Pattern | Folder | Highlights & Architecture |
|---|---|---|
| Live Interactive Demo | sandbox | Deploy on Vercel ➔ — Our full-featured interactive testbed featuring live character switching (ananya, aarav), voice persona switching (af_heart, am_michael), real-time diagnostic probe metrics, and Leva 3D lighting controls. |
| Quickstart | examples/quickstart | Minimal, zero-configuration plug-and-play AI voice avatar deployment with built-in studio lighting & sizing. |
| Local Kiosk | examples/local-kiosk | 100% offline on-device retail & restaurant ordering kiosk with embedded menu reasoning. Demonstrates the On-Device Brain; operates without internet access once model weights are locally cached. |
| Connected App | examples/hybrid-cloud | Illustrates the Connected Brain (onSubmit). Bypasses gigabyte-scale local LLM downloads by routing reasoning to OpenAI, Claude, or corporate APIs while keeping ASR, TTS, and 3D lip blending 100% on-device! |
| Voice Only | examples/voice-only | Voice mode with no avatar, like ChatGPT or Gemini voice: the react-ai-voice-avatar/headless hook, an audio-reactive orb, live captions and interruption. No three.js in the bundle, and no backend needed to try it. |
| Groq Voice | examples/groq-voice | Voice mode on Groq's free tier with your own key: hearing, replies and voice start on Groq with nothing to download, the in-browser hearing and voice download behind the conversation (preloadLocalSpeech), and the page offers to switch when they are ready. Shows Groq's live limits. |
| Headless Custom UI | examples/headless-custom-ui | Still renders the 3D avatar; for no avatar at all, see Voice Only above. Demonstrates hiding built-in DOM overlays (hideStatusPill={true}, showCaptions={false}), streaming transcripts into a custom enterprise UI, and controlling voice outputs imperatively via ref.current?.speak(text). |
<AiVoiceAvatar /> Props| Prop | Type | Default | Description |
|---|---|---|---|
avatarPreset | 'ananya' | 'aarav' | 'default' | 'kiosk' | 'ananya' | Built-in 3D character models featuring both female ('ananya') and male ('aarav') voice concierges out of the box with full ARKit facial blendshapes! |
avatarSize | 'sm' | 'md' | 'lg' | number | 'md' (0.48) | Intuitive model sizing presets or custom decimal scaling multiplier applied directly to the 3D humanoid mesh. |
modelSrc | string | undefined | Absolute local path or remote URL to a custom GLTF/GLB humanoid armature avatar model. |
lightingPreset | 'studio' | 'cyberpunk_violet' | 'cool_azure' | 'warm_amber' | 'clean_white' | 'none' | 'studio' | Pre-built cinematic studio lighting atmospheres directly applied to your 3D viewport without manual Three.js configuration! |
systemPrompt | string | "You are Ananya..." | Conversational persona directives and context injected into active LLMs. |
llmModel | string | by language | Hugging Face id for the local WebGPU reasoning model, used only when onSubmit is absent. Defaults to Qwen2.5-0.5B for English and Gemma 3 1B for Hindi, which Qwen that size cannot speak. Set it to pin one model for every language. |
asrModel | string | by language | Hugging Face id for the local Whisper model. Defaults to Whisper base for English and Whisper small for Hindi, which base transcribes badly. Pass "Xenova/whisper-tiny" for a faster download and worse accuracy. |
ttsEngine | 'kokoro' | 'mms' | 'kokoro' | High-fidelity neural voice synthesis engine executing inside dedicated Web Workers. |
ttsVoice | string | by language | Kokoro voice id. af_* and am_* American, bf_* and bm_* British, hf_* and hm_* Hindi (af_heart, am_michael, bf_emma, hf_alpha, hm_omega). Defaults to one matching ttsLanguage. A voice whose language disagrees is corrected with a warning. |
ttsLanguage | 'en-US' | 'en-GB' | 'hi-IN' | 'en-US' | Conversation language. Selects the voice, the recognition model and the phoneme path. Also sets asrLanguage unless you set that yourself. For any other language, pass onSynthesize and use a cloud voice provider. |
asrLanguage | string | ttsLanguage | Language to transcribe. Follows ttsLanguage by default, since a conversation is almost always held in one language. |
showCaptions | boolean | true | Renders a sleek glassmorphic subtitle overlay displaying spoken interaction dialog. |
hideStatusPill | boolean | false | When true, suppresses the default bottom-left microphone interactive control pill. |
listenMode | 'continuous' | 'push-to-talk' | 'continuous' | continuous keeps the mic hot after the avatar finishes speaking naturally, but explicitly clicking Stop forces it off until tapped again. push-to-talk strictly requires manually tapping to start listening for every single turn. |
loadModels | boolean | true | Set false to render the avatar without downloading any models, then flip it true when the visitor engages. For landing pages and widgets most visitors never talk to. Status stays 'loading' until it is true and the models are up, so show your own call to action meanwhile. |
onModelLoaded | () => void | undefined | Fires once the 3D mesh is parsed and in the scene. Parsing a multi-megabyte GLB leaves the canvas empty for a few seconds; use this to hold a placeholder over it. |
allowInterruption | boolean | true | Lets the user talk over the avatar and cut it off mid-sentence. Turn off for a kiosk or noisy room, where the avatar hearing itself through the speakers is worse than waiting. Ignored in push-to-talk. See Taking turns. |
speechDetection | { positiveSpeechThreshold?, negativeSpeechThreshold?, minSpeechMs?, redemptionMs?, preSpeechPadMs? } | see Taking turns | Tunes how the microphone decides someone is talking. Every field optional. The defaults suit a quiet room; a shop floor needs a higher threshold and a longer minSpeechMs. |
onUserInterrupt | () => void | undefined | Fires when the user talks over the avatar and takes the floor. Only fires if the avatar actually had audio playing. |
gestures | boolean | number | true | Hand and arm gestures while the avatar speaks: one hand or both brought up in front of the chest for each phrase, with small beats on stressed syllables, and the arms back at rest when it stops. false keeps the arms still; a number sets the size, 0 to 1.5. Needs a skeleton with LeftArm, LeftForeArm and LeftHand and the right-hand equivalents; custom avatars without them simply do not gesture. |
onAudioLevelChange | (level: number, source: 'mic' | 'tts' | 'idle') => void | undefined | Real-time audio amplitude (0-1) callbacks for the active stream. Essential for building highly responsive, audio-reactive 3D Visualizers and HUDs! Fires with 0 and 'idle' between turns, so a meter falls to rest rather than freezing. |
onSubmit | (text: string) => Promise<string | AsyncIterable<string> | ReadableStream> | undefined | Connected Brain API: Bypasses local LLMs; routes transcribed user microphone strings to your cloud or custom LLM API endpoint. |
preloadLocalLlm | boolean | false | Only meaningful alongside onSubmit, which otherwise skips the local language model download entirely. Set it to fetch that model in the background while your hosted one answers, so the conversation survives a rate limit, an expired quota or a lost network. See Hosted now, local when warm. |
onLocalLlmReady | () => void | undefined | Fires once the model requested by preloadLocalLlm has loaded. Drop onSubmit here to hand the conversation over; the turns your backend answered are carried across, so the local model knows what was already said. |
onTranscribe | (audio: Float32Array) => Promise<string> | undefined | Replaces local speech recognition with your own service, and Whisper is then not downloaded. Receives one utterance as 16 kHz mono samples. |
onSynthesize | (text: string) => Promise<Float32Array | ArrayBuffer> | undefined | Replaces local voice synthesis with your own service, and Kokoro is then not downloaded. Return an encoded MP3/WAV buffer, or raw 24 kHz PCM. Called a sentence at a time. |
onError | (e: AiVoiceAvatarError) => void | undefined | Fires when a stage fails. Carries stage, message and a severity of degraded or fatal. See Handling failures. |
onTranscriptUpdate | (text: string, speaker: 'user' | 'avatar') => void | undefined | Callback delivering real-time microphone transcriptions and assistant spoken utterance strings. |
onStatusChange | (status: string) => void | undefined | Emits live state transitions (loading, idle, listening, thinking, speaking). |
debug | boolean | false | When true, renders an interactive floating GUI (Leva) to inspect and tune individual 3D blendshapes. |
vadAssetPath | string | undefined | Optional URL or local path override for self-hosting @ricky0123/vad-web ONNX asset binaries in airgapped deployments. |
onnxWasmPath | string | undefined | Optional URL override for self-hosting onnxruntime-web WASM distribution files. |
workerBaseUrl | string | undefined | CSP Escape Hatch: if blob: workers are blocked by your server, fetch pre-compiled Web Workers from this URL directory. |
enableLocalAssetProbe | boolean | false | When true, HEAD-checks /ananya.glb in your own public directory before falling back to the CDN. Off by default: with no local copy the probe 404s, and that 404 lands in every visitor's console. |
statusPillStyle | React.CSSProperties | undefined | Optional custom CSS styling & absolute positioning overrides for the interactive Status Pill overlay. |
accentColor | string | undefined | Custom CSS color string (e.g., #38BDF8) for the active status indicator rings and highlights. |
AiVoiceAvatarHandle)Attach a React ref (useRef<AiVoiceAvatarHandle>(null)) to access imperative real-time controls:
interface AiVoiceAvatarHandle {
/** Command the 3D avatar to speak an arbitrary string with synchronized acoustic lip blending */
speak: (text: string) => void;
/** Manually engage microphone recording and Voice Activity Detection (VAD) */
startListening: () => void;
/** Pause active microphone listening */
stopListening: () => void;
/** Instantly interrupt and halt active voice speech synthesis and clear the audio queue */
interrupt: () => void;
/** Manually submit text to the onSubmit handler, simulating a spoken utterance (useful for text-only fallback) */
sendText: (text: string) => void;
/** Wipe multi-turn conversation memory history and caption overlay states */
clearHistory: () => void;
/** Retrieve live Web Audio API AnalyserNode powering real-time spectral lip sync */
getAnalyser: () => AnalyserNode | undefined;
}
sendText)If your users cannot use a microphone (e.g., noisy environments, privacy concerns, or lack of permissions), you can easily wire up a standard text input field to bypass the speech recognition pipeline entirely!
Simply attach a ref and call sendText() to pass a string directly to your onSubmit handler (or local LLM):
const avatarRef = useRef<AiVoiceAvatarHandle>(null);
// In your UI, attach this to a standard <form> submission:
const handleTextSubmit = (userInput: string) => {
avatarRef.current?.sendText(userInput);
}
When you use sendText, the avatar immediately enters the thinking state and processes the interaction exactly as if the user had spoken it aloud.
useAiVoiceAvatar() (headless)import { useAiVoiceAvatar } from 'react-ai-voice-avatar/headless';
Takes the same options as the component's props, except the ones about the 3D
scene and its overlays (avatarPreset, avatarSize, modelSrc,
lightingPreset, gestures, showCaptions, hideStatusPill, onModelLoaded,
debug and the styling props). onStatusChange is replaced by the returned
status. A few options matter mostly without an avatar:
| Option | Type | Description |
|---|---|---|
onAudioLevelChange | (level: number, source: 'mic' | 'tts' | 'idle') => void | Loudness from 0 to 1, every frame, from whichever side is talking. What an orb or waveform animates from. Write it to the DOM through a ref rather than into state. |
onSpeechStart | (text: string) => void | Each sentence of the reply as it starts playing. Captions that keep pace with the voice. |
onTranscriptUpdate | (text: string, speaker: 'user' | 'avatar') => void | What the user said, once transcribed, and the reply in full. |
onInferenceStart / onInferenceEnd | () => void | Around each turn, from the moment the user stops talking to the end of the reply. |
onTtsEngineChange | (engine: 'kokoro' | 'mms' | 'custom') => void | Fires whenever activeTtsEngine changes. Also a prop on the component. |
preloadLocalSpeech | boolean | Download the in-browser hearing and voice behind onTranscribe and onSynthesize, without waiting for them. Drop the adapters once onLocalSpeechReady fires and the local models take over. Also a prop on the component. |
onLocalSpeechReady | () => void | Fires once, when the models requested by preloadLocalSpeech are loaded. |
loadingProgress | (pct: number, label: string) => void | Download progress per model: 'asr', 'kokoro' (or 'tts' for MMS) and 'llm'. They download in parallel, so keep one figure per label. |
It returns:
| Value | Type | Description |
|---|---|---|
status | 'loading' | 'idle' | 'listening' | 'thinking' | 'speaking' | Where the conversation is. 'loading' until the models are up, and until loadModels is true. |
isLoading, isIdle, isListening, isThinking, isSpeaking | boolean | Shorthands for status. |
isReady | boolean | The models are loaded. startListening does nothing before this. |
isLocalSpeechReady | boolean | The in-browser hearing and voice are loaded, whether or not they are in use. |
activeTtsEngine | 'kokoro' | 'mms' | 'custom' | The voice actually speaking. iPhones and iPads get 'mms', a single plainer voice that ignores ttsVoice, because Kokoro runs Safari out of memory; so does any device where Kokoro fails to load. 'custom' is your onSynthesize. Worth showing if your users will compare devices. |
startListening | () => Promise<void> | Opens the microphone and starts listening. Call it from a click: the first call is where the browser asks for permission. In 'continuous' mode the microphone then stays open across turns. |
stopListening | () => void | Closes the microphone. |
interrupt | () => void | Stops the reply mid-sentence and clears what was queued. |
speak | (text: string) => void | Says the text without a model turn: a greeting, a notification. Plays a sentence at a time. |
sendText | (text: string) => void | Typed input, handled as if it had been spoken. |
clearHistory | () => void | Forgets the conversation so far. |
micError | string | null | Why the microphone could not open, such as a denied permission. null otherwise. |
analyser | AnalyserNode | undefined | The reply's audio, for a frequency visualiser. onAudioLevelChange is simpler if you only need loudness. |
It also returns several refs (currentSpeechTextRef, audioContextRef and
others) that the 3D component uses for lip-sync. A voice UI can ignore them.
We actively welcome community contributions. CONTRIBUTING.md has the local development guide, and ROADMAP.md has what is worth doing next, why, and what is already known about each item — including the measurements behind the open questions.
The largest pieces currently open:
armRig.ts.scripts/verify-pack.mjs).Hindi speech and VAD sensitivity tuning (speechDetection) were previously
listed here and have both shipped.
This library heavily relies on modern Web APIs (WebGPU, WebGL, Web Audio, and Web Workers). It gracefully degrades when certain APIs are unavailable.
| Browser | OS | 3D Rendering (WebGL) | Voice Synthesis (WebGPU/WASM) | Voice Recognition (Web Audio) | Status |
|---|---|---|---|---|---|
| Chrome / Edge | Windows, macOS, Android | ✅ Native | ✅ WebGPU (Ultra Fast) | ✅ Native | 🟢 Tier 1 (Recommended) |
| Safari / iOS | macOS, iOS | ✅ Native | 🔄 Lightweight Models by Default | ✅ Native | 🟡 Supported |
| Firefox | Windows, macOS | ✅ Native | ⚠️ WASM Fallback | ✅ Native | 🟡 Tier 2 (Slower TTS) |
[!NOTE]
- WebGPU is currently enabled by default in Chrome/Edge. On browsers without WebGPU, the library automatically falls back to WASM execution.
- iOS/Safari Preemptive Fallback: Safari and iOS impose strict memory limits that cause 80MB+ models (like Kokoro) to crash the tab. The engine automatically preempts this by forcing the lightweight MMS TTS model (~30MB) on iOS devices, and tracks crash breadcrumbs to prevent OOM reload loops.
- Strict CSP Environments: Safari and Firefox may block
blob:worker execution depending on your Content-Security-Policy headers. If this occurs, host the.worker.jsfiles statically and pass their base path via theworkerBaseUrlprop.
Running Neural Networks in the browser requires capable hardware.
| Deployment Mode | Min RAM | GPU Requirement | Recommended Devices |
|---|---|---|---|
| Connected Brain (ASR + TTS only) | 4GB | None (WASM Fallback ok) | iPhone 11+, Mid-range Android (2021+), Any Laptop |
| Full Local AI (ASR + 500M LLM + TTS) | 8GB | WebGPU Support Preferred | iPhone 13 Pro+, High-end Android (Snapdragon 8 Gen 1+), M1/M2 Macs, Modern PCs |
[!TIP] Mobile Memory Limits: Mobile browsers rigidly enforce memory limits per tab (often terminating tabs exceeding ~1GB). If your mobile app crashes "after some time", ensure you are utilizing the
Connected Brainmode (onSubmitAPI) which offloads the heavy LLM memory footprint to your server while keeping ultra-fast lip-sync and TTS local.
MIT © React AI Voice Avatar Contributors.
TypeScript
74.9%
JavaScript
19.9%
Python
3.5%
Open-source alternative to real-time avatar APIs like HeyGen Interactive Avatar and Tavus. A talking, lip-synced 3D avatar that runs on your user's GPU, driven by your own LLM. No video stream, no per-minute billing. Or voice mode alone, ChatGPT-style, as a headless React hook. 📦 npm: react-ai-voice-avatar
See the codereact-ai-voice-avatar) 🚀🗣️🧬The open-source alternative to real-time avatar APIs, as a React component.
HeyGen Interactive Avatar, Tavus and Soul Machines sell a talking avatar that holds a conversation, rendered on their servers and billed per streaming minute. This does the same job as an npm install: the avatar renders on your user's GPU, so there is no video stream, no per-minute cost, and no third party in the middle of your conversations.
You bring the model. Point it at OpenAI, Anthropic, your own fine-tune or your existing chat endpoint, and keep your keys on your own backend. The package owns the parts that are tedious to build and easy to get wrong: microphone capture, knowing when someone has finished speaking, interrupting the avatar mid-sentence when they talk over it, streaming speech synthesis, and driving 52 ARKit facial blendshapes at 60 FPS so the mouth matches the words.
It can also run with no backend at all. Speech recognition, generation and voice all have in-browser implementations, which makes for a convincing demo and a genuinely offline kiosk. Most production apps will use their own model and keep only speech and lip-sync on the device.
Don't need a face? The same engine is a headless React hook for voice mode, the way ChatGPT and Gemini do it: react-ai-voice-avatar/headless, with no three.js in your bundle. See it ➔ · How ➔
Speech recognition, voice synthesis and lip-sync, all running in your browser tab.
The first visit downloads about 590 MB of models, roughly a minute and a half on a
50 Mbit/s connection. Your browser keeps them, so a return visit is talking in
about 2 seconds with nothing downloaded. Replies on the demo come from a
hosted model on Groq's free tier, shared by everyone trying it, and start about
3 seconds after you stop talking; on a busy day that tier can run out, and
examples/groq-voice
runs the same thing on your own free key.
![]()
react-ai-voice-avatar provides two distinct ways to integrate into your app depending on your design needs. Both share the exact same underlying conversational state machine, Voice Activity Detection (VAD), and turn-taking logic.
Voice mode for your app: talk to it the way you talk to ChatGPT or Gemini, and interrupt it mid-sentence. The hook owns the microphone, knowing when someone has finished speaking, interruption, and streaming the reply into speech. You own the UI, and every frame it hands you a loudness level to animate.
![]()
Try it ➔: a page built on this hook alone. examples/voice-only is the same page as an app to copy.
Import it from react-ai-voice-avatar/headless and Three.js never enters your module graph. Measured on the same Next.js App Router build, one route rendering the 3D avatar and one rendering only the hook:
| Route | First Load JS |
|---|---|
| 3D avatar | 389 kB |
| Headless hook | 117 kB |
npm install react-ai-voice-avatar
import { useRef, useState } from 'react';
// The /headless subpath is what keeps Three.js out of your bundle.
// Importing the hook from the package root pulls the 3D stack in with it.
import { useAiVoiceAvatar } from 'react-ai-voice-avatar/headless';
export function VoiceMode() {
const orb = useRef<HTMLDivElement>(null);
const [micOn, setMicOn] = useState(false);
const voice = useAiVoiceAvatar({
// Your model. Return a string, or a stream so speech starts sooner.
onSubmit: text => fetch('/api/chat', { method: 'POST', body: text }).then(r => r.body),
// 0 to 1 every frame, from whoever is talking. Written straight to the
// DOM: routing it through state would re-render sixty times a second.
onAudioLevelChange: level => {
if (orb.current) orb.current.style.transform = `scale(${1 + level * 0.3})`;
},
});
// The mic stays open across turns, so it is tracked apart from `status`,
// which moves through listening, thinking and speaking.
const toggle = () => {
if (micOn) {
voice.stopListening();
voice.interrupt();
} else {
voice.startListening(); // From a click: this is where the mic prompt appears.
}
setMicOn(!micOn);
};
return (
<>
<div ref={orb} className="orb" data-status={voice.status} />
<button onClick={toggle} disabled={!voice.isReady}>
{voice.isReady ? (micOn ? 'Stop' : 'Talk') : 'Loading…'}
</button>
</>
);
}
For captions, onTranscriptUpdate(text, 'user') gives what the user said and onSpeechStart(text) each sentence of the reply as it starts playing. speak(text) says something without a model turn, such as a greeting, and sendText(text) takes typed input. Everything it returns is in the hook reference.
Each stage runs in the browser or on your backend, and the choice is mostly about the download:
| Setup | Options | Downloaded by each visitor | Audio leaves the device |
|---|---|---|---|
| Cloud speech | onTranscribe + onSubmit + onSynthesize | Nothing until the mic opens, then ~4 MB for the voice detector | Yes, to your providers |
| Local speech (the demo) | onSubmit | ~590 MB once, kept for later visits (~320 MB on iPhone) | No |
| Fully local | none | ~1.3 GB once | Nothing leaves at all |
Cloud speech is the setup to reach for on phones and on pages people visit once: it is ready as soon as the page is. Local speech keeps what people say on their device and costs nothing per minute, and suits an app they come back to, since the models are stored after the first visit. Fully local is for kiosks and offline use. Use loadModels to hold any download until someone actually engages.
onTranscribe and onSynthesize replace the local speech models with your own providers. Each replaces its model entirely: supply one and that model is never downloaded. Drop it later and the local model loads, so a provider outage can fall back to the browser mid-session.
To have that fallback ready rather than downloading it at the moment you need it, pass preloadLocalSpeech: the local hearing and voice download behind the conversation, and onLocalSpeechReady fires when they could take over. examples/groq-voice starts on Groq and offers the switch then.
const voice = useAiVoiceAvatar({
// Instead of Whisper: 16 kHz mono samples of one utterance, to any
// speech-to-text service. Encode them as WAV if the service wants a file.
onTranscribe: async (samples: Float32Array) => {
return await mySpeechToText(samples);
},
// Instead of Kokoro: return an encoded MP3/WAV (ArrayBuffer), which the
// hook decodes, or raw 24 kHz PCM as a Float32Array. Called a sentence at
// a time, so the first plays while the rest are fetched.
onSynthesize: async (text: string) => {
const res = await fetch('/api/tts', { method: 'POST', body: text });
return await res.arrayBuffer();
},
onSubmit: async (text) => fetch('/api/chat', { method: 'POST', body: text }).then(r => r.body),
});
If you want the full visual presence with 60FPS ARKit lip-syncing, use the drop-in 3D component. Under the hood, this is just a wrapper around the useAiVoiceAvatar hook that procedurally maps the audio to a 3D model!
📦 View Package on the Official NPM Registry ➔
# If using the 3D Avatar, you must also install the Three.js ecosystem
npm install react-ai-voice-avatar three @react-three/fiber @react-three/drei
[!NOTE] React 18 Users: Installing the latest
@react-three/dreidefaults to version 10, which demands React 19. If your project runs on React 18, install compatible Three.js React bindings explicitly:npm install @react-three/drei@^9 @react-three/fiber@^8
[!IMPORTANT] React 19.3 and
ERESOLVE:@react-three/fiber@9.7still declares its React peer as>=19 <19.3, so a defaultnpm installalongside React 19.3 or newer fails withERESOLVE unable to resolve dependency tree. This is a Three.js binding constraint, not a limit of this package: our own peer range accepts React 19.3.That range is over-cautious. We build and run a Next.js App Router app against React 19.3.0 with
@react-three/fiber@9.7.0and the avatar renders, loads its models and speaks with no errors. So install past it rather than downgrading:npm install react-ai-voice-avatar three @react-three/fiber @react-three/drei --legacy-peer-depsIf you would rather keep strict peer resolution, pinning React works too:
npm install react@~19.2.0 react-dom@~19.2.0The headless entry point (
react-ai-voice-avatar/headless) pulls in no Three.js at all, so it never hits this and works on any React 18 or 19 version.
To prevent the massive ML assets (WebGPU workers, 3D engines) from bloating your initial page load, use the built-in lazy wrapper. It will automatically code-split the 3D dependencies and render a sleek holographic Skeleton UI while the assets download in the background!
import { AiVoiceAvatarLazy } from 'react-ai-voice-avatar';
// Use it exactly like the normal component!
<AiVoiceAvatarLazy avatarPreset="ananya" />
Lazy loading defers the 3D engine. The speech and language models are the larger cost, and by default they start downloading as soon as the avatar mounts. On a landing page most visitors only look, so render the avatar idle and load the models when someone actually engages:
const [engaged, setEngaged] = useState(false);
<AiVoiceAvatar loadModels={engaged} hideStatusPill={!engaged} />
<button onClick={() => setEngaged(true)}>Talk to it</button>
The live demo's landing page works this way.
AiVoiceAvatar): Recommended for full-screen applications where the avatar is the primary product (e.g., Kiosks, Digital Tutors). The browser aggressively downloads the 3D canvas and ML models immediately so the avatar is ready instantly.AiVoiceAvatarLazy): Recommended for widgets, modals, or sub-routes (e.g., a "Support Desk" chat bubble in a SaaS dashboard). Defers downloading the 1.5MB 3D engine and WebWorkers until the user actually opens the widget.The react-ai-voice-avatar engine is truly zero-config. You do not need to configure Vite optimizeDeps, Next.js Webpack overrides, or manually host Web Worker files—everything is dynamically bundled and executed automatically!
However, because our ONNX WebGPU engine leverages modern multi-threaded SharedArrayBuffer memory pipelines for maximum inference speed, your hosting server can optionally emit standard Cross-Origin Isolation HTTP headers (COOP/COEP) to unlock peak performance. If these headers are not present, the engine automatically falls back to single-threaded WebAssembly without crashing.
To unlock multi-threaded performance, specify these isolation headers in your routing manifests:
vercel.json): Add "headers": [{ "source": "/(.*)", "headers": [{ "key": "Cross-Origin-Opener-Policy", "value": "same-origin" }, { "key": "Cross-Origin-Embedder-Policy", "value": "require-corp" }] }]._headers or netlify.toml): Add /*\n Cross-Origin-Opener-Policy: same-origin\n Cross-Origin-Embedder-Policy: require-corp to public/_headers.Vite (vite.config.ts):
export default defineConfig({
plugins: [react()],
server: {
headers: {
'Cross-Origin-Opener-Policy': 'same-origin',
'Cross-Origin-Embedder-Policy': 'require-corp',
},
},
});
Next.js (next.config.mjs):
export default {
async headers() {
return [
{
source: '/(.*)',
headers: [
{ key: 'Cross-Origin-Opener-Policy', value: 'same-origin' },
{ key: 'Cross-Origin-Embedder-Policy', value: 'require-corp' },
],
},
];
},
};
[!CAUTION] Strict CSP Policies: If your enterprise enforces strict Content Security Policies that block
blob:workers (worker-src 'self'), you can bypass our zero-config Blob loaders by passing theworkerBaseUrlprop to the avatar and hosting the pre-compiled.worker.jsfiles from ourdist/assets/directory yourself.
import React, { useRef, useState } from 'react';
import { Canvas } from '@react-three/fiber';
import { OrbitControls } from '@react-three/drei';
import { AiVoiceAvatar, type AiVoiceAvatarHandle } from 'react-ai-voice-avatar';
export function App() {
const avatarRef = useRef<AiVoiceAvatarHandle>(null);
const [text, setText] = useState('');
return (
<div style={{ width: '100vw', height: '100vh', position: 'relative' }}>
<Canvas camera={{ position: [0, 0.15, 2.2], fov: 32 }}>
<color attach="background" args={['#101116']} />
{/* Subtle studio lighting */}
<pointLight position={[-3, 2, -2]} intensity={25} color="#E67E22" distance={6} />
<pointLight position={[3, 1, -2]} intensity={20} color="#2980B9" distance={6} />
<OrbitControls target={[0, 0.05, 0]} />
{/* Connected Brain: Zero download, instant initialization! */}
<AiVoiceAvatar
ref={avatarRef}
avatarPreset="ananya"
lightingPreset="studio"
ttsEngine="kokoro"
ttsVoice="af_heart"
// Connect your backend here (receives user speech transcript):
onSubmit={async (text) => {
const res = await fetch('/api/chat', {
method: 'POST',
body: JSON.stringify({ prompt: text })
});
return res.body; // Avatar natively reads streams!
}}
/>
</Canvas>
{/* Fallback Text Input for noisy environments */}
<form
onSubmit={(e) => {
e.preventDefault();
if (text.trim() && avatarRef.current) {
avatarRef.current.sendText(text);
setText('');
}
}}
style={{ position: 'absolute', bottom: '20px', left: '50%', transform: 'translateX(-50%)', display: 'flex', gap: '8px', zIndex: 100 }}
>
<input
value={text}
onChange={e => setText(e.target.value)}
placeholder="Type a message..."
style={{ padding: '8px 16px', borderRadius: '20px', border: 'none', background: 'rgba(255,255,255,0.9)', width: '300px' }}
/>
<button type="submit" style={{ padding: '8px 16px', borderRadius: '20px', border: 'none', background: '#3b82f6', color: 'white', cursor: 'pointer' }}>
Send
</button>
</form>
</div>
);
}
The onSubmit prop natively accepts a string, an AsyncIterable<string>, or a ReadableStream. To connect your actual backend, simply drop in one of these copy-paste recipes to parse your streaming format!
[!CAUTION] API Keys Belong on the Server! Never put your OpenAI or Anthropic API keys directly in the frontend browser code. Always route through your own backend endpoint (
/api/chat).
If your backend uses Vercel AI SDK's streamText(...).toTextStreamResponse() or otherwise streams plain raw text, you can pass the stream natively without any parsing!
onSubmit={async (text) => {
const res = await fetch('/api/chat', { method: 'POST', body: JSON.stringify({ prompt: text }) });
return res.body; // Natively supported!
}}
Older versions of the Vercel AI SDK stream data using a specific protocol (e.g., 0:"Hello"). This recipe parses those chunks into clean text with a carry-over buffer for safe network boundaries.
onSubmit={async function* (text) {
const res = await fetch('/api/chat', { method: 'POST', body: JSON.stringify({ prompt: text }) });
if (!res.body) return;
const reader = res.body.getReader();
const decoder = new TextDecoder();
let buffer = '';
while (true) {
const { done, value } = await reader.read();
if (done) break;
buffer += decoder.decode(value, { stream: true });
const lines = buffer.split('\n');
buffer = lines.pop() ?? ''; // keep the trailing partial chunk
for (const line of lines) {
if (line.startsWith('0:')) {
try { yield JSON.parse(line.substring(2)); } catch { /* ignore keep-alive / non-JSON frames */ }
}
}
}
}}
Standard Server-Sent Events (SSE) stream data: {...} blocks. This handles safe parsing across broken network chunk boundaries.
onSubmit={async function* (text) {
const res = await fetch('/api/chat', { method: 'POST', body: JSON.stringify({ prompt: text }) });
if (!res.body) return;
const reader = res.body.getReader();
const decoder = new TextDecoder();
let buffer = '';
while (true) {
const { done, value } = await reader.read();
if (done) break;
buffer += decoder.decode(value, { stream: true });
const lines = buffer.split('\n');
buffer = lines.pop() ?? ''; // keep the trailing partial chunk
for (const line of lines) {
if (line.startsWith('data: ') && line !== 'data: [DONE]') {
try {
const parsed = JSON.parse(line.substring(6));
// AI SDK v5 emits {type:'text-delta', delta:'...'}; OpenAI emits choices[0].delta.content
if (parsed.type === 'text-delta' && parsed.delta) {
yield parsed.delta;
} else if (parsed.choices?.[0]?.delta?.content) {
yield parsed.choices[0].delta.content;
}
} catch { /* ignore keep-alive / non-JSON frames */ }
}
}
}
}}
English and Hindi both run entirely on the device. Set ttsLanguage and the
engine picks a matching voice, a matching speech-recognition model, and a
matching phoneme path.
<AiVoiceAvatar ttsLanguage="hi-IN" />
Hindi needed real work rather than a config flag. Kokoro ships four Hindi voices
inside the same checkpoint as the English ones, but the JavaScript wrapper does
not list them, and the bundled eSpeak build carries English data only and rejects
hi outright. So this package includes its own Devanagari-to-phoneme converter.
Devanagari is close to phonemic, which makes that tractable; the hard part is
schwa deletion, the rule that makes कमल read as "kamal" rather than "kamala", and
getting it wrong produces speech that still sounds like speech while sounding
like someone spelling Hindi out.
Because the converter emits real phonemes, Hindi gets phoneme-driven lip-sync with distinct mouth shapes for the retroflex consonants, not the amplitude-only mouth flapping most engines fall back to outside English. Code-switching works too: English words inside a Hindi sentence are routed to the English phonemiser, so "मुझे coffee चाहिए" is pronounced correctly throughout.
[!NOTE] Local Hindi is a demo, not a product. Models small enough to run in a browser are far weaker in Hindi than in English. It is good enough to show the pipeline working end to end and not good enough to ship. For production Hindi, route hearing and thinking to an API through
onTranscribeandonSubmit; the voice and lip-sync stay local and are genuinely good.
A session is one language at a time. Someone who switches language mid-conversation will be mistranscribed, because the recognition model is told which language to expect and browser Whisper cannot detect it.
The user can talk over the avatar and cut it off mid-sentence. This is on by default, because waiting for a reply to finish is the thing that makes a voice agent feel like a walkie-talkie.
<AiVoiceAvatar
allowInterruption={false} // default: true
onUserInterrupt={() => analytics.track('barge_in')}
/>
Turn it off for a kiosk or a noisy room, where the avatar hearing its own voice
through the speakers and stopping itself is worse than waiting. It is ignored in
push-to-talk, which owns the floor explicitly.
Voice detection scores every 96ms frame for how much it sounds like speech. A
cough, a door and a chair all clear a loudness bar as easily as a word does, so
loudness alone cannot separate them — sustain can. A sound has to keep
scoring as speech for minSpeechMs before the engine treats it as a turn.
Below that bar, a sound only ducks the avatar's voice, reversibly. If it turns out to be a cough the reply resumes from where it paused, at the right place in the sentence and with the mouth still in sync. Nothing about the conversation changed, because nothing was decided on a noise.
The defaults suit a quiet room and headphones. A shop floor is a different problem, and only you know which you have.
<AiVoiceAvatar
speechDetection={{ positiveSpeechThreshold: 0.6, minSpeechMs: 700 }}
/>
| Field | Default | Raise it when | Lower it when |
|---|---|---|---|
positiveSpeechThreshold | 0.5 | Passing noise is mistaken for talking | Quiet speakers go unheard |
negativeSpeechThreshold | 0.35 | Turns end too slowly in a noisy room | Turns end while someone is still talking |
minSpeechMs | 500 | Short noises still start turns | Single-word answers are ignored |
redemptionMs | 1400 | People are cut off while thinking mid-sentence | Replies feel slow to start |
preSpeechPadMs | 800 | The first word is still being clipped | Rarely — this is the audio kept from before the trigger, and it is what stops the first syllable going missing |
Two costs worth knowing before you change anything. A speaker who sits below
positiveSpeechThreshold produces no events at all — raising it trades
quiet voices for quiet rooms. And a filler like "hmm" held long enough to pass
minSpeechMs still reaches transcription; a denylist catches the common ones
after the fact, but sustain cannot tell a long "hmm" from a short word.
Pass onError and your app learns when something breaks, rather than finding out
from a console message it cannot see.
<AiVoiceAvatar
onError={(e) => {
if (e.severity === 'fatal') showFallbackUI(e.stage);
logToSentry(e);
}}
/>
Check severity before reacting. Most failures here are survivable because the
engine falls back: WebGPU to WASM, Kokoro to a smaller voice model. Those arrive
as degraded and the avatar still works, so treating them as fatal would hide a
working experience behind an error screen. A refused microphone is fatal for
listening while typed input still works, which is a judgement only your app can
make.
stage is one of microphone, speech-recognition, language-model,
speech-synthesis, audio-output, worker, model-storage or conversation.
There is also a detail string carrying the engine's internal stage name for bug
reports; it is not stable across versions, so do not branch on it.
model-storage is always degraded: a model loaded and works, but could not be
kept, so the next visit downloads it again. The message says why — usually not
enough free space, which on a phone is the common case.
Downloaded models are stored in the browser's Origin Private File System, so a returning visitor skips the download: on an Apple M4 with WebGPU, a return visit is ready to talk in about 2 seconds, with nothing fetched. The browser's Cache API — what the model loader uses by default — refuses any single file of 256 MiB or more, and both the voice (~310 MB) and the local language model (~750 MB) are larger than that, so before this they were downloaded again on every visit. Where OPFS is unavailable the Cache API is used as before.
Browsers can clear this storage when disk space runs low. If your app is one
people return to, call navigator.storage.persist() from a user gesture to ask
the browser to keep it; the engine does not, because Firefox can answer that
call with a permission prompt, and a library should not put one in front of
your visitors unasked.
You are not locked into our built-in avatars (ananya and aarav)! You can use any custom .glb humanoid model by passing its URL or local path to the modelSrc prop:
<AiVoiceAvatar
modelSrc="/models/my-custom-avatar.glb"
// ...
/>
To ensure the lip-sync and procedural facial dynamics engines work correctly, your custom model must meet the following standard requirements:
.glb (GLTF Binary).jawOpen, eyeBlinkLeft, mouthSmileRight). Our engine automatically traverses your model to find these targets.Head, head, Neck, or neck).ananya and aarav are converted from the Microsoft Rocketbox library, which is MIT licensed like this project, so you can redistribute them without restriction. The library has 115 avatars and any of them can be converted with scripts/convert-rocketbox.py. See assets/avatars/LICENSE.md.
Models exported from Ready Player Me also work, since they carry the same ARKit blendshapes and bone names. Note that Ready Player Me shut down in January 2026, so you can no longer create new avatars there, and existing exports are licensed CC BY-NC-SA rather than MIT.
Four stages, and you choose where each one happens. The defaults are all local, which is why the demo needs no keys, but the interesting production setups are mixed.
flowchart TD
mic(["🎙️ You speak"]) --> vad["Voice detector<br/><i>in the browser</i>"]
vad --> hear["Hearing<br/><i>Whisper in the browser</i><br/>or your onTranscribe"]
hear --> think["Thinking<br/><i>a model in the browser</i><br/>or your onSubmit"]
think --> speak["Speaking<br/><i>Kokoro in the browser</i><br/>or your onSynthesize"]
speak --> face["🔊 Voice, with a lip-synced<br/>3D face or your own UI"]
vad -. "talk over it and it stops" .-> speak
Each stage runs in the browser unless you pass the adapter beside it, which hands that stage to your own service.
| Stage | On the device | Your backend instead |
|---|---|---|
| Hearing — speech to text | Whisper via ONNX | onTranscribe |
| Thinking — the reply | Qwen or Gemma via WebGPU | onSubmit |
| Speaking — text to audio | Kokoro-82M | onSynthesize |
| Face — lip-sync and animation | Always here | Not applicable |
The face never leaves the device, which is the whole point: that is what avatar APIs charge per minute for, and it is the one stage that cannot be outsourced without a video stream.
The common production shape is onSubmit alone. Hearing and speaking stay local,
so no audio ever leaves the browser, while generation goes to whatever model you
already run. That keeps the download to about 600 MB (Whisper base and the Kokoro
voice), keeps your keys on your server, and still gives you a conversation nobody
else can read.
Add onTranscribe and onSynthesize as well and nothing downloads at all, apart
from the ~4 MB voice detector once the microphone opens. Each adapter replaces
its model entirely, so the page is ready as soon as it loads. That is the shape
for phones and for pages people visit once.
Fully local is real, not a demo trick, and it is the right answer for a kiosk, a regulated environment, or anywhere without reliable connectivity. Be aware of the cost: in English the first visit downloads roughly 1.3 GB before anyone can speak (the language model alone is 750 MB), about 2.1 GB in Hindi, and the quality ceiling is whatever a model that size can do.
Measured with scripts/measure-latency.mjs on an
Apple M4 with WebGPU in Chromium: a spoken question through a fake microphone,
the median of five turns after a warm-up turn, timed by the hook's own
callbacks.
| Stage | Time |
|---|---|
| Waiting for silence, to be sure you've finished | 1,400 ms (redemptionMs, adjustable) |
| Hearing: Whisper base transcribes the utterance | 330 ms |
| Thinking, hosted: GPT-OSS 20B on Groq, the demo's route, one round trip | 670 ms |
| Thinking and the first sentence of voice, all in the browser (Qwen 0.5B, Kokoro) | 740 ms |
| Speaking: Kokoro's first sentence ready and playing, with an instant reply | 460 ms |
| Starting up on a return visit, models already stored | about 2 s, nothing downloaded |
So from the moment you stop talking to the first word of the reply:
| Setup | First word after you stop |
|---|---|
| All in the browser | about 2.5 s |
| Hosted replies, local hearing and voice (the demo) | about 2.9 s |
| With an instant reply, the floor for speech alone | about 2.2 s |
The hosted figure adds the route's round trip to the measured speech stages,
since that route answers only the live site. Most of the wait is the pause for
silence, which is what stops the avatar cutting people off while they think.
Lower redemptionMs in speechDetection for a snappier feel, at the cost of
answering half-finished sentences. Without WebGPU, recognition and the voice
run on the CPU and are several times slower.
You do not have to choose once. preloadLocalLlm fetches the in-browser model
behind the conversation while your backend answers, and onLocalLlmReady tells
you when it can take over:
const [localReady, setLocalReady] = useState(false);
<AiVoiceAvatar
// Dropping onSubmit is what hands the conversation over.
onSubmit={localReady ? undefined : askMyBackend}
preloadLocalLlm
onLocalLlmReady={() => setLocalReady(true)}
/>
Nobody waits for a gigabyte to say the first word, and the turns your backend answered are handed to the local model, so it does not restart the conversation from nothing. Useful for a kiosk that must keep working when the wifi drops, a demo on someone else's quota, or a phone, which will never accept the download but can hold a conversation the moment it loads.
The live demo runs exactly this: replies come from a hosted model, and on a desktop with WebGPU the in-browser model downloads during the conversation and takes the floor when it lands. The page says which one is answering at any moment.
Explore the canonical patterns in the examples/ directory:
| Example Pattern | Folder | Highlights & Architecture |
|---|---|---|
| Live Interactive Demo | sandbox | Deploy on Vercel ➔ — Our full-featured interactive testbed featuring live character switching (ananya, aarav), voice persona switching (af_heart, am_michael), real-time diagnostic probe metrics, and Leva 3D lighting controls. |
| Quickstart | examples/quickstart | Minimal, zero-configuration plug-and-play AI voice avatar deployment with built-in studio lighting & sizing. |
| Local Kiosk | examples/local-kiosk | 100% offline on-device retail & restaurant ordering kiosk with embedded menu reasoning. Demonstrates the On-Device Brain; operates without internet access once model weights are locally cached. |
| Connected App | examples/hybrid-cloud | Illustrates the Connected Brain (onSubmit). Bypasses gigabyte-scale local LLM downloads by routing reasoning to OpenAI, Claude, or corporate APIs while keeping ASR, TTS, and 3D lip blending 100% on-device! |
| Voice Only | examples/voice-only | Voice mode with no avatar, like ChatGPT or Gemini voice: the react-ai-voice-avatar/headless hook, an audio-reactive orb, live captions and interruption. No three.js in the bundle, and no backend needed to try it. |
| Groq Voice | examples/groq-voice | Voice mode on Groq's free tier with your own key: hearing, replies and voice start on Groq with nothing to download, the in-browser hearing and voice download behind the conversation (preloadLocalSpeech), and the page offers to switch when they are ready. Shows Groq's live limits. |
| Headless Custom UI | examples/headless-custom-ui | Still renders the 3D avatar; for no avatar at all, see Voice Only above. Demonstrates hiding built-in DOM overlays (hideStatusPill={true}, showCaptions={false}), streaming transcripts into a custom enterprise UI, and controlling voice outputs imperatively via ref.current?.speak(text). |
<AiVoiceAvatar /> Props| Prop | Type | Default | Description |
|---|---|---|---|
avatarPreset | 'ananya' | 'aarav' | 'default' | 'kiosk' | 'ananya' | Built-in 3D character models featuring both female ('ananya') and male ('aarav') voice concierges out of the box with full ARKit facial blendshapes! |
avatarSize | 'sm' | 'md' | 'lg' | number | 'md' (0.48) | Intuitive model sizing presets or custom decimal scaling multiplier applied directly to the 3D humanoid mesh. |
modelSrc | string | undefined | Absolute local path or remote URL to a custom GLTF/GLB humanoid armature avatar model. |
lightingPreset | 'studio' | 'cyberpunk_violet' | 'cool_azure' | 'warm_amber' | 'clean_white' | 'none' | 'studio' | Pre-built cinematic studio lighting atmospheres directly applied to your 3D viewport without manual Three.js configuration! |
systemPrompt | string | "You are Ananya..." | Conversational persona directives and context injected into active LLMs. |
llmModel | string | by language | Hugging Face id for the local WebGPU reasoning model, used only when onSubmit is absent. Defaults to Qwen2.5-0.5B for English and Gemma 3 1B for Hindi, which Qwen that size cannot speak. Set it to pin one model for every language. |
asrModel | string | by language | Hugging Face id for the local Whisper model. Defaults to Whisper base for English and Whisper small for Hindi, which base transcribes badly. Pass "Xenova/whisper-tiny" for a faster download and worse accuracy. |
ttsEngine | 'kokoro' | 'mms' | 'kokoro' | High-fidelity neural voice synthesis engine executing inside dedicated Web Workers. |
ttsVoice | string | by language | Kokoro voice id. af_* and am_* American, bf_* and bm_* British, hf_* and hm_* Hindi (af_heart, am_michael, bf_emma, hf_alpha, hm_omega). Defaults to one matching ttsLanguage. A voice whose language disagrees is corrected with a warning. |
ttsLanguage | 'en-US' | 'en-GB' | 'hi-IN' | 'en-US' | Conversation language. Selects the voice, the recognition model and the phoneme path. Also sets asrLanguage unless you set that yourself. For any other language, pass onSynthesize and use a cloud voice provider. |
asrLanguage | string | ttsLanguage | Language to transcribe. Follows ttsLanguage by default, since a conversation is almost always held in one language. |
showCaptions | boolean | true | Renders a sleek glassmorphic subtitle overlay displaying spoken interaction dialog. |
hideStatusPill | boolean | false | When true, suppresses the default bottom-left microphone interactive control pill. |
listenMode | 'continuous' | 'push-to-talk' | 'continuous' | continuous keeps the mic hot after the avatar finishes speaking naturally, but explicitly clicking Stop forces it off until tapped again. push-to-talk strictly requires manually tapping to start listening for every single turn. |
loadModels | boolean | true | Set false to render the avatar without downloading any models, then flip it true when the visitor engages. For landing pages and widgets most visitors never talk to. Status stays 'loading' until it is true and the models are up, so show your own call to action meanwhile. |
onModelLoaded | () => void | undefined | Fires once the 3D mesh is parsed and in the scene. Parsing a multi-megabyte GLB leaves the canvas empty for a few seconds; use this to hold a placeholder over it. |
allowInterruption | boolean | true | Lets the user talk over the avatar and cut it off mid-sentence. Turn off for a kiosk or noisy room, where the avatar hearing itself through the speakers is worse than waiting. Ignored in push-to-talk. See Taking turns. |
speechDetection | { positiveSpeechThreshold?, negativeSpeechThreshold?, minSpeechMs?, redemptionMs?, preSpeechPadMs? } | see Taking turns | Tunes how the microphone decides someone is talking. Every field optional. The defaults suit a quiet room; a shop floor needs a higher threshold and a longer minSpeechMs. |
onUserInterrupt | () => void | undefined | Fires when the user talks over the avatar and takes the floor. Only fires if the avatar actually had audio playing. |
gestures | boolean | number | true | Hand and arm gestures while the avatar speaks: one hand or both brought up in front of the chest for each phrase, with small beats on stressed syllables, and the arms back at rest when it stops. false keeps the arms still; a number sets the size, 0 to 1.5. Needs a skeleton with LeftArm, LeftForeArm and LeftHand and the right-hand equivalents; custom avatars without them simply do not gesture. |
onAudioLevelChange | (level: number, source: 'mic' | 'tts' | 'idle') => void | undefined | Real-time audio amplitude (0-1) callbacks for the active stream. Essential for building highly responsive, audio-reactive 3D Visualizers and HUDs! Fires with 0 and 'idle' between turns, so a meter falls to rest rather than freezing. |
onSubmit | (text: string) => Promise<string | AsyncIterable<string> | ReadableStream> | undefined | Connected Brain API: Bypasses local LLMs; routes transcribed user microphone strings to your cloud or custom LLM API endpoint. |
preloadLocalLlm | boolean | false | Only meaningful alongside onSubmit, which otherwise skips the local language model download entirely. Set it to fetch that model in the background while your hosted one answers, so the conversation survives a rate limit, an expired quota or a lost network. See Hosted now, local when warm. |
onLocalLlmReady | () => void | undefined | Fires once the model requested by preloadLocalLlm has loaded. Drop onSubmit here to hand the conversation over; the turns your backend answered are carried across, so the local model knows what was already said. |
onTranscribe | (audio: Float32Array) => Promise<string> | undefined | Replaces local speech recognition with your own service, and Whisper is then not downloaded. Receives one utterance as 16 kHz mono samples. |
onSynthesize | (text: string) => Promise<Float32Array | ArrayBuffer> | undefined | Replaces local voice synthesis with your own service, and Kokoro is then not downloaded. Return an encoded MP3/WAV buffer, or raw 24 kHz PCM. Called a sentence at a time. |
onError | (e: AiVoiceAvatarError) => void | undefined | Fires when a stage fails. Carries stage, message and a severity of degraded or fatal. See Handling failures. |
onTranscriptUpdate | (text: string, speaker: 'user' | 'avatar') => void | undefined | Callback delivering real-time microphone transcriptions and assistant spoken utterance strings. |
onStatusChange | (status: string) => void | undefined | Emits live state transitions (loading, idle, listening, thinking, speaking). |
debug | boolean | false | When true, renders an interactive floating GUI (Leva) to inspect and tune individual 3D blendshapes. |
vadAssetPath | string | undefined | Optional URL or local path override for self-hosting @ricky0123/vad-web ONNX asset binaries in airgapped deployments. |
onnxWasmPath | string | undefined | Optional URL override for self-hosting onnxruntime-web WASM distribution files. |
workerBaseUrl | string | undefined | CSP Escape Hatch: if blob: workers are blocked by your server, fetch pre-compiled Web Workers from this URL directory. |
enableLocalAssetProbe | boolean | false | When true, HEAD-checks /ananya.glb in your own public directory before falling back to the CDN. Off by default: with no local copy the probe 404s, and that 404 lands in every visitor's console. |
statusPillStyle | React.CSSProperties | undefined | Optional custom CSS styling & absolute positioning overrides for the interactive Status Pill overlay. |
accentColor | string | undefined | Custom CSS color string (e.g., #38BDF8) for the active status indicator rings and highlights. |
AiVoiceAvatarHandle)Attach a React ref (useRef<AiVoiceAvatarHandle>(null)) to access imperative real-time controls:
interface AiVoiceAvatarHandle {
/** Command the 3D avatar to speak an arbitrary string with synchronized acoustic lip blending */
speak: (text: string) => void;
/** Manually engage microphone recording and Voice Activity Detection (VAD) */
startListening: () => void;
/** Pause active microphone listening */
stopListening: () => void;
/** Instantly interrupt and halt active voice speech synthesis and clear the audio queue */
interrupt: () => void;
/** Manually submit text to the onSubmit handler, simulating a spoken utterance (useful for text-only fallback) */
sendText: (text: string) => void;
/** Wipe multi-turn conversation memory history and caption overlay states */
clearHistory: () => void;
/** Retrieve live Web Audio API AnalyserNode powering real-time spectral lip sync */
getAnalyser: () => AnalyserNode | undefined;
}
sendText)If your users cannot use a microphone (e.g., noisy environments, privacy concerns, or lack of permissions), you can easily wire up a standard text input field to bypass the speech recognition pipeline entirely!
Simply attach a ref and call sendText() to pass a string directly to your onSubmit handler (or local LLM):
const avatarRef = useRef<AiVoiceAvatarHandle>(null);
// In your UI, attach this to a standard <form> submission:
const handleTextSubmit = (userInput: string) => {
avatarRef.current?.sendText(userInput);
}
When you use sendText, the avatar immediately enters the thinking state and processes the interaction exactly as if the user had spoken it aloud.
useAiVoiceAvatar() (headless)import { useAiVoiceAvatar } from 'react-ai-voice-avatar/headless';
Takes the same options as the component's props, except the ones about the 3D
scene and its overlays (avatarPreset, avatarSize, modelSrc,
lightingPreset, gestures, showCaptions, hideStatusPill, onModelLoaded,
debug and the styling props). onStatusChange is replaced by the returned
status. A few options matter mostly without an avatar:
| Option | Type | Description |
|---|---|---|
onAudioLevelChange | (level: number, source: 'mic' | 'tts' | 'idle') => void | Loudness from 0 to 1, every frame, from whichever side is talking. What an orb or waveform animates from. Write it to the DOM through a ref rather than into state. |
onSpeechStart | (text: string) => void | Each sentence of the reply as it starts playing. Captions that keep pace with the voice. |
onTranscriptUpdate | (text: string, speaker: 'user' | 'avatar') => void | What the user said, once transcribed, and the reply in full. |
onInferenceStart / onInferenceEnd | () => void | Around each turn, from the moment the user stops talking to the end of the reply. |
onTtsEngineChange | (engine: 'kokoro' | 'mms' | 'custom') => void | Fires whenever activeTtsEngine changes. Also a prop on the component. |
preloadLocalSpeech | boolean | Download the in-browser hearing and voice behind onTranscribe and onSynthesize, without waiting for them. Drop the adapters once onLocalSpeechReady fires and the local models take over. Also a prop on the component. |
onLocalSpeechReady | () => void | Fires once, when the models requested by preloadLocalSpeech are loaded. |
loadingProgress | (pct: number, label: string) => void | Download progress per model: 'asr', 'kokoro' (or 'tts' for MMS) and 'llm'. They download in parallel, so keep one figure per label. |
It returns:
| Value | Type | Description |
|---|---|---|
status | 'loading' | 'idle' | 'listening' | 'thinking' | 'speaking' | Where the conversation is. 'loading' until the models are up, and until loadModels is true. |
isLoading, isIdle, isListening, isThinking, isSpeaking | boolean | Shorthands for status. |
isReady | boolean | The models are loaded. startListening does nothing before this. |
isLocalSpeechReady | boolean | The in-browser hearing and voice are loaded, whether or not they are in use. |
activeTtsEngine | 'kokoro' | 'mms' | 'custom' | The voice actually speaking. iPhones and iPads get 'mms', a single plainer voice that ignores ttsVoice, because Kokoro runs Safari out of memory; so does any device where Kokoro fails to load. 'custom' is your onSynthesize. Worth showing if your users will compare devices. |
startListening | () => Promise<void> | Opens the microphone and starts listening. Call it from a click: the first call is where the browser asks for permission. In 'continuous' mode the microphone then stays open across turns. |
stopListening | () => void | Closes the microphone. |
interrupt | () => void | Stops the reply mid-sentence and clears what was queued. |
speak | (text: string) => void | Says the text without a model turn: a greeting, a notification. Plays a sentence at a time. |
sendText | (text: string) => void | Typed input, handled as if it had been spoken. |
clearHistory | () => void | Forgets the conversation so far. |
micError | string | null | Why the microphone could not open, such as a denied permission. null otherwise. |
analyser | AnalyserNode | undefined | The reply's audio, for a frequency visualiser. onAudioLevelChange is simpler if you only need loudness. |
It also returns several refs (currentSpeechTextRef, audioContextRef and
others) that the 3D component uses for lip-sync. A voice UI can ignore them.
We actively welcome community contributions. CONTRIBUTING.md has the local development guide, and ROADMAP.md has what is worth doing next, why, and what is already known about each item — including the measurements behind the open questions.
The largest pieces currently open:
armRig.ts.scripts/verify-pack.mjs).Hindi speech and VAD sensitivity tuning (speechDetection) were previously
listed here and have both shipped.
This library heavily relies on modern Web APIs (WebGPU, WebGL, Web Audio, and Web Workers). It gracefully degrades when certain APIs are unavailable.
| Browser | OS | 3D Rendering (WebGL) | Voice Synthesis (WebGPU/WASM) | Voice Recognition (Web Audio) | Status |
|---|---|---|---|---|---|
| Chrome / Edge | Windows, macOS, Android | ✅ Native | ✅ WebGPU (Ultra Fast) | ✅ Native | 🟢 Tier 1 (Recommended) |
| Safari / iOS | macOS, iOS | ✅ Native | 🔄 Lightweight Models by Default | ✅ Native | 🟡 Supported |
| Firefox | Windows, macOS | ✅ Native | ⚠️ WASM Fallback | ✅ Native | 🟡 Tier 2 (Slower TTS) |
[!NOTE]
- WebGPU is currently enabled by default in Chrome/Edge. On browsers without WebGPU, the library automatically falls back to WASM execution.
- iOS/Safari Preemptive Fallback: Safari and iOS impose strict memory limits that cause 80MB+ models (like Kokoro) to crash the tab. The engine automatically preempts this by forcing the lightweight MMS TTS model (~30MB) on iOS devices, and tracks crash breadcrumbs to prevent OOM reload loops.
- Strict CSP Environments: Safari and Firefox may block
blob:worker execution depending on your Content-Security-Policy headers. If this occurs, host the.worker.jsfiles statically and pass their base path via theworkerBaseUrlprop.
Running Neural Networks in the browser requires capable hardware.
| Deployment Mode | Min RAM | GPU Requirement | Recommended Devices |
|---|---|---|---|
| Connected Brain (ASR + TTS only) | 4GB | None (WASM Fallback ok) | iPhone 11+, Mid-range Android (2021+), Any Laptop |
| Full Local AI (ASR + 500M LLM + TTS) | 8GB | WebGPU Support Preferred | iPhone 13 Pro+, High-end Android (Snapdragon 8 Gen 1+), M1/M2 Macs, Modern PCs |
[!TIP] Mobile Memory Limits: Mobile browsers rigidly enforce memory limits per tab (often terminating tabs exceeding ~1GB). If your mobile app crashes "after some time", ensure you are utilizing the
Connected Brainmode (onSubmitAPI) which offloads the heavy LLM memory footprint to your server while keeping ultra-fast lip-sync and TTS local.
MIT © React AI Voice Avatar Contributors.
TypeScript
74.9%
JavaScript
19.9%
Python
3.5%