NavodPeiris/gsplat-talkinghead

LAM Gaussian-Splat Avatars for AI voice agents. Realistic avatars for interview/support/sales agents.

TypeScript

1

13 commits

updated Sep 29, 2026

See the code

See what people are saying

SourceMessageScoreDate

Gaussian Splat Avatars for Realtime AI Agents

2

Oct 1, 2026

README

showcase NPM Version NPM Downloads GitHub Repo stars

gsplat-talkinghead

Lip-Synced Gaussian-Splat avatar components for AI voice agents. Drop it into any React app, pick a provider, and hand it your credentials — everything else is handled internally. No infrastructure provisioning — Gaussian-splat rendering and wav2arkit neural lipsync both run directly in the browser.

Supported providers: OpenAI Realtime, OpenAI GPT-Live, Qwen Realtime (Alibaba Cloud), ElevenLabs Conversational AI, Vapi, LiveKit Agents.

https://github.com/user-attachments/assets/670b67fe-e6f4-4bb6-a58e-66467fd7b977

Requirements

Peer dependencyVersion
@myned-ai/gsplat-flame-avatar-renderer≥ 1.4.0
react≥ 18
react-dom≥ 18
@elevenlabs/react≥ 1.0.2
@vapi-ai/web≥ 2.5.2
livekit-client2.16.1

Optional depending on your provider usecase:

@elevenlabs/react
@vapi-ai/web
livekit-client


Installation

npm install gsplat-talkinghead

Environment setup

Provider secret keys belong on your backend, never in browser code. Each component asks for credentials through a callback (getEphemeralKey, getConversationToken, getToken, createSession): implement it by calling your own backend route, which uses the secret to mint a short-lived token (or forward a WebRTC offer) and returns only that to the browser.

ProviderBackend env (secret)Browser env (public)What your backend serves
OpenAI RealtimeOPENAI_API_KEY—getEphemeralKey → ephemeral key from POST /v1/realtime/client_secrets
OpenAI GPT-LiveOPENAI_API_KEY—createSession → SDP answer from POST /v1/live/sessions
Qwen realtimeDASHSCOPE_API_KEY, DASHSCOPE_WORKSPACE_ID—createSession → SDP answer from Model Studio's WebRTC endpoint
ElevenLabsELEVENLABS_API_KEYagent IDgetConversationToken → conversation token for your agent
Vapi—VAPI_PUBLIC_KEYNothing — the public key is browser-safe; restrict allowed domains in the Vapi dashboard
LiveKitLIVEKIT_API_KEY, LIVEKIT_API_SECRETLIVEKIT_URLgetToken → room access token

Example files (only include the providers you use):

# backend/.env — server only, never shipped to the browser
OPENAI_API_KEY=sk-...
DASHSCOPE_API_KEY=sk-...
DASHSCOPE_WORKSPACE_ID=ws-...
ELEVENLABS_API_KEY=...
LIVEKIT_API_KEY=API...
LIVEKIT_API_SECRET=...
# frontend/.env — public values only (use your bundler's prefix:
# VITE_ for Vite, NEXT_PUBLIC_ for Next.js, REACT_APP_ for Create React App)
VITE_VAPI_PUBLIC_KEY=...
VITE_LIVEKIT_URL=wss://<project>.livekit.cloud

Anything with a bundler prefix is embedded in your JavaScript bundle and visible to users, so never give a secret key one of those prefixes.

To try the providers locally without a backend, see the test-env playground, which reads keys from its own .env (for local testing only).


Providers


OpenAI

Uses the OpenAI Realtime API over WebRTC. Requires a server-side endpoint to mint ephemeral session keys.

import { OpenAIRealtimeAgent } from "gsplat-talkinghead/openai";
import type { OpenAIRealtimeTool } from "gsplat-talkinghead";

const tools: OpenAIRealtimeTool[] = [
  {
    name: "get_product_price",
    description: "Returns the current price of a product.",
    parameters: {
      type: "object",
      properties: {
        product_name: { type: "string", description: "Name of the product" },
      },
      required: ["product_name"],
    },
    handler: ({ product_name }) => {
      const prices: Record<string, string> = {
        "apple iphone 15 pro max": "$1,199",
        "samsung galaxy s23 ultra": "$1,099",
        "sony wh-1000xm5": "$349",
        "dell xps 13": "$999",
        "amazon echo dot": "$49",
      };
      const price =
        prices[(product_name as string).toLowerCase()] ?? "Price not available";
      return { product_name, price };
    },
  },
];

<OpenAIRealtimeAgent
  backgroundImages={["/niceBG.jpg"]}
  agentVoice="nova"
  tools={tools}
  getEphemeralKey={async () => {
    const res = await fetch("/api/realtime-session");
    const { client_secret } = await res.json();
    return client_secret;
  }}
/>;

Backend — ephemeral key endpoint

// app/api/realtime-session/route.ts  (Next.js App Router)
import OpenAI from "openai";

const openai = new OpenAI({
  apiKey: process.env.OPENAI_API_KEY, // server-only secret
});

const sys_prompt = `
# ROLE
You are a product recommendation assistant for Amazon who answers user questions and recommends products based on their preferences.
at initial greeting, say 'Hello! I am your product specialist at Amazon. I can help you find products — feel free to tell me what you are looking for!'
DO NOT repeat it again.

These are currently available products:
1. Apple iPhone 15 Pro Max - iPhone 15 Pro Max delivers premium performance with a lightweight titanium design, stunning 6.7-inch Super Retina XDR display with ProMotion, and the powerful A17 Pro chip. Capture incredible detail with its advanced Pro camera system featuring a 48MP main sensor and 5x optical zoom.
2. Samsung Galaxy S23 Ultra - A high-end Android phone with a stunning display, versatile cameras, and long battery life.
3. Sony WH-1000XM5 Wireless Noise-Canceling Headphones - Premium headphones with industry-leading noise cancellation, exceptional sound quality, and comfortable design.
4. Dell XPS 13 Laptop - A sleek and powerful ultrabook with a stunning InfinityEdge display, Intel Core i7 processor, and long battery life.
5. Amazon Echo Dot (5th Gen) - A compact smart speaker with Alexa voice assistant, perfect for controlling smart home devices, playing music, and getting information.
Always recommend products based on the user's preferences and needs. If the user asks for a specific product, provide information about it and suggest similar alternatives if available.
When the user says goodbye or is done, say "this is the end" to close the session.

# TOOLS
if user asks for product prices, use the get_product_price tool to retrieve the current price of the product and include it in your response.
`;

export async function GET() {
  const session = await openai.realtime.clientSecrets.create({
    session: {
      type: "realtime",
      model: "gpt-realtime-mini-2025-10-06",
      instructions: sys_prompt,
    },
  });

  return Response.json({ client_secret: session.value });
}

Props

PropTypeDefaultDescription
getEphemeralKey() => Promise<string>requiredCalled once per connection. Must resolve to an OpenAI ephemeral key from your backend.
toolsOpenAIRealtimeTool[][]Browser-side function tools: name, description, JSON-schema parameters, and a handler.
agentVoicestring"sage"OpenAI Realtime voice, e.g. alloy ash ballad coral echo sage shimmer verse.

The model and system prompt (instructions) are set by your backend when it mints the ephemeral key, as in the example above.

Runs full wav2arkit neural lipsync — OpenAI's WebRTC session exposes a real remote MediaStream.


OpenAI GPT-Live

gpt-live-1 is OpenAI's full-duplex voice model: it keeps listening while it speaks and hands reasoning and tool calls to a backend model. It uses its own endpoint (POST /v1/live/sessions), not the Realtime API, so it has its own component.

GPT-Live sessions are created with a standard API key, so the WebRTC handshake goes through your backend: the component gives you the browser's SDP offer, your server creates the session and returns the SDP answer.

import { OpenAILiveAgent } from "gsplat-talkinghead/openai";

<OpenAILiveAgent
  avatar="Jane"
  createSession={async (offerSdp) => {
    const res = await fetch("/api/live-session", {
      method: "POST",
      body: offerSdp,
    });
    return res.text(); // SDP answer
  }}
  tools={tools} // same shape as <OpenAIRealtimeAgent tools>
/>;
// Your backend — POST /api/live-session
const res = await fetch("https://api.openai.com/v1/live/sessions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.OPENAI_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    session: {
      model: "gpt-live-1",
      instructions: "You are a helpful assistant.",
      audio: { output: { voice: "willow" } },
      delegation: { type: "responses", responses: { model: "gpt-5.6-terra" } },
    },
    transport: { type: "webrtc", sdp: offerSdp },
  }),
});
const { transport } = await res.json();
return transport.sdp;

Model, voice, instructions and delegation are fixed when the session starts, so they're configured on the backend rather than as component props.

Props

PropTypeDefaultDescription
createSession(offerSdp: string) => Promise<string>requiredSends the SDP offer to your backend and resolves to GPT-Live's SDP answer.
toolsOpenAIRealtimeTool[][]Browser-side function tools. Registered on the Responses delegation backend; needs delegation.type: "responses".

Runs full wav2arkit neural lipsync — agent audio arrives as a WebRTC media track. Unlike OpenAIRealtimeAgent, it doesn't send an opening prompt on connect, so the user may need to speak first.


Qwen realtime (Alibaba Cloud)

qwen3.8-omni-flash-realtime is Alibaba's full-duplex speech-to-speech model, at roughly $0.93 / $1.87 per 1M audio tokens in / out (Singapore). It connects over WebRTC.

The WebRTC SDP exchange is authenticated with your DashScope API key, so it goes through your backend: the component gives you the browser's SDP offer, your server forwards it and returns the answer.

import { QwenRealtimeAgent } from "gsplat-talkinghead/qwen";

<QwenRealtimeAgent
  avatar="Jane"
  instructions="You are a helpful assistant."
  tools={tools} // same shape as <OpenAIRealtimeAgent tools>
  createSession={async (offerSdp) => {
    const res = await fetch("/api/qwen-session", {
      method: "POST",
      body: offerSdp,
    });
    return res.text(); // SDP answer
  }}
/>;
// Your backend — POST /api/qwen-session
const res = await fetch(
  `https://${WORKSPACE_ID}.ap-southeast-1.maas.aliyuncs.com/api/v1/webrtc/realtime?model=qwen3.8-omni-flash-realtime`,
  {
    method: "POST",
    headers: {
      "Content-Type": "application/sdp",
      Authorization: `Bearer ${process.env.DASHSCOPE_API_KEY}`,
    },
    body: offerSdp,
  },
);
return res.text();

Use the cn-beijing host and a Beijing API key for the China (Beijing) region.

Props

PropTypeDefaultDescription
createSession(offerSdp: string) => Promise<string>requiredSends the SDP offer to your backend and resolves to Qwen's SDP answer.
instructionsstring—System prompt.
voicestring"Tina"A voice from Model Studio's voice list.
toolsOpenAIRealtimeTool[][]Browser-side function tools.
turnDetection"server_vad" | "semantic_vad""server_vad"WebRTC only supports server-side turn detection.
transcribeInputbooleanfalseAlso transcribe the user's speech (billed separately).
greetingstring | falsegreet per instructionsInstruction for the agent's opening turn; false waits for the user.

Runs full wav2arkit neural lipsync — agent audio arrives as a WebRTC media track.

The agent speaks first by default (see greeting).


ElevenLabs

Uses the ElevenLabs Conversational AI SDK. Configure your agent in the ElevenLabs dashboard and pass its ID here.

import { ElevenLabsAvatarAgent } from "gsplat-talkinghead/elevenlabs";

// Public agent (no auth required)
<ElevenLabsAvatarAgent
  agentId="your-agent-id"
/>

// Private agent (fetch a short-lived token server-side)
<ElevenLabsAvatarAgent
  backgroundImages={["/niceBG.jpg"]}
  agentId="your-agent-id"
  getConversationToken={async () => {
    const res = await fetch("/api/elevenlabs-token");
    const { token } = await res.json();
    return token;
  }}
/>

Backend — conversation token endpoint

// app/api/elevenlabs-token/route.ts
export async function GET(req: Request) {
  const agentId = new URL(req.url).searchParams.get("agentId");
  const res = await fetch(
    `https://api.elevenlabs.io/v1/convai/conversation/token?agent_id=${agentId}`,
    { headers: { "xi-api-key": process.env.ELEVENLABS_API_KEY! } },
  );
  return Response.json(await res.json());
}

Props

PropTypeDefaultDescription
agentIdstringrequiredAgent ID from the ElevenLabs dashboard.
getConversationToken() => Promise<string>—Required for private agents. Returns a short-lived WebRTC token.
clientToolsRecord<string, (...args) => any>—Client-side tools exposed to the agent.

Lipsync note: the ElevenLabs SDK only exposes frequency-magnitude and volume scalars, never raw waveform audio, so the wav2arkit neural model can't run for this provider. ElevenLabs falls back to coarse, volume-driven mouth movement instead of full lipsync.


Vapi

Uses the Vapi Web SDK. Configure your assistant in the Vapi dashboard or pass an inline configuration object.

import { VapiAvatarAgent } from "gsplat-talkinghead/vapi";

// Using a pre-configured assistant ID
<VapiAvatarAgent
  publicKey="your-vapi-public-key"
  assistantId="your-assistant-id"
/>

// Using an inline assistant config
<VapiAvatarAgent
  backgroundImages={["/niceBG.jpg"]}
  publicKey="your-vapi-public-key"
  assistant={{
    model: {
      provider: "openai",
      model: "gpt-4o-mini",
      messages: [{ role: "system", content: "You are a helpful assistant." }],
    },
    voice: { provider: "11labs", voiceId: "sarah" },
  }}
/>

Props

PropTypeDefaultDescription
publicKeystringrequiredYour Vapi public key (safe to expose in the browser).
assistantIdstring—Pre-configured assistant ID. Mutually exclusive with assistant.
assistantRecord<string, any>—Inline assistant configuration object. Mutually exclusive with assistantId.

Runs full wav2arkit neural lipsync — Vapi's Web SDK exposes a real remote MediaStream.


LiveKit

Uses LiveKit Agents over WebRTC. Your LiveKit agent must be running server-side and the token must grant access to the correct room.

import { LiveKitAvatarAgent } from "gsplat-talkinghead/livekit";

<LiveKitAvatarAgent
  backgroundImages={["/niceBG.jpg"]}
  serverUrl="wss://my-project.livekit.cloud"
  getToken={async () => {
    const res = await fetch("/api/livekit-token");
    const { token } = await res.json();
    return token;
  }}
/>;

Backend — participant token endpoint

// app/api/livekit-token/route.ts
import { AccessToken } from "livekit-server-sdk";

export async function GET() {
  const token = new AccessToken(
    process.env.LIVEKIT_API_KEY!,
    process.env.LIVEKIT_API_SECRET!,
    { identity: `user-${Date.now()}` },
  );
  token.addGrant({ roomJoin: true, room: "agent-room" });
  return Response.json({ token: await token.toJwt() });
}

Props

PropTypeDefaultDescription
serverUrlstringrequiredLiveKit server WebSocket URL, e.g. wss://my-project.livekit.cloud.
getToken() => Promise<string>requiredReturns a short-lived participant token. Generate server-side with the LiveKit server SDK.

Runs full wav2arkit neural lipsync — LiveKit's client exposes a real remote MediaStream.


Shared props

All provider components accept these additional props:

PropTypeDefaultDescription
avatarAvatarPreset"Jane"Built-in avatar: "Jack", "Jane", "John" or "Sasha". See Avatars below.
assetsPathstring—URL/path to a custom Gaussian-splat avatar asset bundle. Takes precedence over avatar.
emotionAvatarEmotion"neutral"Facial emotion over lipsync: "neutral", "happy", "sad" or "thinking". See Emotions.
backgroundImagesstring[][]Image URLs for the background. One is chosen at random each mount. Transparent when omitted.
onSessionEnd() => void—Called when the session ends (end phrase detected, timeout, or user clicked End).
endSessionPhrasestring"this is the end"Case-insensitive substring the component watches for in the agent's transcript to end the session.
sessionTimeoutnumber600000Hard timeout in milliseconds.
classNamestring—Extra CSS class names on the outermost container div.

Avatars

The library ships with four preset avatars — Jack, Jane, John and Sasha — picked with the avatar prop. Jane is used when you pass neither avatar nor assetsPath. Nothing to configure, nothing to host: the bundles are served from this project's GitHub repo (main branch) via jsDelivr's GitHub CDN. They aren't included in the npm package, which keeps the install small.

// Default preset (Jane)
<OpenAIRealtimeAgent getEphemeralKey={...} />

// Another preset
<OpenAIRealtimeAgent avatar="Jack" getEphemeralKey={...} />

The preset list is also exported, e.g. to build an avatar picker:

import { AVATAR_PRESETS, type AvatarPreset } from "gsplat-talkinghead";

AVATAR_PRESETS; // ["Jack", "Jane", "John", "Sasha"]

To serve the presets from your own origin instead of jsDelivr (strict CSP or offline use), copy the repo's gsplat-talkinghead/assets/ folder somewhere public and point the library at it once at startup:

import { configureAvatarPresets } from "gsplat-talkinghead";

configureAvatarPresets({ baseUrl: "/avatars" }); // loads /avatars/Jack.zip, …

To use your own avatar instead, host a Gaussian-splat asset bundle (compatible with @myned-ai/gsplat-flame-avatar-renderer) and pass its URL as assetsPath:

<OpenAIRealtimeAgent
  getEphemeralKey={...}
  assetsPath="https://your-cdn.example.com/avatars/your-avatar-bundle"
/>

Emotions

Every component takes an emotion prop — "neutral", "happy", "sad" or "thinking" — layered on top of lipsync as ARKit blendshape offsets (smile, brows, eye squint/gaze), so the avatar can smile while it talks. Changes blend in over ~0.3 s.

Drive it from your app:

const [emotion, setEmotion] = useState<AvatarEmotion>("neutral");
<OpenAIRealtimeAgent emotion={emotion} ... />

Or let the model set it with the built-in set_emotion tool (for providers with browser-side tools: OpenAIRealtimeAgent, OpenAILiveAgent, QwenRealtimeAgent). It returns to neutral after 8 s by default (createEmotionTool(setEmotion, { resetAfterMs })):

import { createEmotionTool, type AvatarEmotion } from "gsplat-talkinghead";

const [emotion, setEmotion] = useState<AvatarEmotion>("neutral");
const allTools = useMemo(() => [...tools, createEmotionTool(setEmotion)], []);

<OpenAIRealtimeAgent
  emotion={emotion}
  tools={allTools}
  // + in your instructions: "Call set_emotion when your feelings change:
  //   happy for good news, sad for bad news, thinking while you consider options."
/>;

The expressions are defined in EMOTION_BLENDSHAPES (exported) if you want to see or copy the exact offsets.


Advanced: adapter hooks

For full layout control, use AvatarAgent directly with an adapter hook. This lets you compose the avatar into your own UI without the built-in container styles.

import { AvatarAgent } from "gsplat-talkinghead";
import { useVapiAdapter } from "gsplat-talkinghead/vapi";

function MyPage() {
  const adapter = useVapiAdapter({
    publicKey: "your-vapi-public-key",
    assistantId: "your-assistant-id",
  });

  return (
    <div className="my-layout">
      <AvatarAgent adapter={adapter} className="h-[600px]" />
    </div>
  );
}

All adapter hooks follow the same pattern:

useOpenAIRealtimeAdapter(options); // → SessionAdapter
useElevenLabsAdapter(options); // → SessionAdapter
useVapiAdapter(options); // → SessionAdapter
useLiveKitAdapter(options); // → SessionAdapter

A SessionAdapter exposes remoteStream for full neural lipsync, plus two optional fallbacks for platforms that can't produce a MediaStream: subscribeToRemoteAudio (push already-decoded PCM straight into wav2arkit, for a platform whose SDK decodes raw PCM itself) and getRemoteAudioLevel (a 0-1 volume scalar for coarse, volume-only mouth movement — used by ElevenLabs). Implement one of these if you're writing a custom adapter for a platform without a real audio stream.


How it works

Component mounts
      │
      ├── Avatar bundle + wav2arkit model download ──► "Downloading assets" overlay, Start disabled
      │
User clicks Start
      │
      ▼
Provider adapter connects (WebRTC / WebSocket)
      │
      ├── Audio ──► resample to 16kHz ──► wav2arkit (onnxruntime-web) ──► ARKit blendshapes ─┐
      │             (MediaStream tap, or pushed raw PCM — see Advanced: adapter hooks)          │
      │                                                                                         ▼
      │   emotion prop ──► blendshape offsets ──► blink + subtle head sway ──► Gaussian-splat renderer
      │
      ├── Transcript ──► endSessionPhrase check ──► onSessionEnd(), End
      │
      ├── User clicks End ──► End
      │
      └── sessionTimeout ──► onSessionEnd(), End

End ──► face, emotion and head motion reset to rest

The wav2arkit ONNX model (myned-ai/wav2arkit_cpu, Apache-2.0) is fetched and cached in the browser on first use — no server component required.


Package structure

src/
├── index.ts                     ← public exports (AvatarAgent, presets, emotions, types)
├── openai.ts · qwen.ts · vapi.ts · elevenlabs.ts · livekit.ts   ← per-provider entry points
├── types.ts                     ← shared prop types
├── AvatarAgent.tsx              ← platform-agnostic core component
├── OpenAIRealtimeAgent.tsx      ← provider convenience wrappers
├── OpenAILiveAgent.tsx
├── QwenRealtimeAgent.tsx
├── ElevenLabsAvatarAgent.tsx
├── VapiAvatarAgent.tsx
├── LiveKitAvatarAgent.tsx
├── adapters/
│   ├── SessionAdapter.ts        ← adapter interface
│   ├── openai-realtime/ · openai-live/ · qwen/
│   └── elevenlabs/ · vapi/ · livekit/
├── avatar/
│   ├── GaussianAvatarController.ts  ← wraps the renderer: blendshapes, blink, emotion, head sway
│   ├── LazyAvatarController.ts      ← lazy-loads the renderer
│   ├── AvatarContainer.tsx          ← mount point + background image
│   ├── presets.ts                   ← built-in avatars (Jack, Jane, John, Sasha)
│   ├── emotions.ts                  ← emotion blendshapes + set_emotion tool
│   └── transparentCanvas.ts         ← transparent WebGL canvas over the background
├── session/
│   ├── useAgentSession.ts       ← OpenAI Realtime WebRTC session
│   └── codecUtils.ts
├── ui/                          ← Toolbar, StatusBadge, AssetsLoader, AudioBars
├── constants/
│   └── arkit.ts                 ← ARKit blendshape names, neutral pose
└── audio/
    ├── wav2arkit/
    │   ├── modelLoader.ts       ← fetches + caches the ONNX model
    │   ├── resample.ts          ← resamples audio to 16kHz
    │   ├── inferenceEngine.ts   ← onnxruntime-web session + load progress
    │   └── liveLipsync.ts       ← streaming/paced inference pipeline
    ├── useAvatarLipsync.ts      ← wires a MediaStream into wav2arkit
    ├── usePushAudioLipsync.ts   ← wires pushed raw PCM into wav2arkit
    ├── useVolumeFallbackLipsync.ts ← coarse fallback for stream-less adapters
    ├── delayedAudioPlayback.ts  ← delays audio to line up with lipsync
    └── useAudio.ts              ← mic monitoring

Contribution Guide

To contribute to package source code, raise PRs to the main branch.


Sponsoring ❤️

gsplat-talkinghead is free, open-source, and maintained in personal time.

If your product ships AI-powered conversations and this library saves you weeks of WebRTC wrangling, lip-sync work, and provider integration — consider sponsoring. Even a small recurring amount keeps the library maintained, more providers supported, and bugs fixed fast.

Sponsor on GitHub →


Citing this project

If you use gsplat-talkinghead in academic work, a research demo, or a published product, a citation or acknowledgement is appreciated.

BibTeX

@software{peiris2026gsplattalkinghead,
  author  = {Peiris, Navod},
  title   = {gsplat-talkinghead: LAM Gaussian-Splat Avatars for AI voice agents},
  year    = {2026},
  url     = {https://github.com/NavodPeiris/gsplat-talkinghead},
  note    = {npm: gsplat-talkinghead}
}

Plain text

Navod Peiris. gsplat-talkinghead: LAM Gaussian-Splat Avatars for AI voice agents. 2026. https://github.com/NavodPeiris/gsplat-talkinghead

Acknowledgement (for README or paper footnote)

Gaussian-splat avatar and lip-sync powered by gsplat-talkinghead.

NavodPeiris/gsplat-talkinghead

LAM Gaussian-Splat Avatars for AI voice agents. Realistic avatars for interview/support/sales agents.

TypeScript

1

13 commits

updated Sep 29, 2026

See the code

See what people are saying

SourceMessageScoreDate

Gaussian Splat Avatars for Realtime AI Agents

2

Oct 1, 2026

README

showcase NPM Version NPM Downloads GitHub Repo stars

gsplat-talkinghead

Lip-Synced Gaussian-Splat avatar components for AI voice agents. Drop it into any React app, pick a provider, and hand it your credentials — everything else is handled internally. No infrastructure provisioning — Gaussian-splat rendering and wav2arkit neural lipsync both run directly in the browser.

Supported providers: OpenAI Realtime, OpenAI GPT-Live, Qwen Realtime (Alibaba Cloud), ElevenLabs Conversational AI, Vapi, LiveKit Agents.

https://github.com/user-attachments/assets/670b67fe-e6f4-4bb6-a58e-66467fd7b977

Requirements

Peer dependencyVersion
@myned-ai/gsplat-flame-avatar-renderer≥ 1.4.0
react≥ 18
react-dom≥ 18
@elevenlabs/react≥ 1.0.2
@vapi-ai/web≥ 2.5.2
livekit-client2.16.1

Optional depending on your provider usecase:

@elevenlabs/react
@vapi-ai/web
livekit-client


Installation

npm install gsplat-talkinghead

Environment setup

Provider secret keys belong on your backend, never in browser code. Each component asks for credentials through a callback (getEphemeralKey, getConversationToken, getToken, createSession): implement it by calling your own backend route, which uses the secret to mint a short-lived token (or forward a WebRTC offer) and returns only that to the browser.

ProviderBackend env (secret)Browser env (public)What your backend serves
OpenAI RealtimeOPENAI_API_KEY—getEphemeralKey → ephemeral key from POST /v1/realtime/client_secrets
OpenAI GPT-LiveOPENAI_API_KEY—createSession → SDP answer from POST /v1/live/sessions
Qwen realtimeDASHSCOPE_API_KEY, DASHSCOPE_WORKSPACE_ID—createSession → SDP answer from Model Studio's WebRTC endpoint
ElevenLabsELEVENLABS_API_KEYagent IDgetConversationToken → conversation token for your agent
Vapi—VAPI_PUBLIC_KEYNothing — the public key is browser-safe; restrict allowed domains in the Vapi dashboard
LiveKitLIVEKIT_API_KEY, LIVEKIT_API_SECRETLIVEKIT_URLgetToken → room access token

Example files (only include the providers you use):

# backend/.env — server only, never shipped to the browser
OPENAI_API_KEY=sk-...
DASHSCOPE_API_KEY=sk-...
DASHSCOPE_WORKSPACE_ID=ws-...
ELEVENLABS_API_KEY=...
LIVEKIT_API_KEY=API...
LIVEKIT_API_SECRET=...
# frontend/.env — public values only (use your bundler's prefix:
# VITE_ for Vite, NEXT_PUBLIC_ for Next.js, REACT_APP_ for Create React App)
VITE_VAPI_PUBLIC_KEY=...
VITE_LIVEKIT_URL=wss://<project>.livekit.cloud

Anything with a bundler prefix is embedded in your JavaScript bundle and visible to users, so never give a secret key one of those prefixes.

To try the providers locally without a backend, see the test-env playground, which reads keys from its own .env (for local testing only).


Providers


OpenAI

Uses the OpenAI Realtime API over WebRTC. Requires a server-side endpoint to mint ephemeral session keys.

import { OpenAIRealtimeAgent } from "gsplat-talkinghead/openai";
import type { OpenAIRealtimeTool } from "gsplat-talkinghead";

const tools: OpenAIRealtimeTool[] = [
  {
    name: "get_product_price",
    description: "Returns the current price of a product.",
    parameters: {
      type: "object",
      properties: {
        product_name: { type: "string", description: "Name of the product" },
      },
      required: ["product_name"],
    },
    handler: ({ product_name }) => {
      const prices: Record<string, string> = {
        "apple iphone 15 pro max": "$1,199",
        "samsung galaxy s23 ultra": "$1,099",
        "sony wh-1000xm5": "$349",
        "dell xps 13": "$999",
        "amazon echo dot": "$49",
      };
      const price =
        prices[(product_name as string).toLowerCase()] ?? "Price not available";
      return { product_name, price };
    },
  },
];

<OpenAIRealtimeAgent
  backgroundImages={["/niceBG.jpg"]}
  agentVoice="nova"
  tools={tools}
  getEphemeralKey={async () => {
    const res = await fetch("/api/realtime-session");
    const { client_secret } = await res.json();
    return client_secret;
  }}
/>;

Backend — ephemeral key endpoint

// app/api/realtime-session/route.ts  (Next.js App Router)
import OpenAI from "openai";

const openai = new OpenAI({
  apiKey: process.env.OPENAI_API_KEY, // server-only secret
});

const sys_prompt = `
# ROLE
You are a product recommendation assistant for Amazon who answers user questions and recommends products based on their preferences.
at initial greeting, say 'Hello! I am your product specialist at Amazon. I can help you find products — feel free to tell me what you are looking for!'
DO NOT repeat it again.

These are currently available products:
1. Apple iPhone 15 Pro Max - iPhone 15 Pro Max delivers premium performance with a lightweight titanium design, stunning 6.7-inch Super Retina XDR display with ProMotion, and the powerful A17 Pro chip. Capture incredible detail with its advanced Pro camera system featuring a 48MP main sensor and 5x optical zoom.
2. Samsung Galaxy S23 Ultra - A high-end Android phone with a stunning display, versatile cameras, and long battery life.
3. Sony WH-1000XM5 Wireless Noise-Canceling Headphones - Premium headphones with industry-leading noise cancellation, exceptional sound quality, and comfortable design.
4. Dell XPS 13 Laptop - A sleek and powerful ultrabook with a stunning InfinityEdge display, Intel Core i7 processor, and long battery life.
5. Amazon Echo Dot (5th Gen) - A compact smart speaker with Alexa voice assistant, perfect for controlling smart home devices, playing music, and getting information.
Always recommend products based on the user's preferences and needs. If the user asks for a specific product, provide information about it and suggest similar alternatives if available.
When the user says goodbye or is done, say "this is the end" to close the session.

# TOOLS
if user asks for product prices, use the get_product_price tool to retrieve the current price of the product and include it in your response.
`;

export async function GET() {
  const session = await openai.realtime.clientSecrets.create({
    session: {
      type: "realtime",
      model: "gpt-realtime-mini-2025-10-06",
      instructions: sys_prompt,
    },
  });

  return Response.json({ client_secret: session.value });
}

Props

PropTypeDefaultDescription
getEphemeralKey() => Promise<string>requiredCalled once per connection. Must resolve to an OpenAI ephemeral key from your backend.
toolsOpenAIRealtimeTool[][]Browser-side function tools: name, description, JSON-schema parameters, and a handler.
agentVoicestring"sage"OpenAI Realtime voice, e.g. alloy ash ballad coral echo sage shimmer verse.

The model and system prompt (instructions) are set by your backend when it mints the ephemeral key, as in the example above.

Runs full wav2arkit neural lipsync — OpenAI's WebRTC session exposes a real remote MediaStream.


OpenAI GPT-Live

gpt-live-1 is OpenAI's full-duplex voice model: it keeps listening while it speaks and hands reasoning and tool calls to a backend model. It uses its own endpoint (POST /v1/live/sessions), not the Realtime API, so it has its own component.

GPT-Live sessions are created with a standard API key, so the WebRTC handshake goes through your backend: the component gives you the browser's SDP offer, your server creates the session and returns the SDP answer.

import { OpenAILiveAgent } from "gsplat-talkinghead/openai";

<OpenAILiveAgent
  avatar="Jane"
  createSession={async (offerSdp) => {
    const res = await fetch("/api/live-session", {
      method: "POST",
      body: offerSdp,
    });
    return res.text(); // SDP answer
  }}
  tools={tools} // same shape as <OpenAIRealtimeAgent tools>
/>;
// Your backend — POST /api/live-session
const res = await fetch("https://api.openai.com/v1/live/sessions", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.OPENAI_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    session: {
      model: "gpt-live-1",
      instructions: "You are a helpful assistant.",
      audio: { output: { voice: "willow" } },
      delegation: { type: "responses", responses: { model: "gpt-5.6-terra" } },
    },
    transport: { type: "webrtc", sdp: offerSdp },
  }),
});
const { transport } = await res.json();
return transport.sdp;

Model, voice, instructions and delegation are fixed when the session starts, so they're configured on the backend rather than as component props.

Props

PropTypeDefaultDescription
createSession(offerSdp: string) => Promise<string>requiredSends the SDP offer to your backend and resolves to GPT-Live's SDP answer.
toolsOpenAIRealtimeTool[][]Browser-side function tools. Registered on the Responses delegation backend; needs delegation.type: "responses".

Runs full wav2arkit neural lipsync — agent audio arrives as a WebRTC media track. Unlike OpenAIRealtimeAgent, it doesn't send an opening prompt on connect, so the user may need to speak first.


Qwen realtime (Alibaba Cloud)

qwen3.8-omni-flash-realtime is Alibaba's full-duplex speech-to-speech model, at roughly $0.93 / $1.87 per 1M audio tokens in / out (Singapore). It connects over WebRTC.

The WebRTC SDP exchange is authenticated with your DashScope API key, so it goes through your backend: the component gives you the browser's SDP offer, your server forwards it and returns the answer.

import { QwenRealtimeAgent } from "gsplat-talkinghead/qwen";

<QwenRealtimeAgent
  avatar="Jane"
  instructions="You are a helpful assistant."
  tools={tools} // same shape as <OpenAIRealtimeAgent tools>
  createSession={async (offerSdp) => {
    const res = await fetch("/api/qwen-session", {
      method: "POST",
      body: offerSdp,
    });
    return res.text(); // SDP answer
  }}
/>;
// Your backend — POST /api/qwen-session
const res = await fetch(
  `https://${WORKSPACE_ID}.ap-southeast-1.maas.aliyuncs.com/api/v1/webrtc/realtime?model=qwen3.8-omni-flash-realtime`,
  {
    method: "POST",
    headers: {
      "Content-Type": "application/sdp",
      Authorization: `Bearer ${process.env.DASHSCOPE_API_KEY}`,
    },
    body: offerSdp,
  },
);
return res.text();

Use the cn-beijing host and a Beijing API key for the China (Beijing) region.

Props

PropTypeDefaultDescription
createSession(offerSdp: string) => Promise<string>requiredSends the SDP offer to your backend and resolves to Qwen's SDP answer.
instructionsstring—System prompt.
voicestring"Tina"A voice from Model Studio's voice list.
toolsOpenAIRealtimeTool[][]Browser-side function tools.
turnDetection"server_vad" | "semantic_vad""server_vad"WebRTC only supports server-side turn detection.
transcribeInputbooleanfalseAlso transcribe the user's speech (billed separately).
greetingstring | falsegreet per instructionsInstruction for the agent's opening turn; false waits for the user.

Runs full wav2arkit neural lipsync — agent audio arrives as a WebRTC media track.

The agent speaks first by default (see greeting).


ElevenLabs

Uses the ElevenLabs Conversational AI SDK. Configure your agent in the ElevenLabs dashboard and pass its ID here.

import { ElevenLabsAvatarAgent } from "gsplat-talkinghead/elevenlabs";

// Public agent (no auth required)
<ElevenLabsAvatarAgent
  agentId="your-agent-id"
/>

// Private agent (fetch a short-lived token server-side)
<ElevenLabsAvatarAgent
  backgroundImages={["/niceBG.jpg"]}
  agentId="your-agent-id"
  getConversationToken={async () => {
    const res = await fetch("/api/elevenlabs-token");
    const { token } = await res.json();
    return token;
  }}
/>

Backend — conversation token endpoint

// app/api/elevenlabs-token/route.ts
export async function GET(req: Request) {
  const agentId = new URL(req.url).searchParams.get("agentId");
  const res = await fetch(
    `https://api.elevenlabs.io/v1/convai/conversation/token?agent_id=${agentId}`,
    { headers: { "xi-api-key": process.env.ELEVENLABS_API_KEY! } },
  );
  return Response.json(await res.json());
}

Props

PropTypeDefaultDescription
agentIdstringrequiredAgent ID from the ElevenLabs dashboard.
getConversationToken() => Promise<string>—Required for private agents. Returns a short-lived WebRTC token.
clientToolsRecord<string, (...args) => any>—Client-side tools exposed to the agent.

Lipsync note: the ElevenLabs SDK only exposes frequency-magnitude and volume scalars, never raw waveform audio, so the wav2arkit neural model can't run for this provider. ElevenLabs falls back to coarse, volume-driven mouth movement instead of full lipsync.


Vapi

Uses the Vapi Web SDK. Configure your assistant in the Vapi dashboard or pass an inline configuration object.

import { VapiAvatarAgent } from "gsplat-talkinghead/vapi";

// Using a pre-configured assistant ID
<VapiAvatarAgent
  publicKey="your-vapi-public-key"
  assistantId="your-assistant-id"
/>

// Using an inline assistant config
<VapiAvatarAgent
  backgroundImages={["/niceBG.jpg"]}
  publicKey="your-vapi-public-key"
  assistant={{
    model: {
      provider: "openai",
      model: "gpt-4o-mini",
      messages: [{ role: "system", content: "You are a helpful assistant." }],
    },
    voice: { provider: "11labs", voiceId: "sarah" },
  }}
/>

Props

PropTypeDefaultDescription
publicKeystringrequiredYour Vapi public key (safe to expose in the browser).
assistantIdstring—Pre-configured assistant ID. Mutually exclusive with assistant.
assistantRecord<string, any>—Inline assistant configuration object. Mutually exclusive with assistantId.

Runs full wav2arkit neural lipsync — Vapi's Web SDK exposes a real remote MediaStream.


LiveKit

Uses LiveKit Agents over WebRTC. Your LiveKit agent must be running server-side and the token must grant access to the correct room.

import { LiveKitAvatarAgent } from "gsplat-talkinghead/livekit";

<LiveKitAvatarAgent
  backgroundImages={["/niceBG.jpg"]}
  serverUrl="wss://my-project.livekit.cloud"
  getToken={async () => {
    const res = await fetch("/api/livekit-token");
    const { token } = await res.json();
    return token;
  }}
/>;

Backend — participant token endpoint

// app/api/livekit-token/route.ts
import { AccessToken } from "livekit-server-sdk";

export async function GET() {
  const token = new AccessToken(
    process.env.LIVEKIT_API_KEY!,
    process.env.LIVEKIT_API_SECRET!,
    { identity: `user-${Date.now()}` },
  );
  token.addGrant({ roomJoin: true, room: "agent-room" });
  return Response.json({ token: await token.toJwt() });
}

Props

PropTypeDefaultDescription
serverUrlstringrequiredLiveKit server WebSocket URL, e.g. wss://my-project.livekit.cloud.
getToken() => Promise<string>requiredReturns a short-lived participant token. Generate server-side with the LiveKit server SDK.

Runs full wav2arkit neural lipsync — LiveKit's client exposes a real remote MediaStream.


Shared props

All provider components accept these additional props:

PropTypeDefaultDescription
avatarAvatarPreset"Jane"Built-in avatar: "Jack", "Jane", "John" or "Sasha". See Avatars below.
assetsPathstring—URL/path to a custom Gaussian-splat avatar asset bundle. Takes precedence over avatar.
emotionAvatarEmotion"neutral"Facial emotion over lipsync: "neutral", "happy", "sad" or "thinking". See Emotions.
backgroundImagesstring[][]Image URLs for the background. One is chosen at random each mount. Transparent when omitted.
onSessionEnd() => void—Called when the session ends (end phrase detected, timeout, or user clicked End).
endSessionPhrasestring"this is the end"Case-insensitive substring the component watches for in the agent's transcript to end the session.
sessionTimeoutnumber600000Hard timeout in milliseconds.
classNamestring—Extra CSS class names on the outermost container div.

Avatars

The library ships with four preset avatars — Jack, Jane, John and Sasha — picked with the avatar prop. Jane is used when you pass neither avatar nor assetsPath. Nothing to configure, nothing to host: the bundles are served from this project's GitHub repo (main branch) via jsDelivr's GitHub CDN. They aren't included in the npm package, which keeps the install small.

// Default preset (Jane)
<OpenAIRealtimeAgent getEphemeralKey={...} />

// Another preset
<OpenAIRealtimeAgent avatar="Jack" getEphemeralKey={...} />

The preset list is also exported, e.g. to build an avatar picker:

import { AVATAR_PRESETS, type AvatarPreset } from "gsplat-talkinghead";

AVATAR_PRESETS; // ["Jack", "Jane", "John", "Sasha"]

To serve the presets from your own origin instead of jsDelivr (strict CSP or offline use), copy the repo's gsplat-talkinghead/assets/ folder somewhere public and point the library at it once at startup:

import { configureAvatarPresets } from "gsplat-talkinghead";

configureAvatarPresets({ baseUrl: "/avatars" }); // loads /avatars/Jack.zip, …

To use your own avatar instead, host a Gaussian-splat asset bundle (compatible with @myned-ai/gsplat-flame-avatar-renderer) and pass its URL as assetsPath:

<OpenAIRealtimeAgent
  getEphemeralKey={...}
  assetsPath="https://your-cdn.example.com/avatars/your-avatar-bundle"
/>

Emotions

Every component takes an emotion prop — "neutral", "happy", "sad" or "thinking" — layered on top of lipsync as ARKit blendshape offsets (smile, brows, eye squint/gaze), so the avatar can smile while it talks. Changes blend in over ~0.3 s.

Drive it from your app:

const [emotion, setEmotion] = useState<AvatarEmotion>("neutral");
<OpenAIRealtimeAgent emotion={emotion} ... />

Or let the model set it with the built-in set_emotion tool (for providers with browser-side tools: OpenAIRealtimeAgent, OpenAILiveAgent, QwenRealtimeAgent). It returns to neutral after 8 s by default (createEmotionTool(setEmotion, { resetAfterMs })):

import { createEmotionTool, type AvatarEmotion } from "gsplat-talkinghead";

const [emotion, setEmotion] = useState<AvatarEmotion>("neutral");
const allTools = useMemo(() => [...tools, createEmotionTool(setEmotion)], []);

<OpenAIRealtimeAgent
  emotion={emotion}
  tools={allTools}
  // + in your instructions: "Call set_emotion when your feelings change:
  //   happy for good news, sad for bad news, thinking while you consider options."
/>;

The expressions are defined in EMOTION_BLENDSHAPES (exported) if you want to see or copy the exact offsets.


Advanced: adapter hooks

For full layout control, use AvatarAgent directly with an adapter hook. This lets you compose the avatar into your own UI without the built-in container styles.

import { AvatarAgent } from "gsplat-talkinghead";
import { useVapiAdapter } from "gsplat-talkinghead/vapi";

function MyPage() {
  const adapter = useVapiAdapter({
    publicKey: "your-vapi-public-key",
    assistantId: "your-assistant-id",
  });

  return (
    <div className="my-layout">
      <AvatarAgent adapter={adapter} className="h-[600px]" />
    </div>
  );
}

All adapter hooks follow the same pattern:

useOpenAIRealtimeAdapter(options); // → SessionAdapter
useElevenLabsAdapter(options); // → SessionAdapter
useVapiAdapter(options); // → SessionAdapter
useLiveKitAdapter(options); // → SessionAdapter

A SessionAdapter exposes remoteStream for full neural lipsync, plus two optional fallbacks for platforms that can't produce a MediaStream: subscribeToRemoteAudio (push already-decoded PCM straight into wav2arkit, for a platform whose SDK decodes raw PCM itself) and getRemoteAudioLevel (a 0-1 volume scalar for coarse, volume-only mouth movement — used by ElevenLabs). Implement one of these if you're writing a custom adapter for a platform without a real audio stream.


How it works

Component mounts
      │
      ├── Avatar bundle + wav2arkit model download ──► "Downloading assets" overlay, Start disabled
      │
User clicks Start
      │
      ▼
Provider adapter connects (WebRTC / WebSocket)
      │
      ├── Audio ──► resample to 16kHz ──► wav2arkit (onnxruntime-web) ──► ARKit blendshapes ─┐
      │             (MediaStream tap, or pushed raw PCM — see Advanced: adapter hooks)          │
      │                                                                                         ▼
      │   emotion prop ──► blendshape offsets ──► blink + subtle head sway ──► Gaussian-splat renderer
      │
      ├── Transcript ──► endSessionPhrase check ──► onSessionEnd(), End
      │
      ├── User clicks End ──► End
      │
      └── sessionTimeout ──► onSessionEnd(), End

End ──► face, emotion and head motion reset to rest

The wav2arkit ONNX model (myned-ai/wav2arkit_cpu, Apache-2.0) is fetched and cached in the browser on first use — no server component required.


Package structure

src/
├── index.ts                     ← public exports (AvatarAgent, presets, emotions, types)
├── openai.ts · qwen.ts · vapi.ts · elevenlabs.ts · livekit.ts   ← per-provider entry points
├── types.ts                     ← shared prop types
├── AvatarAgent.tsx              ← platform-agnostic core component
├── OpenAIRealtimeAgent.tsx      ← provider convenience wrappers
├── OpenAILiveAgent.tsx
├── QwenRealtimeAgent.tsx
├── ElevenLabsAvatarAgent.tsx
├── VapiAvatarAgent.tsx
├── LiveKitAvatarAgent.tsx
├── adapters/
│   ├── SessionAdapter.ts        ← adapter interface
│   ├── openai-realtime/ · openai-live/ · qwen/
│   └── elevenlabs/ · vapi/ · livekit/
├── avatar/
│   ├── GaussianAvatarController.ts  ← wraps the renderer: blendshapes, blink, emotion, head sway
│   ├── LazyAvatarController.ts      ← lazy-loads the renderer
│   ├── AvatarContainer.tsx          ← mount point + background image
│   ├── presets.ts                   ← built-in avatars (Jack, Jane, John, Sasha)
│   ├── emotions.ts                  ← emotion blendshapes + set_emotion tool
│   └── transparentCanvas.ts         ← transparent WebGL canvas over the background
├── session/
│   ├── useAgentSession.ts       ← OpenAI Realtime WebRTC session
│   └── codecUtils.ts
├── ui/                          ← Toolbar, StatusBadge, AssetsLoader, AudioBars
├── constants/
│   └── arkit.ts                 ← ARKit blendshape names, neutral pose
└── audio/
    ├── wav2arkit/
    │   ├── modelLoader.ts       ← fetches + caches the ONNX model
    │   ├── resample.ts          ← resamples audio to 16kHz
    │   ├── inferenceEngine.ts   ← onnxruntime-web session + load progress
    │   └── liveLipsync.ts       ← streaming/paced inference pipeline
    ├── useAvatarLipsync.ts      ← wires a MediaStream into wav2arkit
    ├── usePushAudioLipsync.ts   ← wires pushed raw PCM into wav2arkit
    ├── useVolumeFallbackLipsync.ts ← coarse fallback for stream-less adapters
    ├── delayedAudioPlayback.ts  ← delays audio to line up with lipsync
    └── useAudio.ts              ← mic monitoring

Contribution Guide

To contribute to package source code, raise PRs to the main branch.


Sponsoring ❤️

gsplat-talkinghead is free, open-source, and maintained in personal time.

If your product ships AI-powered conversations and this library saves you weeks of WebRTC wrangling, lip-sync work, and provider integration — consider sponsoring. Even a small recurring amount keeps the library maintained, more providers supported, and bugs fixed fast.

Sponsor on GitHub →


Citing this project

If you use gsplat-talkinghead in academic work, a research demo, or a published product, a citation or acknowledgement is appreciated.

BibTeX

@software{peiris2026gsplattalkinghead,
  author  = {Peiris, Navod},
  title   = {gsplat-talkinghead: LAM Gaussian-Splat Avatars for AI voice agents},
  year    = {2026},
  url     = {https://github.com/NavodPeiris/gsplat-talkinghead},
  note    = {npm: gsplat-talkinghead}
}

Plain text

Navod Peiris. gsplat-talkinghead: LAM Gaussian-Splat Avatars for AI voice agents. 2026. https://github.com/NavodPeiris/gsplat-talkinghead

Acknowledgement (for README or paper footnote)

Gaussian-splat avatar and lip-sync powered by gsplat-talkinghead.

Languages

TypeScript

90.5%

Python

3.4%

CSS

3.1%

JavaScript

2.3%