myned-ai/wav2arkit_cpu

Model

Wav2ARKit - Audio to Facial Expression (ONNX)

9

5 commits

1 linked in READMEs

updated Feb 22, 2026

See the code

README

Wav2ARKit - Audio to Facial Expression (ONNX)

A fused, end-to-end ONNX model that converts raw audio waveforms directly into 52 ARKit-compatible facial blendshapes. Based on the Facebook Wav2Vec2 and LAM Audio2Expression models, optimized for real-time CPU inference.

Features

FeatureValue
InputRaw 16kHz audio waveform
Output52 ARKit blendshapes @ 30fps
Inference~45ms per second of audio
Speed22× faster than realtime
Size1.8 MB

Quick Start

import onnxruntime as ort
import numpy as np

# Load model
session = ort.InferenceSession("wav2arkit_cpu.onnx", providers=["CPUExecutionProvider"])

# Load audio (16kHz, mono, float32)
# Example: 1 second = 16000 samples
audio = np.random.randn(1, 16000).astype(np.float32)

# Output: (1, 30, 52) - 30 frames at 30fps, 52 blendshapes

Model Specification

Input

NameTypeShapeDescription
audio_waveformfloat32[batch, samples]Raw audio at 16kHz

Output

NameTypeShapeDescription
blendshapesfloat32[batch, frames, 52]ARKit blendshapes [0-1]

Frame Calculation

output_frames = ceil(30 × (num_samples / 16000))

Example: 1 second audio (16000 samples) → 30 frames

ARKit Blendshapes

52 blendshape indices (click to expand)
IdxNameIdxName
0browDownLeft26mouthClose
1browDownRight27mouthDimpleLeft
2browInnerUp28mouthDimpleRight
3browOuterUpLeft29mouthFrownLeft
4browOuterUpRight30mouthFrownRight
5cheekPuff31mouthFunnel
6cheekSquintLeft32mouthLeft
7cheekSquintRight33mouthLowerDownLeft
8eyeBlinkLeft34mouthLowerDownRight
9eyeBlinkRight35mouthPressLeft
10eyeLookDownLeft36mouthPressRight
11eyeLookDownRight37mouthPucker
12eyeLookInLeft38mouthRight
13eyeLookInRight39mouthRollLower
14eyeLookOutLeft40mouthRollUpper
15eyeLookOutRight41mouthShrugLower
16eyeLookUpLeft42mouthShrugUpper
17eyeLookUpRight43mouthSmileLeft
18eyeSquintLeft44mouthSmileRight
19eyeSquintRight45mouthStretchLeft
20eyeWideLeft46mouthStretchRight
21eyeWideRight47mouthUpperUpLeft
22jawForward48mouthUpperUpRight
23jawLeft49noseSneerLeft
24jawOpen50noseSneerRight
25jawRight51tongueOut

Usage Examples

Python with audio file

import onnxruntime as ort
import numpy as np
import soundfile as sf

session = ort.InferenceSession("wav2arkit_cpu.onnx", providers=["CPUExecutionProvider"])

# Load and resample audio to 16kHz if needed
audio, sr = sf.read("speech.wav")
if sr != 16000:
    import librosa
    audio = librosa.resample(audio, orig_sr=sr, target_sr=16000)

# Ensure mono
if len(audio.shape) > 1:
    audio = audio.mean(axis=1)

# Run inference
audio_input = audio.astype(np.float32).reshape(1, -1)
blendshapes = session.run(None, {"audio_waveform": audio_input})[0]

print(f"Duration: {len(audio)/16000:.2f}s → {blendshapes.shape[1]} frames")

C++

#include <onnxruntime_cxx_api.h>

Ort::Env env(ORT_LOGGING_LEVEL_WARNING, "Wav2ARKit");
Ort::Session session(env, L"wav2arkit_cpu.onnx", Ort::SessionOptions{});

std::vector<float> audio(16000);  // 1 second
std::vector<int64_t> shape = {1, 16000};

Ort::MemoryInfo mem = Ort::MemoryInfo::CreateCpu(OrtArenaAllocator, OrtMemTypeDefault);
Ort::Value input = Ort::Value::CreateTensor<float>(mem, audio.data(), audio.size(), shape.data(), shape.size());

const char* input_names[] = {"audio_waveform"};
const char* output_names[] = {"blendshapes"};
auto output = session.Run({}, input_names, &input, 1, output_names, 1);

JavaScript (onnxruntime-web/node)

const ort = require('onnxruntime-node');

const session = await ort.InferenceSession.create('wav2arkit_cpu.onnx');
const audioTensor = new ort.Tensor('float32', audioData, [1, audioData.length]);
const { blendshapes } = await session.run({ audio_waveform: audioTensor });

Architecture

Model Architecture

Note: The identity encoder supports 12 speaker identities (0-11). This ONNX export uses identity 11 baked in for single-speaker inference.

License

Apache 2.0 - Based on:

arkit
audio2expression
avatar
blendshapes
facial-animation
onnx
onnxruntime
realtime
wav2vec2

myned-ai/wav2arkit_cpu

Model

Wav2ARKit - Audio to Facial Expression (ONNX)

9

5 commits

1 linked in READMEs

updated Feb 22, 2026

See the code

README

Wav2ARKit - Audio to Facial Expression (ONNX)

A fused, end-to-end ONNX model that converts raw audio waveforms directly into 52 ARKit-compatible facial blendshapes. Based on the Facebook Wav2Vec2 and LAM Audio2Expression models, optimized for real-time CPU inference.

Features

FeatureValue
InputRaw 16kHz audio waveform
Output52 ARKit blendshapes @ 30fps
Inference~45ms per second of audio
Speed22× faster than realtime
Size1.8 MB

Quick Start

import onnxruntime as ort
import numpy as np

# Load model
session = ort.InferenceSession("wav2arkit_cpu.onnx", providers=["CPUExecutionProvider"])

# Load audio (16kHz, mono, float32)
# Example: 1 second = 16000 samples
audio = np.random.randn(1, 16000).astype(np.float32)

# Output: (1, 30, 52) - 30 frames at 30fps, 52 blendshapes

Model Specification

Input

NameTypeShapeDescription
audio_waveformfloat32[batch, samples]Raw audio at 16kHz

Output

NameTypeShapeDescription
blendshapesfloat32[batch, frames, 52]ARKit blendshapes [0-1]

Frame Calculation

output_frames = ceil(30 × (num_samples / 16000))

Example: 1 second audio (16000 samples) → 30 frames

ARKit Blendshapes

52 blendshape indices (click to expand)
IdxNameIdxName
0browDownLeft26mouthClose
1browDownRight27mouthDimpleLeft
2browInnerUp28mouthDimpleRight
3browOuterUpLeft29mouthFrownLeft
4browOuterUpRight30mouthFrownRight
5cheekPuff31mouthFunnel
6cheekSquintLeft32mouthLeft
7cheekSquintRight33mouthLowerDownLeft
8eyeBlinkLeft34mouthLowerDownRight
9eyeBlinkRight35mouthPressLeft
10eyeLookDownLeft36mouthPressRight
11eyeLookDownRight37mouthPucker
12eyeLookInLeft38mouthRight
13eyeLookInRight39mouthRollLower
14eyeLookOutLeft40mouthRollUpper
15eyeLookOutRight41mouthShrugLower
16eyeLookUpLeft42mouthShrugUpper
17eyeLookUpRight43mouthSmileLeft
18eyeSquintLeft44mouthSmileRight
19eyeSquintRight45mouthStretchLeft
20eyeWideLeft46mouthStretchRight
21eyeWideRight47mouthUpperUpLeft
22jawForward48mouthUpperUpRight
23jawLeft49noseSneerLeft
24jawOpen50noseSneerRight
25jawRight51tongueOut

Usage Examples

Python with audio file

import onnxruntime as ort
import numpy as np
import soundfile as sf

session = ort.InferenceSession("wav2arkit_cpu.onnx", providers=["CPUExecutionProvider"])

# Load and resample audio to 16kHz if needed
audio, sr = sf.read("speech.wav")
if sr != 16000:
    import librosa
    audio = librosa.resample(audio, orig_sr=sr, target_sr=16000)

# Ensure mono
if len(audio.shape) > 1:
    audio = audio.mean(axis=1)

# Run inference
audio_input = audio.astype(np.float32).reshape(1, -1)
blendshapes = session.run(None, {"audio_waveform": audio_input})[0]

print(f"Duration: {len(audio)/16000:.2f}s → {blendshapes.shape[1]} frames")

C++

#include <onnxruntime_cxx_api.h>

Ort::Env env(ORT_LOGGING_LEVEL_WARNING, "Wav2ARKit");
Ort::Session session(env, L"wav2arkit_cpu.onnx", Ort::SessionOptions{});

std::vector<float> audio(16000);  // 1 second
std::vector<int64_t> shape = {1, 16000};

Ort::MemoryInfo mem = Ort::MemoryInfo::CreateCpu(OrtArenaAllocator, OrtMemTypeDefault);
Ort::Value input = Ort::Value::CreateTensor<float>(mem, audio.data(), audio.size(), shape.data(), shape.size());

const char* input_names[] = {"audio_waveform"};
const char* output_names[] = {"blendshapes"};
auto output = session.Run({}, input_names, &input, 1, output_names, 1);

JavaScript (onnxruntime-web/node)

const ort = require('onnxruntime-node');

const session = await ort.InferenceSession.create('wav2arkit_cpu.onnx');
const audioTensor = new ort.Tensor('float32', audioData, [1, audioData.length]);
const { blendshapes } = await session.run({ audio_waveform: audioTensor });

Architecture

Model Architecture

Note: The identity encoder supports 12 speaker identities (0-11). This ONNX export uses identity 11 baked in for single-speaker inference.

License

Apache 2.0 - Based on:

arkit
audio2expression
avatar
blendshapes
facial-animation
onnx
onnxruntime
realtime
wav2vec2