litert-community/wav2vec2-keyword-spotting

Model

Measured on device (edge-compat, w2v2_frontend_fp16): Galaxy S26 · LiteRT 2.2.0 · GPU (ML Drift) · 5.29 ms p50 (2026-08-25); Galaxy S26 · LiteRT 2.2.0 · NPU (QNN/HTP) · did not run: invoke_failed (2026-08-25); Raspberry Pi 5 · LiteRT 2.2.0.dev20260804 · CPU/XNNPACK, 4 threads · 106 ms p50 (2026-08-31); browser · Chromium 151 on M4 Max · LiteRT.js 2.5.3 · WebGPU · 3.37 ms p50 · output matches CPU (2026-08-11). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/wav2vec2-keyword-spot

1

13 commits

3 linked in READMEs

updated Sep 8, 2026

See the code

README

Measured on device (edge-compat, w2v2_frontend_fp16): Galaxy S26 · LiteRT 2.2.0 · GPU (ML Drift) · 5.29 ms p50 (2026-08-25); Galaxy S26 · LiteRT 2.2.0 · NPU (QNN/HTP) · did not run: invoke_failed (2026-08-25); Raspberry Pi 5 · LiteRT 2.2.0.dev20260804 · CPU/XNNPACK, 4 threads · 106 ms p50 (2026-08-31); browser · Chromium 151 on M4 Max · LiteRT.js 2.5.3 · WebGPU · 3.37 ms p50 · output matches CPU (2026-08-11). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/wav2vec2-keyword-spotting__w2v2_frontend_fp16/CARD.md

Measured on device (edge-compat, w2v2_head_fp16): Galaxy S26 · LiteRT 2.2.0 · GPU (ML Drift) · 9.20 ms p50 (2026-08-26); Galaxy S26 · LiteRT 2.2.0 · NPU (QNN/HTP) · 5.67 ms p50 (2026-08-26); Raspberry Pi 5 · LiteRT 2.2.0.dev20260804 · CPU/XNNPACK, 4 threads · 106 ms p50 (2026-08-31); browser · Chromium 151 on M4 Max · LiteRT.js 2.5.3 · WebGPU · 12.2 ms p50 · output matches CPU (2026-08-11). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/wav2vec2-keyword-spotting__w2v2_head_fp16/CARD.md

wav2vec2 Keyword Spotting — LiteRT on-device (all-GPU)

On-device LiteRT conversion of superb/wav2vec2-base-superb-ks (Apache-2.0) — wav2vec2-base keyword spotting, 12 Speech-Commands labels (yes / no / up / down / left / right / on / off / stop / go / unknown / silence). Runs fully on the CompiledModel GPU delegate (LITERT_CL); there is no FFT anywhere (the raw 16 kHz waveform goes straight into a 1D-conv feature extractor — no mel step), and the transformer residual is small enough that the whole model is fp16-exact on GPU (no CPU fallback). Device-verified on a Pixel 8a (Tensor G3).

wav2vec2 KWS — keyword waveform and 12-label probabilities (on-device LiteRT GPU)

Files

FileSizeDelegateIn → Out
w2v2_frontend_fp16.tflite9 MBGPUaudio [1,16000] → feat [1,49,768]
w2v2_head_fp16.tflite181 MBGPUfeat [1,49,768] → logits [1,12]

End-to-end ~19 ms for a 1 s clip (RTF ≈ 0.02); 10/10 keywords correct on real speech, device-vs-CPU logits corr 0.9995.

Two graphs

The model is op-clean for the GPU but the full 1008-node graph exceeds the Mali shader-compile limit (fails to compile fused). Splitting at the conv-frontend / transformer-encoder boundary makes each half compile (frontend 134/134 + head 893/893 LITERT_CL). The frontend output feeds the head:

waveform →[GPU frontend]→ feat[1,49,768] →[GPU head]→ logits[1,12] → argmax
  • frontend = feature_extractor (7 strided 1D convs + GroupNorm) + feature_projection.
  • head = encoder (12 transformer layers + conv positional embedding) + weighted-layer-sum over all 13 hidden states (use_weighted_layer_sum) + projector + mean-pool + classifier.

Minimal usage

Android (Kotlin, CompiledModel GPU)

// staged into filesDir by the sample's install_to_device.sh (the head is 181 MB)
val opts = CompiledModel.Options(Accelerator.GPU)
val fe = CompiledModel.create(File(ctx.filesDir, "w2v2_frontend_fp16.tflite").absolutePath, opts, null)
val head = CompiledModel.create(File(ctx.filesDir, "w2v2_head_fp16.tflite").absolutePath, opts, null)
val feIn = fe.createInputBuffers(); val feOut = fe.createOutputBuffers()
val hdIn = head.createInputBuffers(); val hdOut = head.createOutputBuffers()
feIn[0].writeFloat(audio)                 // [1,16000] 1 s @ 16 kHz, raw [-1,1]
fe.run(feIn, feOut)
hdIn[0].writeFloat(feOut[0].readFloat())  // feat [1,49,768]
head.run(hdIn, hdOut)
val logits = hdOut[0].readFloat()         // [12] -> argmax = keyword

Python (desktop verification)

import numpy as np, soundfile as sf
from ai_edge_litert.interpreter import Interpreter

LABELS = ["yes", "no", "up", "down", "left", "right",
          "on", "off", "stop", "go", "_unknown_", "_silence_"]

wav, _ = sf.read("clip_16k.wav", dtype="float32")            # mono 16 kHz
x = np.zeros((1, 16000), np.float32); n = min(len(wav), 16000); x[0, :n] = wav[:n]
# raw [-1,1] waveform straight in — this checkpoint uses do_normalize=False

fe = Interpreter(model_path="w2v2_frontend_fp16.tflite"); fe.allocate_tensors()
fe.set_tensor(fe.get_input_details()[0]["index"], x); fe.invoke()
feat = fe.get_tensor(fe.get_output_details()[0]["index"])    # [1,49,768]

hd = Interpreter(model_path="w2v2_head_fp16.tflite"); hd.allocate_tensors()
hd.set_tensor(hd.get_input_details()[0]["index"], feat); hd.invoke()
logits = hd.get_tensor(hd.get_output_details()[0]["index"])[0]  # [12]
print(LABELS[int(logits.argmax())])

Re-authoring (litert-torch, parity corr 1.0)

GELU→tanh-GELU · feature-extractor GroupNorm→GN4D · pos-conv weight_norm fold · create_bidirectional_mask→None · weighted-layer-sum accumulated incrementally with baked softmax(layer_weights) constants (the runtime softmax + per-layer scalar gathers otherwise split the Mali partition).

Sample app

A complete Android sample app + the conversion scripts are in the official LiteRT samples repository under compiled_model_api/audio_classification (google-ai-edge/litert-samples). Push these files to the app's filesDir with that sample's install_to_device.sh.

License follows upstream (Apache-2.0).

Performance

Measured on a Pixel 8a (Tensor G3, Android 16) with the standard TFLite benchmark_model tool — 10 warm-up runs then 50 timed runs, reported as the tool's mean.

RuntimeBackendGraph on GPULatency
TFLite benchmark_model (TfLiteGpuDelegateV2) — w2v2_frontend_fp16.tfliteGPU (OpenCL)134 / 13438.4 ms
TFLite benchmark_model (TfLiteGpuDelegateV2) — w2v2_head_fp16.tfliteGPU (OpenCL)52 / 893515.7 ms
TFLite benchmark_model — w2v2_frontend_fp16.tfliteCPU (XNNPACK, 4 threads)—XNNPACK declined the graph
TFLite benchmark_model — w2v2_head_fp16.tfliteCPU (XNNPACK, 4 threads)—123.3 ms

Any on-device figure recorded when this model shipped came from a different runtime. It was taken through LiteRT's own CompiledModel accelerator (logcat reports it as LITERT_CL), which is the path the Kotlin sample app and the LiteRT API use, and it appears elsewhere on this card. The rows above are the classic TFLite OpenCL delegate, measured with a tool anyone can download and re-run. The two are not comparable, so read the rows above as a reproducible floor rather than as this model's speed on LiteRT.

XNNPACK declines these fp16 graphs — it reports failed to delegate DEPTHWISE_CONV_2D and then fails to allocate tensors — so there is no usable CPU number. Disabling XNNPACK falls back to reference kernels, which measured about 20× slower than the GPU on models of this size and would not represent CPU inference anyone would ship.

On this delegate the CPU is the faster choice for w2v2_head_fp16.tflite (123.3 ms on CPU against 515.7 ms on GPU) — worth knowing before you reach for the GPU on a mid-range phone.

Note that the GPU does not take the whole graph here (52 / 893 in w2v2_head_fp16.tflite); the remainder runs on the CPU and the split costs a per-partition round trip.

Snapdragon NPU (Hexagon)

  • w2v2_frontend_fp16.tflite — the GPU runs it at 5.29 ms. The NPU does not — the graph compiles and then fails to run (LiteRtException: Failed to invoke the compiled model).
  • w2v2_head_fp16.tflite — the NPU is 1.62x faster than the GPU (5.67 ms against 9.20 ms) and loads 7.06x faster (179 ms against 1265 ms).
filebackendcompiledinference (median / min)load
w2v2_frontend_fp16.tfliteGPU (Adreno)—5.29 ms / 5.05 ms656 ms
w2v2_head_fp16.tfliteNPU (Hexagon v81)on-device JIT5.67 ms / 5.61 ms179 ms
w2v2_head_fp16.tfliteGPU (Adreno)—9.20 ms / 8.89 ms1265 ms

Measured on a Samsung Galaxy S26 (Snapdragon 8 Elite Gen 5 / SM8850, Hexagon v81, Android 16) with LiteRT CompiledModel 2.2.0, one accelerator per process, 5 warm-up runs then N=50 timed runs, median reported. Every run held thermal status NONE throughout. Headroom 0.61–0.80, where 1.0 is the throttling threshold.

The NPU rows ran the published file unchanged. LiteRT compiled it for the Hexagon on the device at first load. That first compile took 7.5 s here. The load column above is the cached load every later run pays. Recipe and the runtime libraries it needs: NPU guide.

GPU wiring: GPU guide.

Raspberry Pi 5 (CPU)

Measured on a Raspberry Pi 5 Model B Rev 1.1 (8 GB, Raspberry Pi OS 64-bit) with the LiteRT benchmark_model tool from litert-cli-nightly 0.2.0.dev20260805: CPU inference (XNNPACK, 4 threads), 3 invocations per file of 10 warm-up plus 50 timed runs (the tool caps a phase at 150 s, so very slow graphs run fewer — the Runs column is the actual timed total). The latency is the median across invocations; the spread is the min–max over all timed runs. No thermal throttling occurred during these runs (vcgencmd get_throttled stayed 0x0).

FileInference (median)Spread (min–max)RunsPeak memory
w2v2_frontend_fp16.tflite106.2 ms105.7–109.2 ms150138 MB
w2v2_head_fp16.tflite106.3 ms104.6–136.8 ms150600 MB
audio-classification
keyword-spotting
litert
on-device
tflite

litert-community/wav2vec2-keyword-spotting

Model

Measured on device (edge-compat, w2v2_frontend_fp16): Galaxy S26 · LiteRT 2.2.0 · GPU (ML Drift) · 5.29 ms p50 (2026-08-25); Galaxy S26 · LiteRT 2.2.0 · NPU (QNN/HTP) · did not run: invoke_failed (2026-08-25); Raspberry Pi 5 · LiteRT 2.2.0.dev20260804 · CPU/XNNPACK, 4 threads · 106 ms p50 (2026-08-31); browser · Chromium 151 on M4 Max · LiteRT.js 2.5.3 · WebGPU · 3.37 ms p50 · output matches CPU (2026-08-11). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/wav2vec2-keyword-spot

1

13 commits

3 linked in READMEs

updated Sep 8, 2026

See the code

README

Measured on device (edge-compat, w2v2_frontend_fp16): Galaxy S26 · LiteRT 2.2.0 · GPU (ML Drift) · 5.29 ms p50 (2026-08-25); Galaxy S26 · LiteRT 2.2.0 · NPU (QNN/HTP) · did not run: invoke_failed (2026-08-25); Raspberry Pi 5 · LiteRT 2.2.0.dev20260804 · CPU/XNNPACK, 4 threads · 106 ms p50 (2026-08-31); browser · Chromium 151 on M4 Max · LiteRT.js 2.5.3 · WebGPU · 3.37 ms p50 · output matches CPU (2026-08-11). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/wav2vec2-keyword-spotting__w2v2_frontend_fp16/CARD.md

Measured on device (edge-compat, w2v2_head_fp16): Galaxy S26 · LiteRT 2.2.0 · GPU (ML Drift) · 9.20 ms p50 (2026-08-26); Galaxy S26 · LiteRT 2.2.0 · NPU (QNN/HTP) · 5.67 ms p50 (2026-08-26); Raspberry Pi 5 · LiteRT 2.2.0.dev20260804 · CPU/XNNPACK, 4 threads · 106 ms p50 (2026-08-31); browser · Chromium 151 on M4 Max · LiteRT.js 2.5.3 · WebGPU · 12.2 ms p50 · output matches CPU (2026-08-11). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/wav2vec2-keyword-spotting__w2v2_head_fp16/CARD.md

wav2vec2 Keyword Spotting — LiteRT on-device (all-GPU)

On-device LiteRT conversion of superb/wav2vec2-base-superb-ks (Apache-2.0) — wav2vec2-base keyword spotting, 12 Speech-Commands labels (yes / no / up / down / left / right / on / off / stop / go / unknown / silence). Runs fully on the CompiledModel GPU delegate (LITERT_CL); there is no FFT anywhere (the raw 16 kHz waveform goes straight into a 1D-conv feature extractor — no mel step), and the transformer residual is small enough that the whole model is fp16-exact on GPU (no CPU fallback). Device-verified on a Pixel 8a (Tensor G3).

wav2vec2 KWS — keyword waveform and 12-label probabilities (on-device LiteRT GPU)

Files

FileSizeDelegateIn → Out
w2v2_frontend_fp16.tflite9 MBGPUaudio [1,16000] → feat [1,49,768]
w2v2_head_fp16.tflite181 MBGPUfeat [1,49,768] → logits [1,12]

End-to-end ~19 ms for a 1 s clip (RTF ≈ 0.02); 10/10 keywords correct on real speech, device-vs-CPU logits corr 0.9995.

Two graphs

The model is op-clean for the GPU but the full 1008-node graph exceeds the Mali shader-compile limit (fails to compile fused). Splitting at the conv-frontend / transformer-encoder boundary makes each half compile (frontend 134/134 + head 893/893 LITERT_CL). The frontend output feeds the head:

waveform →[GPU frontend]→ feat[1,49,768] →[GPU head]→ logits[1,12] → argmax
  • frontend = feature_extractor (7 strided 1D convs + GroupNorm) + feature_projection.
  • head = encoder (12 transformer layers + conv positional embedding) + weighted-layer-sum over all 13 hidden states (use_weighted_layer_sum) + projector + mean-pool + classifier.

Minimal usage

Android (Kotlin, CompiledModel GPU)

// staged into filesDir by the sample's install_to_device.sh (the head is 181 MB)
val opts = CompiledModel.Options(Accelerator.GPU)
val fe = CompiledModel.create(File(ctx.filesDir, "w2v2_frontend_fp16.tflite").absolutePath, opts, null)
val head = CompiledModel.create(File(ctx.filesDir, "w2v2_head_fp16.tflite").absolutePath, opts, null)
val feIn = fe.createInputBuffers(); val feOut = fe.createOutputBuffers()
val hdIn = head.createInputBuffers(); val hdOut = head.createOutputBuffers()
feIn[0].writeFloat(audio)                 // [1,16000] 1 s @ 16 kHz, raw [-1,1]
fe.run(feIn, feOut)
hdIn[0].writeFloat(feOut[0].readFloat())  // feat [1,49,768]
head.run(hdIn, hdOut)
val logits = hdOut[0].readFloat()         // [12] -> argmax = keyword

Python (desktop verification)

import numpy as np, soundfile as sf
from ai_edge_litert.interpreter import Interpreter

LABELS = ["yes", "no", "up", "down", "left", "right",
          "on", "off", "stop", "go", "_unknown_", "_silence_"]

wav, _ = sf.read("clip_16k.wav", dtype="float32")            # mono 16 kHz
x = np.zeros((1, 16000), np.float32); n = min(len(wav), 16000); x[0, :n] = wav[:n]
# raw [-1,1] waveform straight in — this checkpoint uses do_normalize=False

fe = Interpreter(model_path="w2v2_frontend_fp16.tflite"); fe.allocate_tensors()
fe.set_tensor(fe.get_input_details()[0]["index"], x); fe.invoke()
feat = fe.get_tensor(fe.get_output_details()[0]["index"])    # [1,49,768]

hd = Interpreter(model_path="w2v2_head_fp16.tflite"); hd.allocate_tensors()
hd.set_tensor(hd.get_input_details()[0]["index"], feat); hd.invoke()
logits = hd.get_tensor(hd.get_output_details()[0]["index"])[0]  # [12]
print(LABELS[int(logits.argmax())])

Re-authoring (litert-torch, parity corr 1.0)

GELU→tanh-GELU · feature-extractor GroupNorm→GN4D · pos-conv weight_norm fold · create_bidirectional_mask→None · weighted-layer-sum accumulated incrementally with baked softmax(layer_weights) constants (the runtime softmax + per-layer scalar gathers otherwise split the Mali partition).

Sample app

A complete Android sample app + the conversion scripts are in the official LiteRT samples repository under compiled_model_api/audio_classification (google-ai-edge/litert-samples). Push these files to the app's filesDir with that sample's install_to_device.sh.

License follows upstream (Apache-2.0).

Performance

Measured on a Pixel 8a (Tensor G3, Android 16) with the standard TFLite benchmark_model tool — 10 warm-up runs then 50 timed runs, reported as the tool's mean.

RuntimeBackendGraph on GPULatency
TFLite benchmark_model (TfLiteGpuDelegateV2) — w2v2_frontend_fp16.tfliteGPU (OpenCL)134 / 13438.4 ms
TFLite benchmark_model (TfLiteGpuDelegateV2) — w2v2_head_fp16.tfliteGPU (OpenCL)52 / 893515.7 ms
TFLite benchmark_model — w2v2_frontend_fp16.tfliteCPU (XNNPACK, 4 threads)—XNNPACK declined the graph
TFLite benchmark_model — w2v2_head_fp16.tfliteCPU (XNNPACK, 4 threads)—123.3 ms

Any on-device figure recorded when this model shipped came from a different runtime. It was taken through LiteRT's own CompiledModel accelerator (logcat reports it as LITERT_CL), which is the path the Kotlin sample app and the LiteRT API use, and it appears elsewhere on this card. The rows above are the classic TFLite OpenCL delegate, measured with a tool anyone can download and re-run. The two are not comparable, so read the rows above as a reproducible floor rather than as this model's speed on LiteRT.

XNNPACK declines these fp16 graphs — it reports failed to delegate DEPTHWISE_CONV_2D and then fails to allocate tensors — so there is no usable CPU number. Disabling XNNPACK falls back to reference kernels, which measured about 20× slower than the GPU on models of this size and would not represent CPU inference anyone would ship.

On this delegate the CPU is the faster choice for w2v2_head_fp16.tflite (123.3 ms on CPU against 515.7 ms on GPU) — worth knowing before you reach for the GPU on a mid-range phone.

Note that the GPU does not take the whole graph here (52 / 893 in w2v2_head_fp16.tflite); the remainder runs on the CPU and the split costs a per-partition round trip.

Snapdragon NPU (Hexagon)

  • w2v2_frontend_fp16.tflite — the GPU runs it at 5.29 ms. The NPU does not — the graph compiles and then fails to run (LiteRtException: Failed to invoke the compiled model).
  • w2v2_head_fp16.tflite — the NPU is 1.62x faster than the GPU (5.67 ms against 9.20 ms) and loads 7.06x faster (179 ms against 1265 ms).
filebackendcompiledinference (median / min)load
w2v2_frontend_fp16.tfliteGPU (Adreno)—5.29 ms / 5.05 ms656 ms
w2v2_head_fp16.tfliteNPU (Hexagon v81)on-device JIT5.67 ms / 5.61 ms179 ms
w2v2_head_fp16.tfliteGPU (Adreno)—9.20 ms / 8.89 ms1265 ms

Measured on a Samsung Galaxy S26 (Snapdragon 8 Elite Gen 5 / SM8850, Hexagon v81, Android 16) with LiteRT CompiledModel 2.2.0, one accelerator per process, 5 warm-up runs then N=50 timed runs, median reported. Every run held thermal status NONE throughout. Headroom 0.61–0.80, where 1.0 is the throttling threshold.

The NPU rows ran the published file unchanged. LiteRT compiled it for the Hexagon on the device at first load. That first compile took 7.5 s here. The load column above is the cached load every later run pays. Recipe and the runtime libraries it needs: NPU guide.

GPU wiring: GPU guide.

Raspberry Pi 5 (CPU)

Measured on a Raspberry Pi 5 Model B Rev 1.1 (8 GB, Raspberry Pi OS 64-bit) with the LiteRT benchmark_model tool from litert-cli-nightly 0.2.0.dev20260805: CPU inference (XNNPACK, 4 threads), 3 invocations per file of 10 warm-up plus 50 timed runs (the tool caps a phase at 150 s, so very slow graphs run fewer — the Runs column is the actual timed total). The latency is the median across invocations; the spread is the min–max over all timed runs. No thermal throttling occurred during these runs (vcgencmd get_throttled stayed 0x0).

FileInference (median)Spread (min–max)RunsPeak memory
w2v2_frontend_fp16.tflite106.2 ms105.7–109.2 ms150138 MB
w2v2_head_fp16.tflite106.3 ms104.6–136.8 ms150600 MB
audio-classification
keyword-spotting
litert
on-device
tflite