litert-community/RWKV-7-World-0.1B-LiteRT

Model

Measured on device (edge-compat): browser · Chromium 151 on M4 Max · LiteRT.js 2.5.3 · WebGPU · 15.1 ms p50 · output matches CPU (2026-08-11). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/rwkv-7-world-0.1b/CARD.md

1

12 commits

4 linked in READMEs

updated Sep 8, 2026

See the code

README

Measured on device (edge-compat): browser · Chromium 151 on M4 Max · LiteRT.js 2.5.3 · WebGPU · 15.1 ms p50 · output matches CPU (2026-08-11). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/rwkv-7-world-0.1b/CARD.md

RWKV-7 World 0.1B — on-device text generation (LiteRT GPU)

The first autoregressive language model running its full forward pass on the LiteRT CompiledModel GPU delegate (RNN mode, host-side state; no CPU fallback for any op). RWKV-7 is an RNN: generation feeds one token per step and carries a small fixed-size recurrent state, so the whole model fits a single static GPU graph — no KV cache growth, no dynamic shapes.

  • Architecture: RWKV-x070-World-0.1B-v2.8 — 12 layers, d=768, 12 heads, vocab 65536.
  • Weights: BlinkDL/rwkv-7-world · Apache-2.0.
  • Size: 282 MB (fp16) + 100 MB host-side embedding table.

RWKV-7 on-device generation

Greedy generation on a Pixel 8a; the full per-token forward runs on the GPU.

I/O (per-token step graph)

TensorShapeRole
x_emb (in)[1, 768]embedding row of the current token (host lookup)
att_shift (in/out)[12, 768]per-layer attention token-shift state
ffn_shift (in/out)[12, 768]per-layer FFN token-shift state
wkv (in/out)[144, 64, 64]per-layer-per-head wkv state (12×12 heads)
logits (out)[1, 65536]next-token logits

Host side per step: look the token's row up in the fp16 embedding table (rwkv7_emb_fp16.bin, GATHER is not GPU-compatible; the first LayerNorm is inside the graph), run the graph, argmax the logits, and feed the three output states back in. Prefill = the same loop over the prompt tokens. Tokenizer: RWKV World greedy longest-match trie (rwkv_vocab_v20230424.txt).

GPU conversion

Fully GPU-resident on a Pixel 8a (1863/1863 nodes, 1 partition, ~18 ms/token fp16) via exact re-authorings, no approximation:

  • wkv7 recurrence at T=1 → plain 4D matmul/elementwise ops.
  • GroupNorm(heads) → manual per-head mean/var.
  • F.normalize → x * rsqrt(sum(x²) + eps).
  • softplus → branch-free relu(z) + log1p(exp(-|z|)) (the stock lowering emits GREATER+SELECT, rejected by the GPU delegate).
  • Token embedding lookup host-side.

Verified: sequential step-mode == parallel GPT-mode logits (corr 1.0000000); desktop fp16 CompiledModel corr 1.0000000 vs fp32 PyTorch; on-device 30-token greedy generation tracks desktop fp32 (28/30 tokens identical; the two divergences are fp32 near-ties with logit gap ≤ 0.04).

Minimal usage

Kotlin (Android, LiteRT CompiledModel GPU)

val model = CompiledModel.create(modelPath, CompiledModel.Options(Accelerator.GPU), null)
val inputs = model.createInputBuffers()
val outputs = model.createOutputBuffers()

var att = FloatArray(12 * 768); var ffn = FloatArray(12 * 768)
var wkv = FloatArray(144 * 64 * 64)
for (token in promptIds + generated) {
    inputs[0].writeFloat(embeddingRow(token))   // host fp16-table lookup
    inputs[1].writeFloat(att); inputs[2].writeFloat(ffn); inputs[3].writeFloat(wkv)
    model.run(inputs, outputs)
    val logits = outputs[0].readFloat()          // [65536] -> argmax = next token
    att = outputs[1].readFloat(); ffn = outputs[2].readFloat(); wkv = outputs[3].readFloat()
}

Python (LiteRT CompiledModel API)

import numpy as np
from ai_edge_litert.compiled_model import CompiledModel

model = CompiledModel.from_file("rwkv7_step_fp16.tflite")
inputs = model.create_input_buffers(0)
outputs = model.create_output_buffers(0)

emb = np.fromfile("rwkv7_emb_fp16.bin", "<f2").reshape(65536, 768)
att = np.zeros((12, 768), np.float32)
ffn = np.zeros((12, 768), np.float32)
wkv = np.zeros((144, 64, 64), np.float32)
for token in prompt_ids:
    inputs[0].write(emb[token : token + 1].astype(np.float32))
    inputs[1].write(att.ravel()); inputs[2].write(ffn.ravel()); inputs[3].write(wkv.ravel())
    model.run_by_index(0, inputs, outputs)
    logits = outputs[0].read(65536, np.float32)          # argmax -> next token
    att = outputs[1].read(12 * 768, np.float32)
    ffn = outputs[2].read(12 * 768, np.float32)
    wkv = outputs[3].read(144 * 64 * 64, np.float32)

Files

FileSizeRole
rwkv7_step_fp16.tflite282 MBper-token step graph (fp16 weights)
rwkv7_emb_fp16.bin100 MBembedding table [65536, 768] little-endian fp16, for host lookup
rwkv_vocab_v20230424.txt1.1 MBRWKV World vocabulary

Performance

Measured on a Pixel 8a (Tensor G3, Android 16) with the standard TFLite benchmark_model tool — 10 warm-up runs then 50 timed runs, reported as the tool's mean.

RuntimeBackendGraph on GPULatency
LiteRT CompiledModel (LITERT_CL)GPU1863 / 1863~18 ms
TFLite benchmark_model (TfLiteGpuDelegateV2)GPU (OpenCL)113 / 1863123.1 ms
TFLite benchmark_modelCPU (XNNPACK, 4 threads)—40.2 ms

The two GPU rows are different runtimes, not a contradiction. The LITERT_CL figure is the one recorded when this model shipped, taken through LiteRT's own CompiledModel accelerator — the path the Kotlin sample app and the LiteRT API use. The TfLiteGpuDelegateV2 figure is the classic TFLite OpenCL delegate, measured with a tool anyone can download and re-run. They agree on how much of the graph the GPU takes; they disagree on speed, and the classic delegate is the slower of the two here. Read the TfLiteGpuDelegateV2 row as a reproducible floor, not as this model's speed on LiteRT.

On this delegate the CPU is the faster choice for (40.2 ms on CPU against 123.1 ms on GPU) — worth knowing before you reach for the GPU on a mid-range phone.

Note that the GPU does not take the whole graph here (113 / 1863); the remainder runs on the CPU and the split costs a per-partition round trip.

Snapdragon NPU (Hexagon)

The NPU is 1.70x faster than the GPU (5.37 ms against 9.11 ms) and loads 8.28x faster (270 ms against 2236 ms).

backendinference (median / min)load
NPU (Hexagon v81)5.37 ms / 5.33 ms270 ms
GPU (Adreno)9.11 ms / 8.37 ms2236 ms

Measured on a Samsung Galaxy S26 (Snapdragon 8 Elite Gen 5 / SM8850, Hexagon v81, Android 16), LiteRT CompiledModel 2.2.0, one accelerator per process, 5 warm-up runs then N=50 timed runs, median reported. Every run held thermal status NONE throughout. Headroom 0.77-0.78, where 1.0 is the throttling threshold.

The NPU rows here ran artifacts compiled ahead of time for SM8850 with QAIRT 2.47.0; the GPU rows ran the published files as they are. LiteRT can also compile for the NPU on the device at first load, which is what lets you ship the published file unchanged — that path and the ten runtime libraries it needs are in the NPU recipe, and we did not measure it here. GPU wiring is in the GPU recipe.

License

Apache-2.0 (RWKV / BlinkDL). Converted with litert-torch from the official RWKV-x070-World-0.1B-v2.8 checkpoint.

android
litert
on-device
rwkv
text-generation
tflite

litert-community/RWKV-7-World-0.1B-LiteRT

Model

Measured on device (edge-compat): browser · Chromium 151 on M4 Max · LiteRT.js 2.5.3 · WebGPU · 15.1 ms p50 · output matches CPU (2026-08-11). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/rwkv-7-world-0.1b/CARD.md

1

12 commits

4 linked in READMEs

updated Sep 8, 2026

See the code

README

Measured on device (edge-compat): browser · Chromium 151 on M4 Max · LiteRT.js 2.5.3 · WebGPU · 15.1 ms p50 · output matches CPU (2026-08-11). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/rwkv-7-world-0.1b/CARD.md

RWKV-7 World 0.1B — on-device text generation (LiteRT GPU)

The first autoregressive language model running its full forward pass on the LiteRT CompiledModel GPU delegate (RNN mode, host-side state; no CPU fallback for any op). RWKV-7 is an RNN: generation feeds one token per step and carries a small fixed-size recurrent state, so the whole model fits a single static GPU graph — no KV cache growth, no dynamic shapes.

  • Architecture: RWKV-x070-World-0.1B-v2.8 — 12 layers, d=768, 12 heads, vocab 65536.
  • Weights: BlinkDL/rwkv-7-world · Apache-2.0.
  • Size: 282 MB (fp16) + 100 MB host-side embedding table.

RWKV-7 on-device generation

Greedy generation on a Pixel 8a; the full per-token forward runs on the GPU.

I/O (per-token step graph)

TensorShapeRole
x_emb (in)[1, 768]embedding row of the current token (host lookup)
att_shift (in/out)[12, 768]per-layer attention token-shift state
ffn_shift (in/out)[12, 768]per-layer FFN token-shift state
wkv (in/out)[144, 64, 64]per-layer-per-head wkv state (12×12 heads)
logits (out)[1, 65536]next-token logits

Host side per step: look the token's row up in the fp16 embedding table (rwkv7_emb_fp16.bin, GATHER is not GPU-compatible; the first LayerNorm is inside the graph), run the graph, argmax the logits, and feed the three output states back in. Prefill = the same loop over the prompt tokens. Tokenizer: RWKV World greedy longest-match trie (rwkv_vocab_v20230424.txt).

GPU conversion

Fully GPU-resident on a Pixel 8a (1863/1863 nodes, 1 partition, ~18 ms/token fp16) via exact re-authorings, no approximation:

  • wkv7 recurrence at T=1 → plain 4D matmul/elementwise ops.
  • GroupNorm(heads) → manual per-head mean/var.
  • F.normalize → x * rsqrt(sum(x²) + eps).
  • softplus → branch-free relu(z) + log1p(exp(-|z|)) (the stock lowering emits GREATER+SELECT, rejected by the GPU delegate).
  • Token embedding lookup host-side.

Verified: sequential step-mode == parallel GPT-mode logits (corr 1.0000000); desktop fp16 CompiledModel corr 1.0000000 vs fp32 PyTorch; on-device 30-token greedy generation tracks desktop fp32 (28/30 tokens identical; the two divergences are fp32 near-ties with logit gap ≤ 0.04).

Minimal usage

Kotlin (Android, LiteRT CompiledModel GPU)

val model = CompiledModel.create(modelPath, CompiledModel.Options(Accelerator.GPU), null)
val inputs = model.createInputBuffers()
val outputs = model.createOutputBuffers()

var att = FloatArray(12 * 768); var ffn = FloatArray(12 * 768)
var wkv = FloatArray(144 * 64 * 64)
for (token in promptIds + generated) {
    inputs[0].writeFloat(embeddingRow(token))   // host fp16-table lookup
    inputs[1].writeFloat(att); inputs[2].writeFloat(ffn); inputs[3].writeFloat(wkv)
    model.run(inputs, outputs)
    val logits = outputs[0].readFloat()          // [65536] -> argmax = next token
    att = outputs[1].readFloat(); ffn = outputs[2].readFloat(); wkv = outputs[3].readFloat()
}

Python (LiteRT CompiledModel API)

import numpy as np
from ai_edge_litert.compiled_model import CompiledModel

model = CompiledModel.from_file("rwkv7_step_fp16.tflite")
inputs = model.create_input_buffers(0)
outputs = model.create_output_buffers(0)

emb = np.fromfile("rwkv7_emb_fp16.bin", "<f2").reshape(65536, 768)
att = np.zeros((12, 768), np.float32)
ffn = np.zeros((12, 768), np.float32)
wkv = np.zeros((144, 64, 64), np.float32)
for token in prompt_ids:
    inputs[0].write(emb[token : token + 1].astype(np.float32))
    inputs[1].write(att.ravel()); inputs[2].write(ffn.ravel()); inputs[3].write(wkv.ravel())
    model.run_by_index(0, inputs, outputs)
    logits = outputs[0].read(65536, np.float32)          # argmax -> next token
    att = outputs[1].read(12 * 768, np.float32)
    ffn = outputs[2].read(12 * 768, np.float32)
    wkv = outputs[3].read(144 * 64 * 64, np.float32)

Files

FileSizeRole
rwkv7_step_fp16.tflite282 MBper-token step graph (fp16 weights)
rwkv7_emb_fp16.bin100 MBembedding table [65536, 768] little-endian fp16, for host lookup
rwkv_vocab_v20230424.txt1.1 MBRWKV World vocabulary

Performance

Measured on a Pixel 8a (Tensor G3, Android 16) with the standard TFLite benchmark_model tool — 10 warm-up runs then 50 timed runs, reported as the tool's mean.

RuntimeBackendGraph on GPULatency
LiteRT CompiledModel (LITERT_CL)GPU1863 / 1863~18 ms
TFLite benchmark_model (TfLiteGpuDelegateV2)GPU (OpenCL)113 / 1863123.1 ms
TFLite benchmark_modelCPU (XNNPACK, 4 threads)—40.2 ms

The two GPU rows are different runtimes, not a contradiction. The LITERT_CL figure is the one recorded when this model shipped, taken through LiteRT's own CompiledModel accelerator — the path the Kotlin sample app and the LiteRT API use. The TfLiteGpuDelegateV2 figure is the classic TFLite OpenCL delegate, measured with a tool anyone can download and re-run. They agree on how much of the graph the GPU takes; they disagree on speed, and the classic delegate is the slower of the two here. Read the TfLiteGpuDelegateV2 row as a reproducible floor, not as this model's speed on LiteRT.

On this delegate the CPU is the faster choice for (40.2 ms on CPU against 123.1 ms on GPU) — worth knowing before you reach for the GPU on a mid-range phone.

Note that the GPU does not take the whole graph here (113 / 1863); the remainder runs on the CPU and the split costs a per-partition round trip.

Snapdragon NPU (Hexagon)

The NPU is 1.70x faster than the GPU (5.37 ms against 9.11 ms) and loads 8.28x faster (270 ms against 2236 ms).

backendinference (median / min)load
NPU (Hexagon v81)5.37 ms / 5.33 ms270 ms
GPU (Adreno)9.11 ms / 8.37 ms2236 ms

Measured on a Samsung Galaxy S26 (Snapdragon 8 Elite Gen 5 / SM8850, Hexagon v81, Android 16), LiteRT CompiledModel 2.2.0, one accelerator per process, 5 warm-up runs then N=50 timed runs, median reported. Every run held thermal status NONE throughout. Headroom 0.77-0.78, where 1.0 is the throttling threshold.

The NPU rows here ran artifacts compiled ahead of time for SM8850 with QAIRT 2.47.0; the GPU rows ran the published files as they are. LiteRT can also compile for the NPU on the device at first load, which is what lets you ship the published file unchanged — that path and the ten runtime libraries it needs are in the NPU recipe, and we did not measure it here. GPU wiring is in the GPU recipe.

License

Apache-2.0 (RWKV / BlinkDL). Converted with litert-torch from the official RWKV-x070-World-0.1B-v2.8 checkpoint.

android
litert
on-device
rwkv
text-generation
tflite