litert-community/Qwen3-Reranker-0.6B-LiteRT

Model

Measured on device (edge-compat): Galaxy S26 · LiteRT 2.2.0 · GPU (ML Drift) · 78.5 ms p50 (2026-08-27); Galaxy S26 · LiteRT 2.2.0 · NPU (QNN/HTP) · 47.5 ms p50 (2026-08-27); browser · Chromium 151 on M4 Max · LiteRT.js 2.5.3 · WebGPU · did not run (2026-08-11). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/qwen3-reranker-0.6b/CARD.md

0

13 commits

3 linked in READMEs

updated Sep 8, 2026

See the code

README

Measured on device (edge-compat): Galaxy S26 · LiteRT 2.2.0 · GPU (ML Drift) · 78.5 ms p50 (2026-08-27); Galaxy S26 · LiteRT 2.2.0 · NPU (QNN/HTP) · 47.5 ms p50 (2026-08-27); browser · Chromium 151 on M4 Max · LiteRT.js 2.5.3 · WebGPU · did not run (2026-08-11). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/qwen3-reranker-0.6b/CARD.md

Qwen3-Reranker-0.6B — LiteRT on-device RAG reranker (fully GPU)

Qwen3-Reranker-0.6B (Apache-2.0), the 2025 SOTA small reranker, re-authored to run entirely on the LiteRT CompiledModel GPU (ML Drift). Given a query and candidate documents, it scores each by relevance (P("yes")) and reorders them — the reranking half of an on-device RAG pipeline.

Pairs with litert-community/Qwen3-Embedding-0.6B-LiteRT: embed → retrieve top-k → rerank, all on-device, no server.

On-device RAG reranking on a Pixel 8a

Like the embedder it is a single forward pass (no generation, no KV cache) → a plain .tflite, not a .litertlm. Verified on a Pixel 8a / Tensor G3: all nodes on the GPU delegate, P(yes) parity ref 0.9995 / dev 0.9994 vs the HF fp32 reference.

Files

filepurposeruns on
qwen3rerank_gpu_fp16.tflite28-layer Qwen3 decoder + baked 2-logit head, inputs_embeds[1,256,1024] → logits[1,256,2]GPU
embeddings_fp16.bintied token-embedding table [151669,1024] fp16, for the host-side lookuphost
vocab.json, merges.txtQwen byte-level BPE tokenizerhost

How it scores

prompt = PREFIX + "<Instruct>:… <Query>:… <Document>:…" + SUFFIX     (Qwen3-Reranker template)
       →[host embed lookup]→ inputs_embeds[1,256,1024]
       →[GPU: 28-layer decoder + 2-logit head]→ logits[1,256,2]
       →[softmax over (no,yes) at the last token]→ P(yes) = relevance

The 2-logit head bakes the tied-embedding rows for "no" (2152) and "yes" (9693), so the graph emits [no,yes] directly. The host right-pads and pools the last real token (causal ⇒ it never sees the trailing pad, identical to the official left-pad + attention-mask). Token embedding is a GATHER (GPU-banned) so it is done host-side.

The GPU-clean re-authoring is the same as the embedder (host-embed, GQA cat-repeat to avoid BROADCAST_TO, max-normalized RMSNorm for the deep-stack fp16 overflow, baked RoPE / causal mask).

Minimal usage

Python (reference score with the original model):

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-Reranker-0.6B")
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-Reranker-0.6B").eval()
yes, no = tok.convert_tokens_to_ids("yes"), tok.convert_tokens_to_ids("no")
# … build the PREFIX/SUFFIX prompt, then:
logits = model(**inputs).logits[:, -1, :]
score = torch.softmax(torch.stack([logits[:, no], logits[:, yes]], 1), 1)[:, 1]  # P(yes)

Kotlin (on-device, LiteRT CompiledModel GPU):

val model = CompiledModel.create("qwen3rerank_gpu_fp16.tflite",
    CompiledModel.Options(Accelerator.GPU), null)
// host: build prompt ids -> lookup embeddings_fp16.bin -> inputs_embeds[1,256,1024]
inputs[0].writeFloat(embedLookup(promptIds(query, doc)))
model.run(inputs, outputs)
val logits = outputs[0].readFloat()               // [256,2] = [no,yes] per position
val score = softmaxYes(logits, poolPos)           // P(yes) relevance

Full tokenizer + prompt template + reranking app: see the official LiteRT sample.

Conversion

Reproducible in the official sample's conversion/ (build_qwen3rerank.py, export_embeddings.py, device-parity harness).

Performance

Measured on a Pixel 8a (Tensor G3, Android 16) with the standard TFLite benchmark_model tool — 10 warm-up runs then 50 timed runs, reported as the tool's mean.

RuntimeBackendGraph on GPULatency
TFLite benchmark_model (TfLiteGpuDelegateV2)GPU (OpenCL)108 / 32669277.4 ms
TFLite benchmark_modelCPU (XNNPACK, 4 threads)—4066.1 ms

Any on-device figure recorded when this model shipped came from a different runtime. It was taken through LiteRT's own CompiledModel accelerator (logcat reports it as LITERT_CL), which is the path the Kotlin sample app and the LiteRT API use, and it appears elsewhere on this card. The rows above are the classic TFLite OpenCL delegate, measured with a tool anyone can download and re-run. The two are not comparable, so read the rows above as a reproducible floor rather than as this model's speed on LiteRT.

On this delegate the CPU is the faster choice (4066.1 ms on CPU against 9277.4 ms on GPU) — worth knowing before you reach for the GPU on a mid-range phone.

Note that the GPU does not take the whole graph here (108 / 3266); the remainder runs on the CPU and the split costs a per-partition round trip.

Snapdragon NPU (Hexagon)

The NPU is 1.65x faster than the GPU (47.51 ms against 78.49 ms) and loads 13.35x faster (320 ms against 4265 ms).

backendcompiledinference (median / min)load
NPU (Hexagon v81)AOT (SM8850)47.51 ms / 45.89 ms320 ms
GPU (Adreno)—78.49 ms / 77.55 ms4265 ms

Measured on a Samsung Galaxy S26 (Snapdragon 8 Elite Gen 5 / SM8850, Hexagon v81, Android 16) with LiteRT CompiledModel 2.2.0, one accelerator per process, 5 warm-up runs then N=50 timed runs, median reported. Every run held thermal status NONE throughout. Headroom 0.77–0.78, where 1.0 is the throttling threshold.

The NPU row marked AOT ran an artifact compiled ahead of time for SM8850 (ai-edge-litert 2.2.0 + QAIRT 2.47.0), not the published file. That artifact is not distributed here; the compile is one command in the NPU guide.

GPU wiring: GPU guide.

android
cross-encoder
litert
on-device
qwen3
qwen3-reranker
reranker
reranking
retrieval
text-ranking
tflite

litert-community/Qwen3-Reranker-0.6B-LiteRT

Model

Measured on device (edge-compat): Galaxy S26 · LiteRT 2.2.0 · GPU (ML Drift) · 78.5 ms p50 (2026-08-27); Galaxy S26 · LiteRT 2.2.0 · NPU (QNN/HTP) · 47.5 ms p50 (2026-08-27); browser · Chromium 151 on M4 Max · LiteRT.js 2.5.3 · WebGPU · did not run (2026-08-11). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/qwen3-reranker-0.6b/CARD.md

0

13 commits

3 linked in READMEs

updated Sep 8, 2026

See the code

README

Measured on device (edge-compat): Galaxy S26 · LiteRT 2.2.0 · GPU (ML Drift) · 78.5 ms p50 (2026-08-27); Galaxy S26 · LiteRT 2.2.0 · NPU (QNN/HTP) · 47.5 ms p50 (2026-08-27); browser · Chromium 151 on M4 Max · LiteRT.js 2.5.3 · WebGPU · did not run (2026-08-11). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/qwen3-reranker-0.6b/CARD.md

Qwen3-Reranker-0.6B — LiteRT on-device RAG reranker (fully GPU)

Qwen3-Reranker-0.6B (Apache-2.0), the 2025 SOTA small reranker, re-authored to run entirely on the LiteRT CompiledModel GPU (ML Drift). Given a query and candidate documents, it scores each by relevance (P("yes")) and reorders them — the reranking half of an on-device RAG pipeline.

Pairs with litert-community/Qwen3-Embedding-0.6B-LiteRT: embed → retrieve top-k → rerank, all on-device, no server.

On-device RAG reranking on a Pixel 8a

Like the embedder it is a single forward pass (no generation, no KV cache) → a plain .tflite, not a .litertlm. Verified on a Pixel 8a / Tensor G3: all nodes on the GPU delegate, P(yes) parity ref 0.9995 / dev 0.9994 vs the HF fp32 reference.

Files

filepurposeruns on
qwen3rerank_gpu_fp16.tflite28-layer Qwen3 decoder + baked 2-logit head, inputs_embeds[1,256,1024] → logits[1,256,2]GPU
embeddings_fp16.bintied token-embedding table [151669,1024] fp16, for the host-side lookuphost
vocab.json, merges.txtQwen byte-level BPE tokenizerhost

How it scores

prompt = PREFIX + "<Instruct>:… <Query>:… <Document>:…" + SUFFIX     (Qwen3-Reranker template)
       →[host embed lookup]→ inputs_embeds[1,256,1024]
       →[GPU: 28-layer decoder + 2-logit head]→ logits[1,256,2]
       →[softmax over (no,yes) at the last token]→ P(yes) = relevance

The 2-logit head bakes the tied-embedding rows for "no" (2152) and "yes" (9693), so the graph emits [no,yes] directly. The host right-pads and pools the last real token (causal ⇒ it never sees the trailing pad, identical to the official left-pad + attention-mask). Token embedding is a GATHER (GPU-banned) so it is done host-side.

The GPU-clean re-authoring is the same as the embedder (host-embed, GQA cat-repeat to avoid BROADCAST_TO, max-normalized RMSNorm for the deep-stack fp16 overflow, baked RoPE / causal mask).

Minimal usage

Python (reference score with the original model):

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-Reranker-0.6B")
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-Reranker-0.6B").eval()
yes, no = tok.convert_tokens_to_ids("yes"), tok.convert_tokens_to_ids("no")
# … build the PREFIX/SUFFIX prompt, then:
logits = model(**inputs).logits[:, -1, :]
score = torch.softmax(torch.stack([logits[:, no], logits[:, yes]], 1), 1)[:, 1]  # P(yes)

Kotlin (on-device, LiteRT CompiledModel GPU):

val model = CompiledModel.create("qwen3rerank_gpu_fp16.tflite",
    CompiledModel.Options(Accelerator.GPU), null)
// host: build prompt ids -> lookup embeddings_fp16.bin -> inputs_embeds[1,256,1024]
inputs[0].writeFloat(embedLookup(promptIds(query, doc)))
model.run(inputs, outputs)
val logits = outputs[0].readFloat()               // [256,2] = [no,yes] per position
val score = softmaxYes(logits, poolPos)           // P(yes) relevance

Full tokenizer + prompt template + reranking app: see the official LiteRT sample.

Conversion

Reproducible in the official sample's conversion/ (build_qwen3rerank.py, export_embeddings.py, device-parity harness).

Performance

Measured on a Pixel 8a (Tensor G3, Android 16) with the standard TFLite benchmark_model tool — 10 warm-up runs then 50 timed runs, reported as the tool's mean.

RuntimeBackendGraph on GPULatency
TFLite benchmark_model (TfLiteGpuDelegateV2)GPU (OpenCL)108 / 32669277.4 ms
TFLite benchmark_modelCPU (XNNPACK, 4 threads)—4066.1 ms

Any on-device figure recorded when this model shipped came from a different runtime. It was taken through LiteRT's own CompiledModel accelerator (logcat reports it as LITERT_CL), which is the path the Kotlin sample app and the LiteRT API use, and it appears elsewhere on this card. The rows above are the classic TFLite OpenCL delegate, measured with a tool anyone can download and re-run. The two are not comparable, so read the rows above as a reproducible floor rather than as this model's speed on LiteRT.

On this delegate the CPU is the faster choice (4066.1 ms on CPU against 9277.4 ms on GPU) — worth knowing before you reach for the GPU on a mid-range phone.

Note that the GPU does not take the whole graph here (108 / 3266); the remainder runs on the CPU and the split costs a per-partition round trip.

Snapdragon NPU (Hexagon)

The NPU is 1.65x faster than the GPU (47.51 ms against 78.49 ms) and loads 13.35x faster (320 ms against 4265 ms).

backendcompiledinference (median / min)load
NPU (Hexagon v81)AOT (SM8850)47.51 ms / 45.89 ms320 ms
GPU (Adreno)—78.49 ms / 77.55 ms4265 ms

Measured on a Samsung Galaxy S26 (Snapdragon 8 Elite Gen 5 / SM8850, Hexagon v81, Android 16) with LiteRT CompiledModel 2.2.0, one accelerator per process, 5 warm-up runs then N=50 timed runs, median reported. Every run held thermal status NONE throughout. Headroom 0.77–0.78, where 1.0 is the throttling threshold.

The NPU row marked AOT ran an artifact compiled ahead of time for SM8850 (ai-edge-litert 2.2.0 + QAIRT 2.47.0), not the published file. That artifact is not distributed here; the compile is one command in the NPU guide.

GPU wiring: GPU guide.

android
cross-encoder
litert
on-device
qwen3
qwen3-reranker
reranker
reranking
retrieval
text-ranking
tflite