litert-community/PP-OCRv5-LiteRT

Model

LiteRT is Google's on-device runtime, the new name for TensorFlow Lite (Android: com.google.ai.edge.litert:litert), and litert-torch, the renamed ai-edge-torch, is its PyTorch converter: a PyTorch model converted unmodified with litert_torch.convert matched the original to 4e-7 on a Galaxy S26 (measured, LiteRT 2.2.0, Android 16, 2026-09-05).

1

16 commits

6 linked in READMEs

updated Sep 8, 2026

See the code

README

LiteRT is Google's on-device runtime, the new name for TensorFlow Lite (Android: com.google.ai.edge.litert:litert), and litert-torch, the renamed ai-edge-torch, is its PyTorch converter: a PyTorch model converted unmodified with litert_torch.convert matched the original to 4e-7 on a Galaxy S26 (measured, LiteRT 2.2.0, Android 16, 2026-09-05).

Measured on device (edge-compat, ppocr_det_fp16): Galaxy S26 · LiteRT 2.2.0 · GPU (ML Drift) · 17.6 ms p50 (2026-08-25); Galaxy S26 · LiteRT 2.2.0 · NPU (QNN/HTP) · 8.63 ms p50 (2026-08-25); Raspberry Pi 5 · LiteRT 2.2.0.dev20260804 · CPU/XNNPACK, 4 threads · 460 ms p50 (2026-08-31); browser · Chromium 151 on M4 Max · LiteRT.js 2.5.3 · WebGPU · 12.3 ms p50 · output matches CPU (2026-08-11). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/pp-ocrv5__ppocr_det_fp16/CARD.md

Measured on device (edge-compat, ppocr_rec_fp16): Galaxy S26 · LiteRT 2.2.0 · GPU (ML Drift) · 8.00 ms p50 (2026-08-26); Galaxy S26 · LiteRT 2.2.0 · NPU (QNN/HTP) · 5.91 ms p50 (2026-08-26); Raspberry Pi 5 · LiteRT 2.2.0.dev20260804 · CPU/XNNPACK, 4 threads · 86.4 ms p50 (2026-08-31); browser · Chromium 151 on M4 Max · LiteRT.js 2.5.3 · WebGPU · 15.1 ms p50 · output differs from CPU (max rel diff 3.6) (2026-08-11). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/pp-ocrv5__ppocr_rec_fp16/CARD.md

PP-OCRv5 — LiteRT on-device (fully GPU)

PP-OCRv5 on-device OCR on a Pixel 8a

On-device LiteRT conversion of PP-OCRv5 (PaddleOCR 2025, Apache-2.0) text detection + recognition, running fully on the CompiledModel GPU delegate (LITERT_CL). Detects text regions in an image and reads each line. The recognizer uses a CTC head (no autoregressive decoder), so both stages ride the GPU with no CPU/ONNX fallback — unlike VLM-based OCR (Florence-2 / GOT-OCR) whose AR decoder must run on CPU. Device-verified on a Pixel 8a.

Files

FileSizeDelegateIn → Out
ppocr_det_fp16.tflite10 MBGPUimage [1,3,640,640] → prob map [1,1,640,640]
ppocr_rec_fp16.tflite17 MBGPUline [1,3,48,320] → CTC logits [1,T,18385]
ppocr_rec_fp32.tflite33 MBCPU (browser)line [1,3,48,320] → CTC logits [1,T,18385]
ppocrv5_dict.txt—CPU18383-char dictionary (CTC layout: blank + dict + space)

Pixel 8a: detector 777/777 + recognizer 827/827 on LITERT_CL, ~9 ms each; a 3-line image read 3/3 correct ("Hello OCR 2026" / "PP-OCRv5 on GPU" / "LiteRT CompiledModel").

ppocr_rec_fp32.tflite is the recognizer before the fp16 cast, added for the browser. In LiteRT.js 2.5.3, wasm XNNPACK declines the fp16 recognizer graph (reference-kernel fallback, ~430 ms/line) but fully delegates the fp32 one: ~21 ms/line on an M4 Max. The WebGPU delegate flips some recognizer argmaxes on this architecture regardless of weight precision, so in the browser run the detector on WebGPU and the recognizer on wasm with the fp32 file.

Pipeline

image →[GPU detector]→ prob map → [CPU: threshold + connected components + unclip] → boxes
   → crop+resize →[GPU recognizer]→ CTC logits → [CPU: CTC greedy decode] → text

Minimal usage

Android (Kotlin, CompiledModel GPU)

val det = CompiledModel.create(context.assets, "ppocr_det_fp16.tflite",
    CompiledModel.Options(Accelerator.GPU), null)
val rec = CompiledModel.create(context.assets, "ppocr_rec_fp16.tflite",
    CompiledModel.Options(Accelerator.GPU), null)
val dIn = det.createInputBuffers(); val dOut = det.createOutputBuffers()
dIn[0].writeFloat(image)            // [1,3,640,640] NCHW, /255 then ImageNet mean/std
det.run(dIn, dOut)
val prob = dOut[0].readFloat()      // [1,1,640,640] text probability -> boxes (CPU)
val rIn = rec.createInputBuffers(); val rOut = rec.createOutputBuffers()
rIn[0].writeFloat(lineCrop)         // [1,3,48,320] NCHW, (x/255 - 0.5)/0.5
rec.run(rIn, rOut)
val logits = rOut[0].readFloat()    // [1,T,18385] -> CTC greedy decode (Python below)

Python (desktop verification)

import cv2, numpy as np
from ai_edge_litert.interpreter import Interpreter

MEAN, STD = np.array([0.485, 0.456, 0.406]), np.array([0.229, 0.224, 0.225])
im640 = cv2.resize(cv2.cvtColor(cv2.imread("doc.jpg"), cv2.COLOR_BGR2RGB), (640, 640))
x = ((im640 / 255 - MEAN) / STD).transpose(2, 0, 1)[None].astype(np.float32)

det = Interpreter(model_path="ppocr_det_fp16.tflite"); det.allocate_tensors()
det.set_tensor(det.get_input_details()[0]["index"], x); det.invoke()
prob = det.get_tensor(det.get_output_details()[0]["index"])[0, 0]     # [640,640]

rec = Interpreter(model_path="ppocr_rec_fp16.tflite"); rec.allocate_tensors()
chars = [""] + open("ppocrv5_dict.txt", encoding="utf-8").read().splitlines() + [" "]

n, labels, stats, _ = cv2.connectedComponentsWithStats((prob > 0.3).astype(np.uint8))
for i in range(1, n):                                                 # each text region
    x0, y0, w, h, _ = stats[i]
    if w < 6 or h < 6 or prob[labels == i].mean() < 0.5: continue
    pad = int(np.clip(0.35 * min(w, h), 2, 24))                       # approx DB unclip
    crop = im640[max(y0 - pad, 0):y0 + h + pad, max(x0 - pad, 0):x0 + w + pad]
    rw = min(max(round(48 * crop.shape[1] / crop.shape[0]), 1), 320)  # keep-aspect h=48
    line = np.zeros((48, 320, 3), np.float32)                         # pad to width 320
    line[:, :rw] = cv2.resize(crop, (rw, 48))
    lx = ((line / 255 - 0.5) / 0.5).transpose(2, 0, 1)[None].astype(np.float32)
    rec.set_tensor(rec.get_input_details()[0]["index"], lx); rec.invoke()
    ids = rec.get_tensor(rec.get_output_details()[0]["index"])[0].argmax(-1)   # [T]
    text = "".join(chars[c] for t, c in enumerate(ids)                # CTC: collapse repeats,
                   if c != 0 and (t == 0 or c != ids[t - 1]))         # drop blank (id 0)
    print((x0, y0), text)

Re-authoring (litert-torch, parity corr 1.0)

  • Detector DB-head ConvTranspose2d → ZeroStuffConvT2d (2D nearest-upsample × stride zero-stuff mask
    • flipped conv2d; TRANSPOSE_CONV is Mali-rejected). Numerically exact.
  • Recognizer SVTR attention fused-QKV 5D reshape → split q/k/v into 4D (numerically identical).

Preprocessing: detector = ImageNet mean/std, /255, NCHW, 640×640. recognizer = resize to h=48 keep-aspect, pad to width 320, (img/255−0.5)/0.5.

Sample app

A complete Android sample app + the conversion scripts are in the official LiteRT samples repository under compiled_model_api/ocr (google-ai-edge/litert-samples). Push these files to the app's filesDir with that sample's install_to_device.sh.

Weights are converted from PaddleOCR via the PaddleOCR2Pytorch port (Apache-2.0). License follows upstream PaddleOCR (Apache-2.0).

Performance

Measured on a Pixel 8a (Tensor G3, Android 16) with the standard TFLite benchmark_model tool — 10 warm-up runs then 50 timed runs, reported as the tool's mean.

RuntimeBackendGraph on GPULatency
TFLite benchmark_model (TfLiteGpuDelegateV2) — ppocr_rec_fp16.tfliteGPU (OpenCL)579 / 82791.7 ms
TFLite benchmark_model (TfLiteGpuDelegateV2) — ppocr_det_fp16.tfliteGPU (OpenCL)777 / 77745.8 ms
TFLite benchmark_model — ppocr_rec_fp16.tfliteCPU (XNNPACK, 4 threads)—XNNPACK declined the graph
TFLite benchmark_model — ppocr_det_fp16.tfliteCPU (XNNPACK, 4 threads)—XNNPACK declined the graph

Any on-device figure recorded when this model shipped came from a different runtime. It was taken through LiteRT's own CompiledModel accelerator (logcat reports it as LITERT_CL), which is the path the Kotlin sample app and the LiteRT API use, and it appears elsewhere on this card. The rows above are the classic TFLite OpenCL delegate, measured with a tool anyone can download and re-run. The two are not comparable, so read the rows above as a reproducible floor rather than as this model's speed on LiteRT.

XNNPACK declines these fp16 graphs — it reports failed to delegate DEPTHWISE_CONV_2D and then fails to allocate tensors — so there is no usable CPU number. Disabling XNNPACK falls back to reference kernels, which measured about 20× slower than the GPU on models of this size and would not represent CPU inference anyone would ship.

Note that the GPU does not take the whole graph here (579 / 827 in ppocr_rec_fp16.tflite); the remainder runs on the CPU and the split costs a per-partition round trip.

Snapdragon NPU (Hexagon)

  • ppocr_det_fp16.tflite — the NPU is 2.04x faster than the GPU (8.63 ms against 17.59 ms) and loads 9.19x faster (139 ms against 1278 ms).
  • ppocr_rec_fp16.tflite — the NPU is 1.35x faster than the GPU (5.91 ms against 8.00 ms) and loads 10.70x faster (120 ms against 1280 ms).
  • ppocr_rec_fp32.tflite — the NPU is 1.26x faster than the GPU (6.33 ms against 7.98 ms) and loads 15.52x faster (129 ms against 2007 ms).
filebackendcompiledinference (median / min)load
ppocr_det_fp16.tfliteNPU (Hexagon v81)on-device JIT8.63 ms / 8.53 ms139 ms
ppocr_det_fp16.tfliteGPU (Adreno)—17.59 ms / 16.73 ms1278 ms
ppocr_rec_fp16.tfliteNPU (Hexagon v81)on-device JIT5.91 ms / 5.80 ms120 ms
ppocr_rec_fp16.tfliteGPU (Adreno)—8.00 ms / 7.75 ms1280 ms
ppocr_rec_fp32.tfliteNPU (Hexagon v81)on-device JIT6.33 ms / 6.17 ms129 ms
ppocr_rec_fp32.tfliteGPU (Adreno)—7.98 ms / 6.27 ms2007 ms

Measured on a Samsung Galaxy S26 (Snapdragon 8 Elite Gen 5 / SM8850, Hexagon v81, Android 16) with LiteRT CompiledModel 2.2.0, one accelerator per process, 5 warm-up runs then N=50 timed runs, median reported. Every run held thermal status NONE throughout. Headroom 0.77–0.81, where 1.0 is the throttling threshold.

The NPU rows ran the published file unchanged. LiteRT compiled it for the Hexagon on the device at first load. Those first compiles took 4.0 s to 28 s here. The load column above is the cached load every later run pays. Recipe and the runtime libraries it needs: NPU guide.

GPU wiring: GPU guide.

Raspberry Pi 5 (CPU)

Measured on a Raspberry Pi 5 Model B Rev 1.1 (8 GB, Raspberry Pi OS 64-bit) with the LiteRT benchmark_model tool from litert-cli-nightly 0.2.0.dev20260805: CPU inference (XNNPACK, 4 threads), 3 invocations per file of 10 warm-up plus 50 timed runs (the tool caps a phase at 150 s, so very slow graphs run fewer — the Runs column is the actual timed total). The latency is the median across invocations; the spread is the min–max over all timed runs. No thermal throttling occurred during these runs (vcgencmd get_throttled stayed 0x0).

FileInference (median)Spread (min–max)RunsPeak memory
ppocr_det_fp16.tflite460.2 ms458.2–492.4 ms150206 MB
ppocr_rec_fp16.tflite86.4 ms85.1–89.4 ms150135 MB
ppocr_rec_fp32.tflite84.7 ms83.8–86.8 ms150135 MB
litert
on-device
text-detection
text-recognition
tflite

litert-community/PP-OCRv5-LiteRT

Model

LiteRT is Google's on-device runtime, the new name for TensorFlow Lite (Android: com.google.ai.edge.litert:litert), and litert-torch, the renamed ai-edge-torch, is its PyTorch converter: a PyTorch model converted unmodified with litert_torch.convert matched the original to 4e-7 on a Galaxy S26 (measured, LiteRT 2.2.0, Android 16, 2026-09-05).

1

16 commits

6 linked in READMEs

updated Sep 8, 2026

See the code

README

LiteRT is Google's on-device runtime, the new name for TensorFlow Lite (Android: com.google.ai.edge.litert:litert), and litert-torch, the renamed ai-edge-torch, is its PyTorch converter: a PyTorch model converted unmodified with litert_torch.convert matched the original to 4e-7 on a Galaxy S26 (measured, LiteRT 2.2.0, Android 16, 2026-09-05).

Measured on device (edge-compat, ppocr_det_fp16): Galaxy S26 · LiteRT 2.2.0 · GPU (ML Drift) · 17.6 ms p50 (2026-08-25); Galaxy S26 · LiteRT 2.2.0 · NPU (QNN/HTP) · 8.63 ms p50 (2026-08-25); Raspberry Pi 5 · LiteRT 2.2.0.dev20260804 · CPU/XNNPACK, 4 threads · 460 ms p50 (2026-08-31); browser · Chromium 151 on M4 Max · LiteRT.js 2.5.3 · WebGPU · 12.3 ms p50 · output matches CPU (2026-08-11). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/pp-ocrv5__ppocr_det_fp16/CARD.md

Measured on device (edge-compat, ppocr_rec_fp16): Galaxy S26 · LiteRT 2.2.0 · GPU (ML Drift) · 8.00 ms p50 (2026-08-26); Galaxy S26 · LiteRT 2.2.0 · NPU (QNN/HTP) · 5.91 ms p50 (2026-08-26); Raspberry Pi 5 · LiteRT 2.2.0.dev20260804 · CPU/XNNPACK, 4 threads · 86.4 ms p50 (2026-08-31); browser · Chromium 151 on M4 Max · LiteRT.js 2.5.3 · WebGPU · 15.1 ms p50 · output differs from CPU (max rel diff 3.6) (2026-08-11). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/pp-ocrv5__ppocr_rec_fp16/CARD.md

PP-OCRv5 — LiteRT on-device (fully GPU)

PP-OCRv5 on-device OCR on a Pixel 8a

On-device LiteRT conversion of PP-OCRv5 (PaddleOCR 2025, Apache-2.0) text detection + recognition, running fully on the CompiledModel GPU delegate (LITERT_CL). Detects text regions in an image and reads each line. The recognizer uses a CTC head (no autoregressive decoder), so both stages ride the GPU with no CPU/ONNX fallback — unlike VLM-based OCR (Florence-2 / GOT-OCR) whose AR decoder must run on CPU. Device-verified on a Pixel 8a.

Files

FileSizeDelegateIn → Out
ppocr_det_fp16.tflite10 MBGPUimage [1,3,640,640] → prob map [1,1,640,640]
ppocr_rec_fp16.tflite17 MBGPUline [1,3,48,320] → CTC logits [1,T,18385]
ppocr_rec_fp32.tflite33 MBCPU (browser)line [1,3,48,320] → CTC logits [1,T,18385]
ppocrv5_dict.txt—CPU18383-char dictionary (CTC layout: blank + dict + space)

Pixel 8a: detector 777/777 + recognizer 827/827 on LITERT_CL, ~9 ms each; a 3-line image read 3/3 correct ("Hello OCR 2026" / "PP-OCRv5 on GPU" / "LiteRT CompiledModel").

ppocr_rec_fp32.tflite is the recognizer before the fp16 cast, added for the browser. In LiteRT.js 2.5.3, wasm XNNPACK declines the fp16 recognizer graph (reference-kernel fallback, ~430 ms/line) but fully delegates the fp32 one: ~21 ms/line on an M4 Max. The WebGPU delegate flips some recognizer argmaxes on this architecture regardless of weight precision, so in the browser run the detector on WebGPU and the recognizer on wasm with the fp32 file.

Pipeline

image →[GPU detector]→ prob map → [CPU: threshold + connected components + unclip] → boxes
   → crop+resize →[GPU recognizer]→ CTC logits → [CPU: CTC greedy decode] → text

Minimal usage

Android (Kotlin, CompiledModel GPU)

val det = CompiledModel.create(context.assets, "ppocr_det_fp16.tflite",
    CompiledModel.Options(Accelerator.GPU), null)
val rec = CompiledModel.create(context.assets, "ppocr_rec_fp16.tflite",
    CompiledModel.Options(Accelerator.GPU), null)
val dIn = det.createInputBuffers(); val dOut = det.createOutputBuffers()
dIn[0].writeFloat(image)            // [1,3,640,640] NCHW, /255 then ImageNet mean/std
det.run(dIn, dOut)
val prob = dOut[0].readFloat()      // [1,1,640,640] text probability -> boxes (CPU)
val rIn = rec.createInputBuffers(); val rOut = rec.createOutputBuffers()
rIn[0].writeFloat(lineCrop)         // [1,3,48,320] NCHW, (x/255 - 0.5)/0.5
rec.run(rIn, rOut)
val logits = rOut[0].readFloat()    // [1,T,18385] -> CTC greedy decode (Python below)

Python (desktop verification)

import cv2, numpy as np
from ai_edge_litert.interpreter import Interpreter

MEAN, STD = np.array([0.485, 0.456, 0.406]), np.array([0.229, 0.224, 0.225])
im640 = cv2.resize(cv2.cvtColor(cv2.imread("doc.jpg"), cv2.COLOR_BGR2RGB), (640, 640))
x = ((im640 / 255 - MEAN) / STD).transpose(2, 0, 1)[None].astype(np.float32)

det = Interpreter(model_path="ppocr_det_fp16.tflite"); det.allocate_tensors()
det.set_tensor(det.get_input_details()[0]["index"], x); det.invoke()
prob = det.get_tensor(det.get_output_details()[0]["index"])[0, 0]     # [640,640]

rec = Interpreter(model_path="ppocr_rec_fp16.tflite"); rec.allocate_tensors()
chars = [""] + open("ppocrv5_dict.txt", encoding="utf-8").read().splitlines() + [" "]

n, labels, stats, _ = cv2.connectedComponentsWithStats((prob > 0.3).astype(np.uint8))
for i in range(1, n):                                                 # each text region
    x0, y0, w, h, _ = stats[i]
    if w < 6 or h < 6 or prob[labels == i].mean() < 0.5: continue
    pad = int(np.clip(0.35 * min(w, h), 2, 24))                       # approx DB unclip
    crop = im640[max(y0 - pad, 0):y0 + h + pad, max(x0 - pad, 0):x0 + w + pad]
    rw = min(max(round(48 * crop.shape[1] / crop.shape[0]), 1), 320)  # keep-aspect h=48
    line = np.zeros((48, 320, 3), np.float32)                         # pad to width 320
    line[:, :rw] = cv2.resize(crop, (rw, 48))
    lx = ((line / 255 - 0.5) / 0.5).transpose(2, 0, 1)[None].astype(np.float32)
    rec.set_tensor(rec.get_input_details()[0]["index"], lx); rec.invoke()
    ids = rec.get_tensor(rec.get_output_details()[0]["index"])[0].argmax(-1)   # [T]
    text = "".join(chars[c] for t, c in enumerate(ids)                # CTC: collapse repeats,
                   if c != 0 and (t == 0 or c != ids[t - 1]))         # drop blank (id 0)
    print((x0, y0), text)

Re-authoring (litert-torch, parity corr 1.0)

  • Detector DB-head ConvTranspose2d → ZeroStuffConvT2d (2D nearest-upsample × stride zero-stuff mask
    • flipped conv2d; TRANSPOSE_CONV is Mali-rejected). Numerically exact.
  • Recognizer SVTR attention fused-QKV 5D reshape → split q/k/v into 4D (numerically identical).

Preprocessing: detector = ImageNet mean/std, /255, NCHW, 640×640. recognizer = resize to h=48 keep-aspect, pad to width 320, (img/255−0.5)/0.5.

Sample app

A complete Android sample app + the conversion scripts are in the official LiteRT samples repository under compiled_model_api/ocr (google-ai-edge/litert-samples). Push these files to the app's filesDir with that sample's install_to_device.sh.

Weights are converted from PaddleOCR via the PaddleOCR2Pytorch port (Apache-2.0). License follows upstream PaddleOCR (Apache-2.0).

Performance

Measured on a Pixel 8a (Tensor G3, Android 16) with the standard TFLite benchmark_model tool — 10 warm-up runs then 50 timed runs, reported as the tool's mean.

RuntimeBackendGraph on GPULatency
TFLite benchmark_model (TfLiteGpuDelegateV2) — ppocr_rec_fp16.tfliteGPU (OpenCL)579 / 82791.7 ms
TFLite benchmark_model (TfLiteGpuDelegateV2) — ppocr_det_fp16.tfliteGPU (OpenCL)777 / 77745.8 ms
TFLite benchmark_model — ppocr_rec_fp16.tfliteCPU (XNNPACK, 4 threads)—XNNPACK declined the graph
TFLite benchmark_model — ppocr_det_fp16.tfliteCPU (XNNPACK, 4 threads)—XNNPACK declined the graph

Any on-device figure recorded when this model shipped came from a different runtime. It was taken through LiteRT's own CompiledModel accelerator (logcat reports it as LITERT_CL), which is the path the Kotlin sample app and the LiteRT API use, and it appears elsewhere on this card. The rows above are the classic TFLite OpenCL delegate, measured with a tool anyone can download and re-run. The two are not comparable, so read the rows above as a reproducible floor rather than as this model's speed on LiteRT.

XNNPACK declines these fp16 graphs — it reports failed to delegate DEPTHWISE_CONV_2D and then fails to allocate tensors — so there is no usable CPU number. Disabling XNNPACK falls back to reference kernels, which measured about 20× slower than the GPU on models of this size and would not represent CPU inference anyone would ship.

Note that the GPU does not take the whole graph here (579 / 827 in ppocr_rec_fp16.tflite); the remainder runs on the CPU and the split costs a per-partition round trip.

Snapdragon NPU (Hexagon)

  • ppocr_det_fp16.tflite — the NPU is 2.04x faster than the GPU (8.63 ms against 17.59 ms) and loads 9.19x faster (139 ms against 1278 ms).
  • ppocr_rec_fp16.tflite — the NPU is 1.35x faster than the GPU (5.91 ms against 8.00 ms) and loads 10.70x faster (120 ms against 1280 ms).
  • ppocr_rec_fp32.tflite — the NPU is 1.26x faster than the GPU (6.33 ms against 7.98 ms) and loads 15.52x faster (129 ms against 2007 ms).
filebackendcompiledinference (median / min)load
ppocr_det_fp16.tfliteNPU (Hexagon v81)on-device JIT8.63 ms / 8.53 ms139 ms
ppocr_det_fp16.tfliteGPU (Adreno)—17.59 ms / 16.73 ms1278 ms
ppocr_rec_fp16.tfliteNPU (Hexagon v81)on-device JIT5.91 ms / 5.80 ms120 ms
ppocr_rec_fp16.tfliteGPU (Adreno)—8.00 ms / 7.75 ms1280 ms
ppocr_rec_fp32.tfliteNPU (Hexagon v81)on-device JIT6.33 ms / 6.17 ms129 ms
ppocr_rec_fp32.tfliteGPU (Adreno)—7.98 ms / 6.27 ms2007 ms

Measured on a Samsung Galaxy S26 (Snapdragon 8 Elite Gen 5 / SM8850, Hexagon v81, Android 16) with LiteRT CompiledModel 2.2.0, one accelerator per process, 5 warm-up runs then N=50 timed runs, median reported. Every run held thermal status NONE throughout. Headroom 0.77–0.81, where 1.0 is the throttling threshold.

The NPU rows ran the published file unchanged. LiteRT compiled it for the Hexagon on the device at first load. Those first compiles took 4.0 s to 28 s here. The load column above is the cached load every later run pays. Recipe and the runtime libraries it needs: NPU guide.

GPU wiring: GPU guide.

Raspberry Pi 5 (CPU)

Measured on a Raspberry Pi 5 Model B Rev 1.1 (8 GB, Raspberry Pi OS 64-bit) with the LiteRT benchmark_model tool from litert-cli-nightly 0.2.0.dev20260805: CPU inference (XNNPACK, 4 threads), 3 invocations per file of 10 warm-up plus 50 timed runs (the tool caps a phase at 150 s, so very slow graphs run fewer — the Runs column is the actual timed total). The latency is the median across invocations; the spread is the min–max over all timed runs. No thermal throttling occurred during these runs (vcgencmd get_throttled stayed 0x0).

FileInference (median)Spread (min–max)RunsPeak memory
ppocr_det_fp16.tflite460.2 ms458.2–492.4 ms150206 MB
ppocr_rec_fp16.tflite86.4 ms85.1–89.4 ms150135 MB
ppocr_rec_fp32.tflite84.7 ms83.8–86.8 ms150135 MB
litert
on-device
text-detection
text-recognition
tflite