Measured on device (edge-compat): browser · Chromium 151 on M4 Max · LiteRT.js 2.5.3 · WebGPU · 285 ms p50 · output differs from CPU (max rel diff 2.3e+10) (2026-08-20). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/tipsv2-b14-dpt/CARD.md
2
16 commits
3 linked in READMEs
updated Sep 8, 2026
Measured on device (edge-compat): browser · Chromium 151 on M4 Max · LiteRT.js 2.5.3 · WebGPU · 285 ms p50 · output differs from CPU (max rel diff 2.3e+10) (2026-08-20). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/tipsv2-b14-dpt/CARD.md
TIPSv2 (Google DeepMind, CVPR 2026) is a
DINOv2-style ViT-B/14 vision-language backbone; google/tipsv2-b14-dpt
adds three DPT heads trained on the frozen backbone. This is that model as one LiteRT
graph that runs fully on the CompiledModel GPU delegate (no CPU fallback) and returns
all three dense outputs from a single 448×448 image:

Left to right: input, metric depth (inverse-depth Spectral colormap), surface normals (xyz → RGB), ADE20K segmentation — all three from the on-device fp16 GPU run (Pixel 8a).
[1, 3, 448, 448] NCHW, RGB in [0, 1] — no ImageNet mean/std (TIPSv2 convention).[1, 1, 448, 448] depth in metres[1, 3, 448, 448] unit surface normals[1, 150, 256, 256] segmentation logits at the DPT head's native grid —
argmax over the 150 channels host-side, then upscale (the official pipeline bilinearly
upsamples the logits to 448 first; argmax agreement between the two is 99.96 %).Fully GPU-resident on a Pixel 8a (1434/1434 nodes, 1 partition, ~0.9 s for all three heads). Re-authored from the HF weights with exact rewrites:
[1, heads, N, d] with a manual
softmax(qkᵀ/√d)·v — the delegate rejects the native 5D head-split.attn.proj / mlp.fc2; SafeLayerNorm (deviation pre-scaled by
1/64 before squaring so the fp16 variance cannot overflow); tanh-GELU (no ERF kernel —
the only non-exact rewrite, corr 0.999998 vs the official model).cat(patch, cls.expand) @ W → patch @ W_a + cls @ W_b (exact, no
BROADCAST_TO); ConvTranspose2d(k=s) → zero-stuff + Conv2d (exact; TRANSPOSE_CONV is
rejected on Pixel 8a); the fusion blocks' bilinear ×2 align_corners=True (banned on the GPU)
→ two constant-RHS matmuls U·X·Uᵀ (exact).relu(l)/Σrelu(l)), so
power-of-2 scales are folded into its weights/biases (per-level convs 1/4096…1/64, each fusion
out_conv ×1/4, project 1/32, head 1/16 — matching scales at every residual add) keeping
every stage ≲100. Bit-exact in fp32.Parity (desktop fp32 re-authored vs official PyTorch): depth corr 0.999998 · normals 0.999999 ·
seg argmax 99.96 %. Device fp16 (Pixel 8a GPU) vs desktop fp32: depth corr 0.99986 · normals
0.99990 · seg argmax 99.3 %. Conversion script: build_tipsv2.py (litert-torch).
// 318 MB — stage into filesDir (adb push + run-as cp) rather than APK assets
val model = CompiledModel.create(File(context.filesDir, "tipsv2_b14_dpt_fp16.tflite").absolutePath,
CompiledModel.Options(Accelerator.GPU), null)
val inputs = model.createInputBuffers()
val outputs = model.createOutputBuffers()
inputs[0].writeFloat(imageNchw01) // [1,3,448,448] RGB / 255, no mean/std
model.run(inputs, outputs)
val depth = outputs[0].readFloat() // [448*448] metres
val normals = outputs[1].readFloat() // [3*448*448] unit vectors
val seg = outputs[2].readFloat() // [150*256*256] logits -> argmax per pixel host-side
import numpy as np
from ai_edge_litert.compiled_model import CompiledModel
model = CompiledModel.from_file("tipsv2_b14_dpt_fp16.tflite")
inputs = model.create_input_buffers(0)
outputs = model.create_output_buffers(0)
inputs[0].write(np.ascontiguousarray(image, np.float32)) # [1,3,448,448] in [0,1]
model.run_by_index(0, inputs, outputs)
depth = outputs[0].read(448 * 448, np.float32).reshape(448, 448) # metres
normals = outputs[1].read(3 * 448 * 448, np.float32).reshape(3, 448, 448)
seg = outputs[2].read(150 * 256 * 256, np.float32).reshape(150, 256, 256)
labels = seg.argmax(0) # ADE20K ids (256x256)
Measured on a Pixel 8a (Tensor G3 / Mali-G715) through LiteRT's own CompiledModel
accelerator (LITERT_CL), the path the Kotlin sample and the LiteRT API use.
| Runtime | Backend | Graph on GPU | Latency |
|---|---|---|---|
LiteRT CompiledModel (LITERT_CL) | GPU | 1434 / 1434 | ~0.92 s / image (all three heads; wall-clock difference between 1 and 6 back-to-back runs) |
GPU compile + load ≈ 5 s on first use. The three heads are ~3/4 of the compute (256-channel 3×3 convs up to 256×256); a single-head build would be correspondingly faster.
The TIPSv2 Android sample (Kotlin, CompiledModel GPU): photo picker → input | depth,
normals | segmentation with an ADE20K legend. The conversion script build_tipsv2.py
(re-authoring + parity + litert-torch convert + fp16 + device harness) is included in this repo.
The NPU is faster — once the GELU is written as an op. The original file loses on
the NPU (326.9 ms vs 282.5 ms on the GPU), and the cause is the tanh-GELU elementwise
decomposition (24 sites: ViT blocks + DPT readout), which falls off the Hexagon
compiler's fast path. tipsv2_b14_dpt_erf_fp16.tflite is the same model with the
GELU emitted as the builtin GELU op (exact erf); all three outputs match the original
file at corr 0.99998+ (depth 0.999984, normals 0.999999, seg 0.999993):
| file | NPU (Hexagon v81) | GPU (Adreno) |
|---|---|---|
tipsv2_b14_dpt_fp16.tflite (tanh-GELU) | 326.9 ms | 282.5 ms |
tipsv2_b14_dpt_erf_fp16.tflite (GELU op) | 142.0 ms | 273.3 ms |
Pick by accelerator: the erf file for the NPU (2.30x). It also ran at full speed on
this Adreno GPU; the Mali path is unverified with the builtin GELU op — keep the
original file for Mali. The erf build is build_tipsv2.py with the gelu function
set to torch.nn.functional.gelu.
Measured on a Samsung Galaxy S26 (Snapdragon 8 Elite Gen 5 / SM8850, Hexagon v81,
Android 16), LiteRT CompiledModel 2.2.0, on-device JIT compile (first load
compiles in ~2-4 min and caches; later loads ~0.4-0.8 s), one accelerator per process,
warm-up then N=50 timed runs, median reported, every row at thermal status NONE —
which also settles the direction this section previously called provisional.
Latencies were taken on the fp32 build of the same graph; fp16 weight storage measured
within noise on the models where the fold test was run (DINOv2, zipformer).
The runtime libraries the NPU needs are in the NPU recipe; GPU wiring is in the GPU recipe.
Measured on a Raspberry Pi 5 Model B Rev 1.1 (8 GB, Raspberry Pi OS 64-bit) with the LiteRT benchmark_model tool from litert-cli-nightly 0.2.0.dev20260805: CPU inference (XNNPACK, 4 threads), 3 invocations per file of 10 warm-up plus 50 timed runs (the tool caps a phase at 150 s, so very slow graphs run fewer — the Runs column is the actual timed total). The latency is the median across invocations; the spread is the min–max over all timed runs. No thermal throttling occurred during these runs (vcgencmd get_throttled stayed 0x0).
| File | Inference (median) | Spread (min–max) | Runs | Peak memory |
|---|---|---|---|---|
tipsv2_b14_dpt_erf_fp16.tflite | 10,642.7 ms | 10,582.8–10,762.4 ms | 45 | 1291 MB |
tipsv2_b14_dpt_fp16.tflite | 11,185.4 ms | 11,096.1–11,293.2 ms | 42 | 1292 MB |
Apache-2.0 (TIPSv2 / Google DeepMind). Converted with litert-torch. Hero photo: Pexels (free license).
Measured on device (edge-compat): browser · Chromium 151 on M4 Max · LiteRT.js 2.5.3 · WebGPU · 285 ms p50 · output differs from CPU (max rel diff 2.3e+10) (2026-08-20). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/tipsv2-b14-dpt/CARD.md
2
16 commits
3 linked in READMEs
updated Sep 8, 2026
Measured on device (edge-compat): browser · Chromium 151 on M4 Max · LiteRT.js 2.5.3 · WebGPU · 285 ms p50 · output differs from CPU (max rel diff 2.3e+10) (2026-08-20). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/tipsv2-b14-dpt/CARD.md
TIPSv2 (Google DeepMind, CVPR 2026) is a
DINOv2-style ViT-B/14 vision-language backbone; google/tipsv2-b14-dpt
adds three DPT heads trained on the frozen backbone. This is that model as one LiteRT
graph that runs fully on the CompiledModel GPU delegate (no CPU fallback) and returns
all three dense outputs from a single 448×448 image:

Left to right: input, metric depth (inverse-depth Spectral colormap), surface normals (xyz → RGB), ADE20K segmentation — all three from the on-device fp16 GPU run (Pixel 8a).
[1, 3, 448, 448] NCHW, RGB in [0, 1] — no ImageNet mean/std (TIPSv2 convention).[1, 1, 448, 448] depth in metres[1, 3, 448, 448] unit surface normals[1, 150, 256, 256] segmentation logits at the DPT head's native grid —
argmax over the 150 channels host-side, then upscale (the official pipeline bilinearly
upsamples the logits to 448 first; argmax agreement between the two is 99.96 %).Fully GPU-resident on a Pixel 8a (1434/1434 nodes, 1 partition, ~0.9 s for all three heads). Re-authored from the HF weights with exact rewrites:
[1, heads, N, d] with a manual
softmax(qkᵀ/√d)·v — the delegate rejects the native 5D head-split.attn.proj / mlp.fc2; SafeLayerNorm (deviation pre-scaled by
1/64 before squaring so the fp16 variance cannot overflow); tanh-GELU (no ERF kernel —
the only non-exact rewrite, corr 0.999998 vs the official model).cat(patch, cls.expand) @ W → patch @ W_a + cls @ W_b (exact, no
BROADCAST_TO); ConvTranspose2d(k=s) → zero-stuff + Conv2d (exact; TRANSPOSE_CONV is
rejected on Pixel 8a); the fusion blocks' bilinear ×2 align_corners=True (banned on the GPU)
→ two constant-RHS matmuls U·X·Uᵀ (exact).relu(l)/Σrelu(l)), so
power-of-2 scales are folded into its weights/biases (per-level convs 1/4096…1/64, each fusion
out_conv ×1/4, project 1/32, head 1/16 — matching scales at every residual add) keeping
every stage ≲100. Bit-exact in fp32.Parity (desktop fp32 re-authored vs official PyTorch): depth corr 0.999998 · normals 0.999999 ·
seg argmax 99.96 %. Device fp16 (Pixel 8a GPU) vs desktop fp32: depth corr 0.99986 · normals
0.99990 · seg argmax 99.3 %. Conversion script: build_tipsv2.py (litert-torch).
// 318 MB — stage into filesDir (adb push + run-as cp) rather than APK assets
val model = CompiledModel.create(File(context.filesDir, "tipsv2_b14_dpt_fp16.tflite").absolutePath,
CompiledModel.Options(Accelerator.GPU), null)
val inputs = model.createInputBuffers()
val outputs = model.createOutputBuffers()
inputs[0].writeFloat(imageNchw01) // [1,3,448,448] RGB / 255, no mean/std
model.run(inputs, outputs)
val depth = outputs[0].readFloat() // [448*448] metres
val normals = outputs[1].readFloat() // [3*448*448] unit vectors
val seg = outputs[2].readFloat() // [150*256*256] logits -> argmax per pixel host-side
import numpy as np
from ai_edge_litert.compiled_model import CompiledModel
model = CompiledModel.from_file("tipsv2_b14_dpt_fp16.tflite")
inputs = model.create_input_buffers(0)
outputs = model.create_output_buffers(0)
inputs[0].write(np.ascontiguousarray(image, np.float32)) # [1,3,448,448] in [0,1]
model.run_by_index(0, inputs, outputs)
depth = outputs[0].read(448 * 448, np.float32).reshape(448, 448) # metres
normals = outputs[1].read(3 * 448 * 448, np.float32).reshape(3, 448, 448)
seg = outputs[2].read(150 * 256 * 256, np.float32).reshape(150, 256, 256)
labels = seg.argmax(0) # ADE20K ids (256x256)
Measured on a Pixel 8a (Tensor G3 / Mali-G715) through LiteRT's own CompiledModel
accelerator (LITERT_CL), the path the Kotlin sample and the LiteRT API use.
| Runtime | Backend | Graph on GPU | Latency |
|---|---|---|---|
LiteRT CompiledModel (LITERT_CL) | GPU | 1434 / 1434 | ~0.92 s / image (all three heads; wall-clock difference between 1 and 6 back-to-back runs) |
GPU compile + load ≈ 5 s on first use. The three heads are ~3/4 of the compute (256-channel 3×3 convs up to 256×256); a single-head build would be correspondingly faster.
The TIPSv2 Android sample (Kotlin, CompiledModel GPU): photo picker → input | depth,
normals | segmentation with an ADE20K legend. The conversion script build_tipsv2.py
(re-authoring + parity + litert-torch convert + fp16 + device harness) is included in this repo.
The NPU is faster — once the GELU is written as an op. The original file loses on
the NPU (326.9 ms vs 282.5 ms on the GPU), and the cause is the tanh-GELU elementwise
decomposition (24 sites: ViT blocks + DPT readout), which falls off the Hexagon
compiler's fast path. tipsv2_b14_dpt_erf_fp16.tflite is the same model with the
GELU emitted as the builtin GELU op (exact erf); all three outputs match the original
file at corr 0.99998+ (depth 0.999984, normals 0.999999, seg 0.999993):
| file | NPU (Hexagon v81) | GPU (Adreno) |
|---|---|---|
tipsv2_b14_dpt_fp16.tflite (tanh-GELU) | 326.9 ms | 282.5 ms |
tipsv2_b14_dpt_erf_fp16.tflite (GELU op) | 142.0 ms | 273.3 ms |
Pick by accelerator: the erf file for the NPU (2.30x). It also ran at full speed on
this Adreno GPU; the Mali path is unverified with the builtin GELU op — keep the
original file for Mali. The erf build is build_tipsv2.py with the gelu function
set to torch.nn.functional.gelu.
Measured on a Samsung Galaxy S26 (Snapdragon 8 Elite Gen 5 / SM8850, Hexagon v81,
Android 16), LiteRT CompiledModel 2.2.0, on-device JIT compile (first load
compiles in ~2-4 min and caches; later loads ~0.4-0.8 s), one accelerator per process,
warm-up then N=50 timed runs, median reported, every row at thermal status NONE —
which also settles the direction this section previously called provisional.
Latencies were taken on the fp32 build of the same graph; fp16 weight storage measured
within noise on the models where the fold test was run (DINOv2, zipformer).
The runtime libraries the NPU needs are in the NPU recipe; GPU wiring is in the GPU recipe.
Measured on a Raspberry Pi 5 Model B Rev 1.1 (8 GB, Raspberry Pi OS 64-bit) with the LiteRT benchmark_model tool from litert-cli-nightly 0.2.0.dev20260805: CPU inference (XNNPACK, 4 threads), 3 invocations per file of 10 warm-up plus 50 timed runs (the tool caps a phase at 150 s, so very slow graphs run fewer — the Runs column is the actual timed total). The latency is the median across invocations; the spread is the min–max over all timed runs. No thermal throttling occurred during these runs (vcgencmd get_throttled stayed 0x0).
| File | Inference (median) | Spread (min–max) | Runs | Peak memory |
|---|---|---|---|---|
tipsv2_b14_dpt_erf_fp16.tflite | 10,642.7 ms | 10,582.8–10,762.4 ms | 45 | 1291 MB |
tipsv2_b14_dpt_fp16.tflite | 11,185.4 ms | 11,096.1–11,293.2 ms | 42 | 1292 MB |
Apache-2.0 (TIPSv2 / Google DeepMind). Converted with litert-torch. Hero photo: Pexels (free license).