GLiNER2.5 Multi for LiteRT — Android GPU FP32
0
2 commits
3 linked in READMEs
updated Oct 2, 2026
GLiNER2.5 Multi extracts person, organization, location, product and date entities from English and Japanese text. These files are the official fastino/gliner2.5-multi-v1 weights converted for LiteRT with LiteRT Torch. Tested on a Samsung Galaxy S26 (SM-S942Q, SM8850, Android 16) with LiteRT 2.2.0 and explicit GPU FP32 computation: at every shipped window, the spans match the official gliner2 fp32 implementation exactly on every bilingual test input that fits (70, 75 and 80 inputs; span F1 1.000). The phone's Hexagon NPU matched 69/70, 72/75 and 77/80 span sets, so it is reported below but not offered as a supported path.

The image renders the actual output of gliner25_multi_s128_wfp16.tflite with the
fp16 embedding table (LiteRT CompiledModel, desktop CPU) on the two sentences in
host_assets/example.json. The person, organization and product names are
invented.
Three window sizes are shipped. Use the smallest window whose N holds the encoded tokens (the five-label schema prompt plus the text) and whose T holds the processed text slots. Longer inputs are rejected, not truncated.
| File | Window N / text slots T | Packed output floats | Bytes | Role |
|---|---|---|---|---|
gliner25_multi_s128_wfp16.tflite | 128 / 48 | 78,110 | 191,876,560 | recommended |
gliner25_multi_s256_wfp16.tflite | 256 / 192 | 292,958 | 211,168,608 | recommended |
gliner25_multi_s512_wfp16.tflite | 512 / 384 | 579,422 | 250,248,384 | recommended |
fp32/gliner25_multi_s128_fp32.tflite | 128 / 48 | 78,110 | 363,985,160 | fp32 reference |
fp32/gliner25_multi_s256_fp32.tflite | 256 / 192 | 292,958 | 383,277,208 | fp32 reference |
fp32/gliner25_multi_s512_fp32.tflite | 512 / 384 | 579,422 | 422,356,980 | fp32 reference |
wfp16 files store the 96 FULLY_CONNECTED weight tensors as float16 and keep
everything else, including inputs and outputs, in float32. The fp32/ files are the
same graphs with float32 weights. Weight storage does not set the GPU compute
precision: request FP32 explicitly. At the default precision the outputs stay
finite, but one of 70 span sets changes.
Each graph is the dense part of the upstream BoundaryExtractor: the mDeBERTa-v3-base encoder, the boundary encoder, the boundary query head and the per-token projections. It takes word-embedding rows and returns one packed float32 tensor of 17 logical outputs. The host splits and tokenizes the text and looks up the rows before the graph, then runs the upstream sparse decoder after it.
The caller picks the word splitter per request: whitespace for space-delimited
text such as English, char for Japanese and other CJK text. There is no language
detector; mixed-language text needs an explicit choice too. Capacity counts
processed text slots and encoded tokens, not source characters. char gives each
non-space character its own slot and keeps runs of ASCII letters and digits
together. Text ending in 。 takes one extra slot, because the processor appends
. to text that does not end in ., ! or ?.
host_assets/ holds everything the host needs:
| File | Purpose |
|---|---|
word_embeddings_fp16.bin | recommended: [250112,768] float16, row-major, no header, 384,172,032 B; upcast selected rows to float32 |
word_embeddings_fp32.bin | reference: float32, 768,344,064 B, bit-identical to the checkpoint tensor |
sparse_decoder_fp32.safetensors, decoder_parameters.json | the 16 upstream sparse-decoder tensors (860,676 B) and their list |
tokenizer.json, tokenizer_config.json, config.json, encoder_config/config.json | exact files from the pinned checkpoint |
graph_contract_s{128,256,512}.json | input shapes and output-slice offsets |
runtime/ | Python host runtime: input construction, unpacking, upstream decoding |
example.json | both example sentences with encoded inputs and official spans |
Install requirements-lock.txt and run from the repository root. --seq auto (the
default) picks the smallest fitting window; --model fp32 and --table fp32 select
the references.
python examples/run_example.py --splitter whitespace \
--text "Mira Velsan presented the Lumenquill tablet for Asterfold Labs in Bristol on October 12."
The same steps by hand:
import json, os, sys
from pathlib import Path
import numpy as np
os.environ["HF_HUB_OFFLINE"] = "1"
sys.path.insert(0, str(Path("host_assets/runtime").resolve()))
from host_runtime import HostRuntime
from ai_edge_litert.compiled_model import CompiledModel, CpuOptions, HardwareAccelerator, Options
host = HostRuntime(Path("host_assets"), table="fp16")
text = "星瀬澄香は京都で霧灯研究社の新製品『星糸端末』を紹介した。"
inputs, captured = host.prepare(text, 128, "char")
model = CompiledModel.from_file("gliner25_multi_s128_wfp16.tflite", options=Options(
hardware_accelerators=HardwareAccelerator.CPU, cpu_options=CpuOptions(num_threads=4)))
sig = "serving_default"
ins = {f"args_{i}": model.create_input_buffer_by_name(sig, f"args_{i}") for i in range(5)}
outs = {"output_0": model.create_output_buffer_by_name(sig, "output_0")}
for i, x in enumerate(inputs):
ins[f"args_{i}"].write(np.ascontiguousarray(x, dtype=np.float32))
model.run_by_name(sig, ins, outs)
packed = outs["output_0"].read(78110, np.float32).reshape(1, 1, 1, 78110)
print(json.dumps(host.decode(captured, packed, inputs)["entities"], ensure_ascii=False))
# {"person": [{"text": "星瀬澄香", "confidence": 0.9864593744277954, "start": 0, "end": 4}], ...
Add com.google.ai.edge.litert:litert:2.2.0, stage the model file in app-private
storage and keep one Environment per process. For s128 the inputs are args_0
[1,128,768], args_1 [1,128], args_2 [1,48,128], args_3 [1,5,128] and
args_4 [1,48], filled as HOST_CONTRACT.md describes.
import com.google.ai.edge.litert.Accelerator
import com.google.ai.edge.litert.CompiledModel
import com.google.ai.edge.litert.Environment
import java.io.File
fun runDensePrefix(env: Environment, modelDir: File, inputs: List<FloatArray>): FloatArray {
val options = CompiledModel.Options(Accelerator.GPU).apply {
gpuOptions = CompiledModel.GpuOptions(precision = CompiledModel.GpuOptions.Precision.FP32)
}
val path = File(modelDir, "gliner25_multi_s128_wfp16.tflite").path
CompiledModel.create(path, options, env).use { model ->
val ins = (0..4).associate { "args_$it" to model.createInputBuffer("args_$it", "serving_default") }
val outs = mapOf("output_0" to model.createOutputBuffer("output_0", "serving_default"))
try {
inputs.forEachIndexed { i, x -> ins.getValue("args_$i").writeFloat(x) }
model.run(ins, outs, "serving_default")
return outs.getValue("output_0").readFloat()
} finally {
(ins.values + outs.values).forEach { it.close() }
}
}
}
// start_logits [1,5,49] begins at float 46976 (graph_contract_s128.json)
fun startLogits(packed: FloatArray): Array<FloatArray> =
Array(5) { q -> packed.copyOfRange(46976 + q * 49, 46976 + (q + 1) * 49) }
The host steps (splitter, tokenizer, embedding lookup, sparse decoder) ship in
Python only. An Android port can be checked against host_assets/example.json,
which holds the encoded ids, routing positions and official spans.
HOST_CONTRACT.md is the complete specification. In short:
gliner2 processor does: schema tokens
for the five labels, [SEP_TEXT], then the split text. Tokenizing the joined
string or adding BOS/CLS does not reproduce it. Right-pad the ids with 0 to N.args_0 holds the raw table rows for every position, padding included; the
graph applies the embedding LayerNorm. The other four inputs are 0/1 masks and
one-hot routing rows.graph_contract_s*.json) into label, text, start, end and confidence per span at
threshold 0.5. Offsets count Unicode code points: text[start:end] is the entity.The reference is the official gliner2 2.0.0 fp32 CPU implementation at revision
235cf92d, threshold 0.5, same splitter. The 80 test inputs (35 English, 45
Japanese) use the five-label schema; 70 / 75 / 80 fit the three windows. A span
counts only when label, start and end all match, so F1 1.000 means the files
reproduce the official output, not that every entity is right. Upstream misses carry
over: on ten Japanese check sentences the upstream model found 4 of 10 expected
organizations, and so do the converted files.
Desktop LiteRT CPU (ai-edge-litert 2.1.6 CompiledModel, macOS arm64, fp32 table):
| Window | Inputs | wfp16 F1 (max confidence drift) | fp32 F1 (max confidence drift) |
|---|---|---|---|
| 128 | 70 | 1.000 (1.3e-3) | 1.000 (1.2e-5) |
| 256 | 75 | 1.000 (1.3e-3) | 1.000 (1.2e-5) |
| 512 | 80 | 1.000 (1.3e-3) | 1.000 (1.2e-5) |
Every input returns the identical span set, also with the fp16 table (drift at most 1.4e-3).
On the Galaxy S26 (LiteRT 2.2.0 Kotlin CompiledModel API, wfp16 graphs unless noted), only the graph ran on the phone. Inputs were built with the fp16 table, and outputs were decoded on the desktop with the Python runtime. A pass needs every official span set and a confidence drift of at most 0.01. GPU and NPU each took the whole graph as one partition.
| Window | Accelerator, precision | Identical span sets | Max confidence drift | Delegated (logcat) | Load + compile | Opening ms | Sustained ms |
|---|---|---|---|---|---|---|---|
| 128 | GPU, explicit FP32 | 70/70 | 1.4e-3 | 1148/1148 nodes (LITERT_CL) | 1.5 s | 24.5 | 24.3 |
| 256 | GPU, same | 75/75 | 1.4e-3 | 1149/1149 | 1.4 s | 63.5 | 93.0 |
| 512 | GPU, same | 80/80 | 1.4e-3 | 1149/1149 | 1.7 s | 200.9 | 422.5 |
| 128 | NPU (Hexagon HTP, JIT, BURST), fp16 compute | 69/70 | 4.8e-2 | 1148/1148 ops, 1 node (DispatchDelegate) | 5.4 s | 8.3 | 8.3 |
| 256 | NPU, same | 72/75 | 4.8e-2 | 1149/1149 ops, 1 node | 45.7 s | 34.5 | 34.4 |
| 512 | NPU, same | 77/80 | 4.8e-2 | 1149/1149 ops, 1 node | 167.1 s | 186.9 | 186.5 |
| 128 | CPU (reference), XNNPACK, 4 threads | 70/70 | 1.4e-3 | 1147/1148 (XNNPACK) | 0.3 s | 32.7 | 49.5 |
| 256 | CPU, same | 75/75 | 1.4e-3 | 1148/1149 | 0.3 s | 74.6 | 128.5 |
| 512 | CPU, same | 80/80 | 1.4e-3 | 1148/1149 | 0.3 s | 210.1 | 407.3 |
| 128 | GPU, explicit FP32, fp32/ graph | 70/70 | 3.2e-4 | 1052/1052 | 1.5 s | 24.7 | 24.3 |
| 128 | GPU, default precision (fp16 compute), diagnostic only | 69/70 | 3.3e-2 | 1148/1148 | 1.2 s | 13.9 | 15.0 |
Every NPU difference sits at the 0.5 threshold: it dropped two spans the official model returns at 0.502 and 0.503 and added one at 0.519 that the official model keeps below 0.5. All NPU outputs were finite, but the drift exceeds 0.01, so the NPU is reported, not validated. The default-precision GPU row flipped one span at the threshold, which is why FP32 is required.
Each input ran 2 warm-ups and 5 timed repetitions of input write → run → output
readback. Opening is the first input's median after the phone cooled to thermal
status 0 (battery 34.1–37.6 °C). For NPU s256 and s512, the first compile had warmed
it to status 2 (39.3 °C and 41.3 °C). Sustained is the median over all inputs of the
back-to-back job, during which the phone reached at most status 2 and 43.6 °C. Each
load is a first load in a fresh app process; for the NPU it includes the on-device
compile. Latency is a single-device sample, not a benchmark.
whitespace) and Japanese (char) only.fastino/gliner2.5-multi-v1, revision
235cf92d6d4318da9bfca0d08975c8fa7250d13b (BoundaryExtractor, encoder
microsoft/mdeberta-v3-base), loaded with gliner2 2.0.0 and transformers
4.57.6.litert-torch 0.9.3 (torch 2.12.1), fixed shapes. The dense prefix
was re-expressed for the GPU delegate without changing its math (matmul routing,
float masks, rank-4 attention, baked relative-position buckets).ai-edge-quantizer 0.8.0 FLOAT_CASTING on FULLY_CONNECTED
weights only (96 tensors).ai-edge-litert 2.1.6; complete pins are in
requirements-lock.txt.License: the checkpoint is Apache-2.0, and these converted files are released under
the same license (LICENSE). The mDeBERTa-v3 encoder is MIT
(licenses/DeBERTa-MIT.txt). The gliner2 and transformers code used by the host
runtime is Apache-2.0 (licenses/), with notices in NOTICE. Upstream: the
model, the
GLiNER2 repository and the paper
arXiv:2507.18546. Runtime:
LiteRT.
GLiNER2.5 Multi for LiteRT — Android GPU FP32
0
2 commits
3 linked in READMEs
updated Oct 2, 2026
GLiNER2.5 Multi extracts person, organization, location, product and date entities from English and Japanese text. These files are the official fastino/gliner2.5-multi-v1 weights converted for LiteRT with LiteRT Torch. Tested on a Samsung Galaxy S26 (SM-S942Q, SM8850, Android 16) with LiteRT 2.2.0 and explicit GPU FP32 computation: at every shipped window, the spans match the official gliner2 fp32 implementation exactly on every bilingual test input that fits (70, 75 and 80 inputs; span F1 1.000). The phone's Hexagon NPU matched 69/70, 72/75 and 77/80 span sets, so it is reported below but not offered as a supported path.

The image renders the actual output of gliner25_multi_s128_wfp16.tflite with the
fp16 embedding table (LiteRT CompiledModel, desktop CPU) on the two sentences in
host_assets/example.json. The person, organization and product names are
invented.
Three window sizes are shipped. Use the smallest window whose N holds the encoded tokens (the five-label schema prompt plus the text) and whose T holds the processed text slots. Longer inputs are rejected, not truncated.
| File | Window N / text slots T | Packed output floats | Bytes | Role |
|---|---|---|---|---|
gliner25_multi_s128_wfp16.tflite | 128 / 48 | 78,110 | 191,876,560 | recommended |
gliner25_multi_s256_wfp16.tflite | 256 / 192 | 292,958 | 211,168,608 | recommended |
gliner25_multi_s512_wfp16.tflite | 512 / 384 | 579,422 | 250,248,384 | recommended |
fp32/gliner25_multi_s128_fp32.tflite | 128 / 48 | 78,110 | 363,985,160 | fp32 reference |
fp32/gliner25_multi_s256_fp32.tflite | 256 / 192 | 292,958 | 383,277,208 | fp32 reference |
fp32/gliner25_multi_s512_fp32.tflite | 512 / 384 | 579,422 | 422,356,980 | fp32 reference |
wfp16 files store the 96 FULLY_CONNECTED weight tensors as float16 and keep
everything else, including inputs and outputs, in float32. The fp32/ files are the
same graphs with float32 weights. Weight storage does not set the GPU compute
precision: request FP32 explicitly. At the default precision the outputs stay
finite, but one of 70 span sets changes.
Each graph is the dense part of the upstream BoundaryExtractor: the mDeBERTa-v3-base encoder, the boundary encoder, the boundary query head and the per-token projections. It takes word-embedding rows and returns one packed float32 tensor of 17 logical outputs. The host splits and tokenizes the text and looks up the rows before the graph, then runs the upstream sparse decoder after it.
The caller picks the word splitter per request: whitespace for space-delimited
text such as English, char for Japanese and other CJK text. There is no language
detector; mixed-language text needs an explicit choice too. Capacity counts
processed text slots and encoded tokens, not source characters. char gives each
non-space character its own slot and keeps runs of ASCII letters and digits
together. Text ending in 。 takes one extra slot, because the processor appends
. to text that does not end in ., ! or ?.
host_assets/ holds everything the host needs:
| File | Purpose |
|---|---|
word_embeddings_fp16.bin | recommended: [250112,768] float16, row-major, no header, 384,172,032 B; upcast selected rows to float32 |
word_embeddings_fp32.bin | reference: float32, 768,344,064 B, bit-identical to the checkpoint tensor |
sparse_decoder_fp32.safetensors, decoder_parameters.json | the 16 upstream sparse-decoder tensors (860,676 B) and their list |
tokenizer.json, tokenizer_config.json, config.json, encoder_config/config.json | exact files from the pinned checkpoint |
graph_contract_s{128,256,512}.json | input shapes and output-slice offsets |
runtime/ | Python host runtime: input construction, unpacking, upstream decoding |
example.json | both example sentences with encoded inputs and official spans |
Install requirements-lock.txt and run from the repository root. --seq auto (the
default) picks the smallest fitting window; --model fp32 and --table fp32 select
the references.
python examples/run_example.py --splitter whitespace \
--text "Mira Velsan presented the Lumenquill tablet for Asterfold Labs in Bristol on October 12."
The same steps by hand:
import json, os, sys
from pathlib import Path
import numpy as np
os.environ["HF_HUB_OFFLINE"] = "1"
sys.path.insert(0, str(Path("host_assets/runtime").resolve()))
from host_runtime import HostRuntime
from ai_edge_litert.compiled_model import CompiledModel, CpuOptions, HardwareAccelerator, Options
host = HostRuntime(Path("host_assets"), table="fp16")
text = "星瀬澄香は京都で霧灯研究社の新製品『星糸端末』を紹介した。"
inputs, captured = host.prepare(text, 128, "char")
model = CompiledModel.from_file("gliner25_multi_s128_wfp16.tflite", options=Options(
hardware_accelerators=HardwareAccelerator.CPU, cpu_options=CpuOptions(num_threads=4)))
sig = "serving_default"
ins = {f"args_{i}": model.create_input_buffer_by_name(sig, f"args_{i}") for i in range(5)}
outs = {"output_0": model.create_output_buffer_by_name(sig, "output_0")}
for i, x in enumerate(inputs):
ins[f"args_{i}"].write(np.ascontiguousarray(x, dtype=np.float32))
model.run_by_name(sig, ins, outs)
packed = outs["output_0"].read(78110, np.float32).reshape(1, 1, 1, 78110)
print(json.dumps(host.decode(captured, packed, inputs)["entities"], ensure_ascii=False))
# {"person": [{"text": "星瀬澄香", "confidence": 0.9864593744277954, "start": 0, "end": 4}], ...
Add com.google.ai.edge.litert:litert:2.2.0, stage the model file in app-private
storage and keep one Environment per process. For s128 the inputs are args_0
[1,128,768], args_1 [1,128], args_2 [1,48,128], args_3 [1,5,128] and
args_4 [1,48], filled as HOST_CONTRACT.md describes.
import com.google.ai.edge.litert.Accelerator
import com.google.ai.edge.litert.CompiledModel
import com.google.ai.edge.litert.Environment
import java.io.File
fun runDensePrefix(env: Environment, modelDir: File, inputs: List<FloatArray>): FloatArray {
val options = CompiledModel.Options(Accelerator.GPU).apply {
gpuOptions = CompiledModel.GpuOptions(precision = CompiledModel.GpuOptions.Precision.FP32)
}
val path = File(modelDir, "gliner25_multi_s128_wfp16.tflite").path
CompiledModel.create(path, options, env).use { model ->
val ins = (0..4).associate { "args_$it" to model.createInputBuffer("args_$it", "serving_default") }
val outs = mapOf("output_0" to model.createOutputBuffer("output_0", "serving_default"))
try {
inputs.forEachIndexed { i, x -> ins.getValue("args_$i").writeFloat(x) }
model.run(ins, outs, "serving_default")
return outs.getValue("output_0").readFloat()
} finally {
(ins.values + outs.values).forEach { it.close() }
}
}
}
// start_logits [1,5,49] begins at float 46976 (graph_contract_s128.json)
fun startLogits(packed: FloatArray): Array<FloatArray> =
Array(5) { q -> packed.copyOfRange(46976 + q * 49, 46976 + (q + 1) * 49) }
The host steps (splitter, tokenizer, embedding lookup, sparse decoder) ship in
Python only. An Android port can be checked against host_assets/example.json,
which holds the encoded ids, routing positions and official spans.
HOST_CONTRACT.md is the complete specification. In short:
gliner2 processor does: schema tokens
for the five labels, [SEP_TEXT], then the split text. Tokenizing the joined
string or adding BOS/CLS does not reproduce it. Right-pad the ids with 0 to N.args_0 holds the raw table rows for every position, padding included; the
graph applies the embedding LayerNorm. The other four inputs are 0/1 masks and
one-hot routing rows.graph_contract_s*.json) into label, text, start, end and confidence per span at
threshold 0.5. Offsets count Unicode code points: text[start:end] is the entity.The reference is the official gliner2 2.0.0 fp32 CPU implementation at revision
235cf92d, threshold 0.5, same splitter. The 80 test inputs (35 English, 45
Japanese) use the five-label schema; 70 / 75 / 80 fit the three windows. A span
counts only when label, start and end all match, so F1 1.000 means the files
reproduce the official output, not that every entity is right. Upstream misses carry
over: on ten Japanese check sentences the upstream model found 4 of 10 expected
organizations, and so do the converted files.
Desktop LiteRT CPU (ai-edge-litert 2.1.6 CompiledModel, macOS arm64, fp32 table):
| Window | Inputs | wfp16 F1 (max confidence drift) | fp32 F1 (max confidence drift) |
|---|---|---|---|
| 128 | 70 | 1.000 (1.3e-3) | 1.000 (1.2e-5) |
| 256 | 75 | 1.000 (1.3e-3) | 1.000 (1.2e-5) |
| 512 | 80 | 1.000 (1.3e-3) | 1.000 (1.2e-5) |
Every input returns the identical span set, also with the fp16 table (drift at most 1.4e-3).
On the Galaxy S26 (LiteRT 2.2.0 Kotlin CompiledModel API, wfp16 graphs unless noted), only the graph ran on the phone. Inputs were built with the fp16 table, and outputs were decoded on the desktop with the Python runtime. A pass needs every official span set and a confidence drift of at most 0.01. GPU and NPU each took the whole graph as one partition.
| Window | Accelerator, precision | Identical span sets | Max confidence drift | Delegated (logcat) | Load + compile | Opening ms | Sustained ms |
|---|---|---|---|---|---|---|---|
| 128 | GPU, explicit FP32 | 70/70 | 1.4e-3 | 1148/1148 nodes (LITERT_CL) | 1.5 s | 24.5 | 24.3 |
| 256 | GPU, same | 75/75 | 1.4e-3 | 1149/1149 | 1.4 s | 63.5 | 93.0 |
| 512 | GPU, same | 80/80 | 1.4e-3 | 1149/1149 | 1.7 s | 200.9 | 422.5 |
| 128 | NPU (Hexagon HTP, JIT, BURST), fp16 compute | 69/70 | 4.8e-2 | 1148/1148 ops, 1 node (DispatchDelegate) | 5.4 s | 8.3 | 8.3 |
| 256 | NPU, same | 72/75 | 4.8e-2 | 1149/1149 ops, 1 node | 45.7 s | 34.5 | 34.4 |
| 512 | NPU, same | 77/80 | 4.8e-2 | 1149/1149 ops, 1 node | 167.1 s | 186.9 | 186.5 |
| 128 | CPU (reference), XNNPACK, 4 threads | 70/70 | 1.4e-3 | 1147/1148 (XNNPACK) | 0.3 s | 32.7 | 49.5 |
| 256 | CPU, same | 75/75 | 1.4e-3 | 1148/1149 | 0.3 s | 74.6 | 128.5 |
| 512 | CPU, same | 80/80 | 1.4e-3 | 1148/1149 | 0.3 s | 210.1 | 407.3 |
| 128 | GPU, explicit FP32, fp32/ graph | 70/70 | 3.2e-4 | 1052/1052 | 1.5 s | 24.7 | 24.3 |
| 128 | GPU, default precision (fp16 compute), diagnostic only | 69/70 | 3.3e-2 | 1148/1148 | 1.2 s | 13.9 | 15.0 |
Every NPU difference sits at the 0.5 threshold: it dropped two spans the official model returns at 0.502 and 0.503 and added one at 0.519 that the official model keeps below 0.5. All NPU outputs were finite, but the drift exceeds 0.01, so the NPU is reported, not validated. The default-precision GPU row flipped one span at the threshold, which is why FP32 is required.
Each input ran 2 warm-ups and 5 timed repetitions of input write → run → output
readback. Opening is the first input's median after the phone cooled to thermal
status 0 (battery 34.1–37.6 °C). For NPU s256 and s512, the first compile had warmed
it to status 2 (39.3 °C and 41.3 °C). Sustained is the median over all inputs of the
back-to-back job, during which the phone reached at most status 2 and 43.6 °C. Each
load is a first load in a fresh app process; for the NPU it includes the on-device
compile. Latency is a single-device sample, not a benchmark.
whitespace) and Japanese (char) only.fastino/gliner2.5-multi-v1, revision
235cf92d6d4318da9bfca0d08975c8fa7250d13b (BoundaryExtractor, encoder
microsoft/mdeberta-v3-base), loaded with gliner2 2.0.0 and transformers
4.57.6.litert-torch 0.9.3 (torch 2.12.1), fixed shapes. The dense prefix
was re-expressed for the GPU delegate without changing its math (matmul routing,
float masks, rank-4 attention, baked relative-position buckets).ai-edge-quantizer 0.8.0 FLOAT_CASTING on FULLY_CONNECTED
weights only (96 tensors).ai-edge-litert 2.1.6; complete pins are in
requirements-lock.txt.License: the checkpoint is Apache-2.0, and these converted files are released under
the same license (LICENSE). The mDeBERTa-v3 encoder is MIT
(licenses/DeBERTa-MIT.txt). The gliner2 and transformers code used by the host
runtime is Apache-2.0 (licenses/), with notices in NOTICE. Upstream: the
model, the
GLiNER2 repository and the paper
arXiv:2507.18546. Runtime:
LiteRT.