litert-community/GLiNER2.5-Multi-LiteRT

Model

GLiNER2.5 Multi for LiteRT — Android GPU FP32

0

2 commits

3 linked in READMEs

updated Oct 2, 2026

See the code

README

GLiNER2.5 Multi for LiteRT — Android GPU FP32

GLiNER2.5 Multi extracts person, organization, location, product and date entities from English and Japanese text. These files are the official fastino/gliner2.5-multi-v1 weights converted for LiteRT with LiteRT Torch. Tested on a Samsung Galaxy S26 (SM-S942Q, SM8850, Android 16) with LiteRT 2.2.0 and explicit GPU FP32 computation: at every shipped window, the spans match the official gliner2 fp32 implementation exactly on every bilingual test input that fits (70, 75 and 80 inputs; span F1 1.000). The phone's Hexagon NPU matched 69/70, 72/75 and 77/80 span sets, so it is reported below but not offered as a supported path.

Entity spans on the two example sentences

The image renders the actual output of gliner25_multi_s128_wfp16.tflite with the fp16 embedding table (LiteRT CompiledModel, desktop CPU) on the two sentences in host_assets/example.json. The person, organization and product names are invented.

Files and supported configuration

Three window sizes are shipped. Use the smallest window whose N holds the encoded tokens (the five-label schema prompt plus the text) and whose T holds the processed text slots. Longer inputs are rejected, not truncated.

FileWindow N / text slots TPacked output floatsBytesRole
gliner25_multi_s128_wfp16.tflite128 / 4878,110191,876,560recommended
gliner25_multi_s256_wfp16.tflite256 / 192292,958211,168,608recommended
gliner25_multi_s512_wfp16.tflite512 / 384579,422250,248,384recommended
fp32/gliner25_multi_s128_fp32.tflite128 / 4878,110363,985,160fp32 reference
fp32/gliner25_multi_s256_fp32.tflite256 / 192292,958383,277,208fp32 reference
fp32/gliner25_multi_s512_fp32.tflite512 / 384579,422422,356,980fp32 reference

wfp16 files store the 96 FULLY_CONNECTED weight tensors as float16 and keep everything else, including inputs and outputs, in float32. The fp32/ files are the same graphs with float32 weights. Weight storage does not set the GPU compute precision: request FP32 explicitly. At the default precision the outputs stay finite, but one of 70 span sets changes.

Each graph is the dense part of the upstream BoundaryExtractor: the mDeBERTa-v3-base encoder, the boundary encoder, the boundary query head and the per-token projections. It takes word-embedding rows and returns one packed float32 tensor of 17 logical outputs. The host splits and tokenizes the text and looks up the rows before the graph, then runs the upstream sparse decoder after it.

The caller picks the word splitter per request: whitespace for space-delimited text such as English, char for Japanese and other CJK text. There is no language detector; mixed-language text needs an explicit choice too. Capacity counts processed text slots and encoded tokens, not source characters. char gives each non-space character its own slot and keeps runs of ASCII letters and digits together. Text ending in 。 takes one extra slot, because the processor appends . to text that does not end in ., ! or ?.

host_assets/ holds everything the host needs:

FilePurpose
word_embeddings_fp16.binrecommended: [250112,768] float16, row-major, no header, 384,172,032 B; upcast selected rows to float32
word_embeddings_fp32.binreference: float32, 768,344,064 B, bit-identical to the checkpoint tensor
sparse_decoder_fp32.safetensors, decoder_parameters.jsonthe 16 upstream sparse-decoder tensors (860,676 B) and their list
tokenizer.json, tokenizer_config.json, config.json, encoder_config/config.jsonexact files from the pinned checkpoint
graph_contract_s{128,256,512}.jsoninput shapes and output-slice offsets
runtime/Python host runtime: input construction, unpacking, upstream decoding
example.jsonboth example sentences with encoded inputs and official spans

Minimal usage

Python — complete pipeline, desktop CPU

Install requirements-lock.txt and run from the repository root. --seq auto (the default) picks the smallest fitting window; --model fp32 and --table fp32 select the references.

python examples/run_example.py --splitter whitespace \
  --text "Mira Velsan presented the Lumenquill tablet for Asterfold Labs in Bristol on October 12."

The same steps by hand:

import json, os, sys
from pathlib import Path
import numpy as np
os.environ["HF_HUB_OFFLINE"] = "1"
sys.path.insert(0, str(Path("host_assets/runtime").resolve()))
from host_runtime import HostRuntime
from ai_edge_litert.compiled_model import CompiledModel, CpuOptions, HardwareAccelerator, Options

host = HostRuntime(Path("host_assets"), table="fp16")
text = "星瀬澄香は京都で霧灯研究社の新製品『星糸端末』を紹介した。"
inputs, captured = host.prepare(text, 128, "char")

model = CompiledModel.from_file("gliner25_multi_s128_wfp16.tflite", options=Options(
    hardware_accelerators=HardwareAccelerator.CPU, cpu_options=CpuOptions(num_threads=4)))
sig = "serving_default"
ins = {f"args_{i}": model.create_input_buffer_by_name(sig, f"args_{i}") for i in range(5)}
outs = {"output_0": model.create_output_buffer_by_name(sig, "output_0")}
for i, x in enumerate(inputs):
    ins[f"args_{i}"].write(np.ascontiguousarray(x, dtype=np.float32))
model.run_by_name(sig, ins, outs)
packed = outs["output_0"].read(78110, np.float32).reshape(1, 1, 1, 78110)
print(json.dumps(host.decode(captured, packed, inputs)["entities"], ensure_ascii=False))
# {"person": [{"text": "星瀬澄香", "confidence": 0.9864593744277954, "start": 0, "end": 4}], ...

Kotlin — Android GPU with explicit FP32

Add com.google.ai.edge.litert:litert:2.2.0, stage the model file in app-private storage and keep one Environment per process. For s128 the inputs are args_0 [1,128,768], args_1 [1,128], args_2 [1,48,128], args_3 [1,5,128] and args_4 [1,48], filled as HOST_CONTRACT.md describes.

import com.google.ai.edge.litert.Accelerator
import com.google.ai.edge.litert.CompiledModel
import com.google.ai.edge.litert.Environment
import java.io.File

fun runDensePrefix(env: Environment, modelDir: File, inputs: List<FloatArray>): FloatArray {
  val options = CompiledModel.Options(Accelerator.GPU).apply {
    gpuOptions = CompiledModel.GpuOptions(precision = CompiledModel.GpuOptions.Precision.FP32)
  }
  val path = File(modelDir, "gliner25_multi_s128_wfp16.tflite").path
  CompiledModel.create(path, options, env).use { model ->
    val ins = (0..4).associate { "args_$it" to model.createInputBuffer("args_$it", "serving_default") }
    val outs = mapOf("output_0" to model.createOutputBuffer("output_0", "serving_default"))
    try {
      inputs.forEachIndexed { i, x -> ins.getValue("args_$i").writeFloat(x) }
      model.run(ins, outs, "serving_default")
      return outs.getValue("output_0").readFloat()
    } finally {
      (ins.values + outs.values).forEach { it.close() }
    }
  }
}

// start_logits [1,5,49] begins at float 46976 (graph_contract_s128.json)
fun startLogits(packed: FloatArray): Array<FloatArray> =
  Array(5) { q -> packed.copyOfRange(46976 + q * 49, 46976 + (q + 1) * 49) }

The host steps (splitter, tokenizer, embedding lookup, sparse decoder) ship in Python only. An Android port can be checked against host_assets/example.json, which holds the encoded ids, routing positions and official spans.

Host contract

HOST_CONTRACT.md is the complete specification. In short:

  1. Build the sequence the way the upstream gliner2 processor does: schema tokens for the five labels, [SEP_TEXT], then the split text. Tokenizing the joined string or adding BOS/CLS does not reproduce it. Right-pad the ids with 0 to N.
  2. args_0 holds the raw table rows for every position, padding included; the graph applies the embedding LayerNorm. The other four inputs are 0/1 masks and one-hot routing rows.
  3. The upstream sparse decoder turns the 17 output slices (offsets in graph_contract_s*.json) into label, text, start, end and confidence per span at threshold 0.5. Offsets count Unicode code points: text[start:end] is the entity.

Measured quality and performance

The reference is the official gliner2 2.0.0 fp32 CPU implementation at revision 235cf92d, threshold 0.5, same splitter. The 80 test inputs (35 English, 45 Japanese) use the five-label schema; 70 / 75 / 80 fit the three windows. A span counts only when label, start and end all match, so F1 1.000 means the files reproduce the official output, not that every entity is right. Upstream misses carry over: on ten Japanese check sentences the upstream model found 4 of 10 expected organizations, and so do the converted files.

Desktop LiteRT CPU (ai-edge-litert 2.1.6 CompiledModel, macOS arm64, fp32 table):

WindowInputswfp16 F1 (max confidence drift)fp32 F1 (max confidence drift)
128701.000 (1.3e-3)1.000 (1.2e-5)
256751.000 (1.3e-3)1.000 (1.2e-5)
512801.000 (1.3e-3)1.000 (1.2e-5)

Every input returns the identical span set, also with the fp16 table (drift at most 1.4e-3).

On the Galaxy S26 (LiteRT 2.2.0 Kotlin CompiledModel API, wfp16 graphs unless noted), only the graph ran on the phone. Inputs were built with the fp16 table, and outputs were decoded on the desktop with the Python runtime. A pass needs every official span set and a confidence drift of at most 0.01. GPU and NPU each took the whole graph as one partition.

WindowAccelerator, precisionIdentical span setsMax confidence driftDelegated (logcat)Load + compileOpening msSustained ms
128GPU, explicit FP3270/701.4e-31148/1148 nodes (LITERT_CL)1.5 s24.524.3
256GPU, same75/751.4e-31149/11491.4 s63.593.0
512GPU, same80/801.4e-31149/11491.7 s200.9422.5
128NPU (Hexagon HTP, JIT, BURST), fp16 compute69/704.8e-21148/1148 ops, 1 node (DispatchDelegate)5.4 s8.38.3
256NPU, same72/754.8e-21149/1149 ops, 1 node45.7 s34.534.4
512NPU, same77/804.8e-21149/1149 ops, 1 node167.1 s186.9186.5
128CPU (reference), XNNPACK, 4 threads70/701.4e-31147/1148 (XNNPACK)0.3 s32.749.5
256CPU, same75/751.4e-31148/11490.3 s74.6128.5
512CPU, same80/801.4e-31148/11490.3 s210.1407.3
128GPU, explicit FP32, fp32/ graph70/703.2e-41052/10521.5 s24.724.3
128GPU, default precision (fp16 compute), diagnostic only69/703.3e-21148/11481.2 s13.915.0

Every NPU difference sits at the 0.5 threshold: it dropped two spans the official model returns at 0.502 and 0.503 and added one at 0.519 that the official model keeps below 0.5. All NPU outputs were finite, but the drift exceeds 0.01, so the NPU is reported, not validated. The default-precision GPU row flipped one span at the threshold, which is why FP32 is required.

Each input ran 2 warm-ups and 5 timed repetitions of input write → run → output readback. Opening is the first input's median after the phone cooled to thermal status 0 (battery 34.1–37.6 °C). For NPU s256 and s512, the first compile had warmed it to status 2 (39.3 °C and 41.3 °C). Sustained is the median over all inputs of the back-to-back job, during which the phone reached at most status 2 and 43.6 °C. Each load is a first load in a fresh app process; for the NPU it includes the on-device compile. Latency is a single-device sample, not a benchmark.

Limits

  • Schema: only the five labels above, in that order, at threshold 0.5. Other labels, relations, classifications and record schemas were not validated.
  • Languages: English (whitespace) and Japanese (char) only.
  • Length: one window per call. Splitting longer text was not validated.
  • Devices: tested on the Galaxy S26 only.
  • NPU: measured above, not validated.

Provenance, conversion and license

  • Source: fastino/gliner2.5-multi-v1, revision 235cf92d6d4318da9bfca0d08975c8fa7250d13b (BoundaryExtractor, encoder microsoft/mdeberta-v3-base), loaded with gliner2 2.0.0 and transformers 4.57.6.
  • Conversion: litert-torch 0.9.3 (torch 2.12.1), fixed shapes. The dense prefix was re-expressed for the GPU delegate without changing its math (matmul routing, float masks, rank-4 attention, baked relative-position buckets).
  • Two fp32-exact rewrites keep the same files finite under fp16 delegates: the attention-mask fill is -10000 instead of the float32 minimum, and each of the 29 active LayerNorms is computed on x·2⁻³ with eps·2⁻⁶. On 20 inputs per window the rewritten torch graph matched the original bit for bit (60/60 comparisons, maximum difference 0.0).
  • Weight storage: ai-edge-quantizer 0.8.0 FLOAT_CASTING on FULLY_CONNECTED weights only (96 tensors).
  • Desktop checks used ai-edge-litert 2.1.6; complete pins are in requirements-lock.txt.

License: the checkpoint is Apache-2.0, and these converted files are released under the same license (LICENSE). The mDeBERTa-v3 encoder is MIT (licenses/DeBERTa-MIT.txt). The gliner2 and transformers code used by the host runtime is Apache-2.0 (licenses/), with notices in NOTICE. Upstream: the model, the GLiNER2 repository and the paper arXiv:2507.18546. Runtime: LiteRT.

android
gliner2
information-extraction
japanese
litert
mdeberta-v3
multilingual
named-entity-recognition
tflite
token-classification

litert-community/GLiNER2.5-Multi-LiteRT

Model

GLiNER2.5 Multi for LiteRT — Android GPU FP32

0

2 commits

3 linked in READMEs

updated Oct 2, 2026

See the code

README

GLiNER2.5 Multi for LiteRT — Android GPU FP32

GLiNER2.5 Multi extracts person, organization, location, product and date entities from English and Japanese text. These files are the official fastino/gliner2.5-multi-v1 weights converted for LiteRT with LiteRT Torch. Tested on a Samsung Galaxy S26 (SM-S942Q, SM8850, Android 16) with LiteRT 2.2.0 and explicit GPU FP32 computation: at every shipped window, the spans match the official gliner2 fp32 implementation exactly on every bilingual test input that fits (70, 75 and 80 inputs; span F1 1.000). The phone's Hexagon NPU matched 69/70, 72/75 and 77/80 span sets, so it is reported below but not offered as a supported path.

Entity spans on the two example sentences

The image renders the actual output of gliner25_multi_s128_wfp16.tflite with the fp16 embedding table (LiteRT CompiledModel, desktop CPU) on the two sentences in host_assets/example.json. The person, organization and product names are invented.

Files and supported configuration

Three window sizes are shipped. Use the smallest window whose N holds the encoded tokens (the five-label schema prompt plus the text) and whose T holds the processed text slots. Longer inputs are rejected, not truncated.

FileWindow N / text slots TPacked output floatsBytesRole
gliner25_multi_s128_wfp16.tflite128 / 4878,110191,876,560recommended
gliner25_multi_s256_wfp16.tflite256 / 192292,958211,168,608recommended
gliner25_multi_s512_wfp16.tflite512 / 384579,422250,248,384recommended
fp32/gliner25_multi_s128_fp32.tflite128 / 4878,110363,985,160fp32 reference
fp32/gliner25_multi_s256_fp32.tflite256 / 192292,958383,277,208fp32 reference
fp32/gliner25_multi_s512_fp32.tflite512 / 384579,422422,356,980fp32 reference

wfp16 files store the 96 FULLY_CONNECTED weight tensors as float16 and keep everything else, including inputs and outputs, in float32. The fp32/ files are the same graphs with float32 weights. Weight storage does not set the GPU compute precision: request FP32 explicitly. At the default precision the outputs stay finite, but one of 70 span sets changes.

Each graph is the dense part of the upstream BoundaryExtractor: the mDeBERTa-v3-base encoder, the boundary encoder, the boundary query head and the per-token projections. It takes word-embedding rows and returns one packed float32 tensor of 17 logical outputs. The host splits and tokenizes the text and looks up the rows before the graph, then runs the upstream sparse decoder after it.

The caller picks the word splitter per request: whitespace for space-delimited text such as English, char for Japanese and other CJK text. There is no language detector; mixed-language text needs an explicit choice too. Capacity counts processed text slots and encoded tokens, not source characters. char gives each non-space character its own slot and keeps runs of ASCII letters and digits together. Text ending in 。 takes one extra slot, because the processor appends . to text that does not end in ., ! or ?.

host_assets/ holds everything the host needs:

FilePurpose
word_embeddings_fp16.binrecommended: [250112,768] float16, row-major, no header, 384,172,032 B; upcast selected rows to float32
word_embeddings_fp32.binreference: float32, 768,344,064 B, bit-identical to the checkpoint tensor
sparse_decoder_fp32.safetensors, decoder_parameters.jsonthe 16 upstream sparse-decoder tensors (860,676 B) and their list
tokenizer.json, tokenizer_config.json, config.json, encoder_config/config.jsonexact files from the pinned checkpoint
graph_contract_s{128,256,512}.jsoninput shapes and output-slice offsets
runtime/Python host runtime: input construction, unpacking, upstream decoding
example.jsonboth example sentences with encoded inputs and official spans

Minimal usage

Python — complete pipeline, desktop CPU

Install requirements-lock.txt and run from the repository root. --seq auto (the default) picks the smallest fitting window; --model fp32 and --table fp32 select the references.

python examples/run_example.py --splitter whitespace \
  --text "Mira Velsan presented the Lumenquill tablet for Asterfold Labs in Bristol on October 12."

The same steps by hand:

import json, os, sys
from pathlib import Path
import numpy as np
os.environ["HF_HUB_OFFLINE"] = "1"
sys.path.insert(0, str(Path("host_assets/runtime").resolve()))
from host_runtime import HostRuntime
from ai_edge_litert.compiled_model import CompiledModel, CpuOptions, HardwareAccelerator, Options

host = HostRuntime(Path("host_assets"), table="fp16")
text = "星瀬澄香は京都で霧灯研究社の新製品『星糸端末』を紹介した。"
inputs, captured = host.prepare(text, 128, "char")

model = CompiledModel.from_file("gliner25_multi_s128_wfp16.tflite", options=Options(
    hardware_accelerators=HardwareAccelerator.CPU, cpu_options=CpuOptions(num_threads=4)))
sig = "serving_default"
ins = {f"args_{i}": model.create_input_buffer_by_name(sig, f"args_{i}") for i in range(5)}
outs = {"output_0": model.create_output_buffer_by_name(sig, "output_0")}
for i, x in enumerate(inputs):
    ins[f"args_{i}"].write(np.ascontiguousarray(x, dtype=np.float32))
model.run_by_name(sig, ins, outs)
packed = outs["output_0"].read(78110, np.float32).reshape(1, 1, 1, 78110)
print(json.dumps(host.decode(captured, packed, inputs)["entities"], ensure_ascii=False))
# {"person": [{"text": "星瀬澄香", "confidence": 0.9864593744277954, "start": 0, "end": 4}], ...

Kotlin — Android GPU with explicit FP32

Add com.google.ai.edge.litert:litert:2.2.0, stage the model file in app-private storage and keep one Environment per process. For s128 the inputs are args_0 [1,128,768], args_1 [1,128], args_2 [1,48,128], args_3 [1,5,128] and args_4 [1,48], filled as HOST_CONTRACT.md describes.

import com.google.ai.edge.litert.Accelerator
import com.google.ai.edge.litert.CompiledModel
import com.google.ai.edge.litert.Environment
import java.io.File

fun runDensePrefix(env: Environment, modelDir: File, inputs: List<FloatArray>): FloatArray {
  val options = CompiledModel.Options(Accelerator.GPU).apply {
    gpuOptions = CompiledModel.GpuOptions(precision = CompiledModel.GpuOptions.Precision.FP32)
  }
  val path = File(modelDir, "gliner25_multi_s128_wfp16.tflite").path
  CompiledModel.create(path, options, env).use { model ->
    val ins = (0..4).associate { "args_$it" to model.createInputBuffer("args_$it", "serving_default") }
    val outs = mapOf("output_0" to model.createOutputBuffer("output_0", "serving_default"))
    try {
      inputs.forEachIndexed { i, x -> ins.getValue("args_$i").writeFloat(x) }
      model.run(ins, outs, "serving_default")
      return outs.getValue("output_0").readFloat()
    } finally {
      (ins.values + outs.values).forEach { it.close() }
    }
  }
}

// start_logits [1,5,49] begins at float 46976 (graph_contract_s128.json)
fun startLogits(packed: FloatArray): Array<FloatArray> =
  Array(5) { q -> packed.copyOfRange(46976 + q * 49, 46976 + (q + 1) * 49) }

The host steps (splitter, tokenizer, embedding lookup, sparse decoder) ship in Python only. An Android port can be checked against host_assets/example.json, which holds the encoded ids, routing positions and official spans.

Host contract

HOST_CONTRACT.md is the complete specification. In short:

  1. Build the sequence the way the upstream gliner2 processor does: schema tokens for the five labels, [SEP_TEXT], then the split text. Tokenizing the joined string or adding BOS/CLS does not reproduce it. Right-pad the ids with 0 to N.
  2. args_0 holds the raw table rows for every position, padding included; the graph applies the embedding LayerNorm. The other four inputs are 0/1 masks and one-hot routing rows.
  3. The upstream sparse decoder turns the 17 output slices (offsets in graph_contract_s*.json) into label, text, start, end and confidence per span at threshold 0.5. Offsets count Unicode code points: text[start:end] is the entity.

Measured quality and performance

The reference is the official gliner2 2.0.0 fp32 CPU implementation at revision 235cf92d, threshold 0.5, same splitter. The 80 test inputs (35 English, 45 Japanese) use the five-label schema; 70 / 75 / 80 fit the three windows. A span counts only when label, start and end all match, so F1 1.000 means the files reproduce the official output, not that every entity is right. Upstream misses carry over: on ten Japanese check sentences the upstream model found 4 of 10 expected organizations, and so do the converted files.

Desktop LiteRT CPU (ai-edge-litert 2.1.6 CompiledModel, macOS arm64, fp32 table):

WindowInputswfp16 F1 (max confidence drift)fp32 F1 (max confidence drift)
128701.000 (1.3e-3)1.000 (1.2e-5)
256751.000 (1.3e-3)1.000 (1.2e-5)
512801.000 (1.3e-3)1.000 (1.2e-5)

Every input returns the identical span set, also with the fp16 table (drift at most 1.4e-3).

On the Galaxy S26 (LiteRT 2.2.0 Kotlin CompiledModel API, wfp16 graphs unless noted), only the graph ran on the phone. Inputs were built with the fp16 table, and outputs were decoded on the desktop with the Python runtime. A pass needs every official span set and a confidence drift of at most 0.01. GPU and NPU each took the whole graph as one partition.

WindowAccelerator, precisionIdentical span setsMax confidence driftDelegated (logcat)Load + compileOpening msSustained ms
128GPU, explicit FP3270/701.4e-31148/1148 nodes (LITERT_CL)1.5 s24.524.3
256GPU, same75/751.4e-31149/11491.4 s63.593.0
512GPU, same80/801.4e-31149/11491.7 s200.9422.5
128NPU (Hexagon HTP, JIT, BURST), fp16 compute69/704.8e-21148/1148 ops, 1 node (DispatchDelegate)5.4 s8.38.3
256NPU, same72/754.8e-21149/1149 ops, 1 node45.7 s34.534.4
512NPU, same77/804.8e-21149/1149 ops, 1 node167.1 s186.9186.5
128CPU (reference), XNNPACK, 4 threads70/701.4e-31147/1148 (XNNPACK)0.3 s32.749.5
256CPU, same75/751.4e-31148/11490.3 s74.6128.5
512CPU, same80/801.4e-31148/11490.3 s210.1407.3
128GPU, explicit FP32, fp32/ graph70/703.2e-41052/10521.5 s24.724.3
128GPU, default precision (fp16 compute), diagnostic only69/703.3e-21148/11481.2 s13.915.0

Every NPU difference sits at the 0.5 threshold: it dropped two spans the official model returns at 0.502 and 0.503 and added one at 0.519 that the official model keeps below 0.5. All NPU outputs were finite, but the drift exceeds 0.01, so the NPU is reported, not validated. The default-precision GPU row flipped one span at the threshold, which is why FP32 is required.

Each input ran 2 warm-ups and 5 timed repetitions of input write → run → output readback. Opening is the first input's median after the phone cooled to thermal status 0 (battery 34.1–37.6 °C). For NPU s256 and s512, the first compile had warmed it to status 2 (39.3 °C and 41.3 °C). Sustained is the median over all inputs of the back-to-back job, during which the phone reached at most status 2 and 43.6 °C. Each load is a first load in a fresh app process; for the NPU it includes the on-device compile. Latency is a single-device sample, not a benchmark.

Limits

  • Schema: only the five labels above, in that order, at threshold 0.5. Other labels, relations, classifications and record schemas were not validated.
  • Languages: English (whitespace) and Japanese (char) only.
  • Length: one window per call. Splitting longer text was not validated.
  • Devices: tested on the Galaxy S26 only.
  • NPU: measured above, not validated.

Provenance, conversion and license

  • Source: fastino/gliner2.5-multi-v1, revision 235cf92d6d4318da9bfca0d08975c8fa7250d13b (BoundaryExtractor, encoder microsoft/mdeberta-v3-base), loaded with gliner2 2.0.0 and transformers 4.57.6.
  • Conversion: litert-torch 0.9.3 (torch 2.12.1), fixed shapes. The dense prefix was re-expressed for the GPU delegate without changing its math (matmul routing, float masks, rank-4 attention, baked relative-position buckets).
  • Two fp32-exact rewrites keep the same files finite under fp16 delegates: the attention-mask fill is -10000 instead of the float32 minimum, and each of the 29 active LayerNorms is computed on x·2⁻³ with eps·2⁻⁶. On 20 inputs per window the rewritten torch graph matched the original bit for bit (60/60 comparisons, maximum difference 0.0).
  • Weight storage: ai-edge-quantizer 0.8.0 FLOAT_CASTING on FULLY_CONNECTED weights only (96 tensors).
  • Desktop checks used ai-edge-litert 2.1.6; complete pins are in requirements-lock.txt.

License: the checkpoint is Apache-2.0, and these converted files are released under the same license (LICENSE). The mDeBERTa-v3 encoder is MIT (licenses/DeBERTa-MIT.txt). The gliner2 and transformers code used by the host runtime is Apache-2.0 (licenses/), with notices in NOTICE. Upstream: the model, the GLiNER2 repository and the paper arXiv:2507.18546. Runtime: LiteRT.

android
gliner2
information-extraction
japanese
litert
mdeberta-v3
multilingual
named-entity-recognition
tflite
token-classification