litert-community/GLiNER2.5-Decide-LiteRT

Model

GLiNER2.5-Decide for LiteRT — Android GPU FP32

11

6 commits

3 linked in READMEs

updated Oct 4, 2026

See the code

README

GLiNER2.5-Decide for LiteRT — Android GPU FP32

Run GLiNER2.5-Decide on a phone GPU: pass a text and any set of labels at call time (intent, routing, sentiment, yes/no gates, multi-label tags), get one decision per task from a single forward pass. The files are the official fastino/GLiNER2.5-Decide weights converted with LiteRT Torch. The validated Android configuration is LiteRT 2.2.0 with explicit GPU FP32 computation on a Samsung Galaxy S26 (SM-S942Q, SM8850, Android 16). In the Android sample on that phone, the decisions equal the official gliner2 fp32 result on all 126 tested (request, window) pairs; on desktop LiteRT CPU they equal it on all 361 test requests. Other Android GPU families have not been validated here.

Decisions and label probabilities for the example request

The image shows the request in host_assets/example.json (an invented name) and the probability of every label. The values are model output: the s128 graph returned them in the Android sample on the Galaxy S26 GPU (FP32), and examples/run_example.py prints the same values to the shown precision on desktop CPU with the s256 graph. It is a rendering of model output, not a screenshot.

Files

Three encoded-token windows are shipped. Use the smallest window that holds the request: the task schemas plus the text must fit in N encoded tokens, with at most 32 labels in total.

FileWindow NBytesRole
gliner25_decide_s128_wfp16.tflite128660,383,872default
gliner25_decide_s256_wfp16.tflite256710,715,520default
gliner25_decide_s512_wfp16.tflite512811,378,816default
gliner25_decide_s128_npu_wfp16.tflite128660,302,784Qualcomm NPU (see NPU)
fp32/gliner25_decide_s128_fp32.tflite1281,268,514,036float32 reference
fp32/gliner25_decide_s256_fp32.tflite2561,318,845,684float32 reference
fp32/gliner25_decide_s512_fp32.tflite5121,419,508,980float32 reference

wfp16 files store the 146 FULLY_CONNECTED weight tensors as float16 with a DEQUANTIZE to float32 (1,780 operators); activations and every other constant stay float32. The fp32 files are the same graphs with float32 weights (1,634 operators). The three default graphs, the float16 embedding table and tokenizer.json are 2,452,978,688 bytes together. The npu file is the s128 graph changed for float16 hardware (see NPU): 170 float16 FULLY_CONNECTED weights and 1,874 operators, with the inputs and outputs of graph_contract_s128.json.

Each graph is the DeBERTa-v3-large encoder and the classification head of the checkpoint. The host looks up the word embeddings before the graph and turns the logits into decisions after it.

File in host_assets/Purpose
word_embeddings_fp16.bin[128011,1024] float16 table, 262,166,528 B; default, upcast to float32 on lookup
word_embeddings_fp32.binthe same table in float32, 524,333,056 B; reference
tokenizer.json (8,333,952 B), tokenizer_config.json, special_tokens_map.json, config.json, encoder_config/config.jsonexact files of the pinned checkpoint
graph_contract_s{128,256,512}.jsonsignature names, shapes, bytes and sha256 of both storages
runtime/Python host: decide_inputs.py builds the inputs with gliner2 2.0.0, host_decide.py decides
example.jsonthe worked request with the official result and logits
Graph input / outputShapeContent
args_0 inputs_embeds[1,N,1024] float32table rows at the right-padded token ids (padding rows included)
args_1 attention_mask[1,N] float321 for encoded tokens, 0 for padding
args_2 label_routing[1,32,N] float32row j one-hot at the j-th [L] label marker; unused rows zero
output_0 logits[1,1,1,32] float32one logit per label slot, in request order

HOST_CONTRACT.md specifies the encoded sequence, the marker positions and the decision rules.

Minimal usage

Python — desktop CPU

Install requirements-lock.txt (Python 3.12) and run from the downloaded repository. The host runtime builds the inputs with the pinned gliner2 2.0.0 schema and tokenizer code, so no checkpoint download is needed. python examples/run_example.py runs the same request and compares it with the official result stored in example.json.

import sys
import numpy as np
from ai_edge_litert.compiled_model import (
    CompiledModel, CpuOptions, HardwareAccelerator, Options)

sys.path.insert(0, "host_assets/runtime")
from decide_inputs import DecideHost  # gliner2 2.0.0 schema + tokenizer
from host_decide import decide

host = DecideHost("host_assets")  # float16 table, upcast to float32 on lookup
text = ("Hi, this is Mirela Tovantis. The wireless headphones I ordered arrived on "
        "Tuesday with a cracked case, and the left earbud will not charge. Please "
        "send a replacement or refund me before the weekend.")
tasks = {
    "intent": ["refund_request", "replacement_request", "order_status",
               "technical_support", "other"],
    "urgency": ["low", "medium", "high"],
    "sentiment": ["positive", "neutral", "negative"],
    "topics": {"labels": ["shipping", "product_quality", "billing", "battery"],
               "multi_label": True, "cls_threshold": 0.4},
}
request = host.prepare_classification(text, tasks)  # smallest window: s128

model = CompiledModel.from_file(
    f"gliner25_decide_s{request.seq}_wfp16.tflite",
    options=Options(hardware_accelerators=HardwareAccelerator.CPU,
                    cpu_options=CpuOptions(num_threads=4)))
inputs, outputs = model.create_input_buffers(0), model.create_output_buffers(0)
for buffer, array in zip(inputs, request.ordered_inputs()):  # args_0..args_2
    buffer.write(array)
model.run_by_index(0, inputs, outputs)
logits = outputs[0].read(32, np.float32)
decisions, probabilities, _ = decide(logits, request.tasks)
print(decisions)
# {'intent': 'refund_request', 'urgency': 'high', 'sentiment': 'negative',
#  'topics': ['shipping', 'product_quality', 'battery']}

Kotlin — Android GPU with explicit FP32

Use implementation("com.google.ai.edge.litert:litert:2.2.0"), stage the model file in the app's private directory, and keep one Environment for the process. The three inputs are built as HOST_CONTRACT.md describes; the buffers follow the signature order args_0..args_2.

import com.google.ai.edge.litert.Accelerator
import com.google.ai.edge.litert.CompiledModel
import com.google.ai.edge.litert.Environment
import java.io.File
import kotlin.math.exp

/** Runs one request through the s128 graph on the GPU and returns the 32 logits. */
fun decideLogits(
  env: Environment,
  modelDir: File,
  embeds: FloatArray, // [1,128,1024], float16 table rows upcast to float32
  mask: FloatArray, // [1,128]
  routing: FloatArray, // [1,32,128]
): FloatArray {
  val options =
    CompiledModel.Options(Accelerator.GPU).apply {
      // Explicit FP32: the default GPU precision changed 28 of 42 decisions.
      gpuOptions = CompiledModel.GpuOptions(precision = CompiledModel.GpuOptions.Precision.FP32)
    }
  val path = File(modelDir, "gliner25_decide_s128_wfp16.tflite").path
  CompiledModel.create(path, options, env).use { model ->
    val inputs = model.createInputBuffers()
    val outputs = model.createOutputBuffers()
    try {
      inputs[0].writeFloat(embeds)
      inputs[1].writeFloat(mask)
      inputs[2].writeFloat(routing)
      model.run(inputs, outputs)
      return outputs[0].readFloat() // slots 0..K-1 hold the K labels in request order
    } finally {
      (inputs + outputs).forEach { it.close() }
    }
  }
}

/** Single-label task: softmax over its slice of the logits, then the lowest-index maximum. */
fun decideSingleLabel(logits: FloatArray, offset: Int, labels: List<String>): Pair<String, Double> {
  val part = logits.copyOfRange(offset, offset + labels.size)
  val max = part.max()
  val weights = part.map { exp((it - max).toDouble()) }
  val best = weights.indices.maxBy { weights[it] }
  return labels[best] to weights[best] / weights.sum()
}

The complete Kotlin host is in android/: word splitter, SentencePiece Unigram tokenizer, gliner2 schema and input construction, the memory-mapped float16 table, and a float32 port of the decision rules including multi-label thresholds (DecideDecoder.kt), inside a Jetpack Compose app. Type a text and one task per line, tap Classify, and read each task's decision with its probability. Build and install steps are in android/README.md.

GPU precision

Weight storage and GPU computation precision are separate settings. With the runtime's default GPU precision, the s128 graph (float32 weights) compiled and returned finite logits on the Galaxy S26, but only 14 of 42 decisions matched the official result. With GpuOptions(precision = FP32) all 42 matched. Use the explicit option above for every file.

NPU

gliner25_decide_s128_npu_wfp16.tflite runs on the Qualcomm NPU of the Galaxy S26 (Hexagon v81 HTP) and gives the official decision on all 42 phone requests. The default s128 file does not: on the same NPU it compiles and runs, but 13 of 42 decisions match.

The NPU computes in float16. Three changes keep the graph within float16's range and precision. Each is exact in float32: on the 42 requests, the changed graph's outputs equal the default graph's bit for bit, in torch float32 and on desktop LiteRT CPU.

  • LayerNorm pre-scale. 22 of the 49 LayerNorms (the feed-forward output norms of layers 0–21) compute on x·2⁻ᵏ with eps·4⁻ᵏ, k from 1 to 5. In float32 their input reaches |x| = 1,423, so the squared deviation reaches 2.0e6 in 15 of 24 layers, above the float16 maximum of 65,504.
  • Attention mask. Masked positions add −1e4 instead of float32's lowest value, which has no float16 representation.
  • Feed-forward split. The first FULLY_CONNECTED of each feed-forward block runs as two halves joined by CONCATENATION before the GELU. Without the split, the NPU compiled the block's FC → GELU → FC chain into a path that added 19 to 25 times the float16 rounding error per block, and one decision flipped.
Galaxy S26, LiteRT 2.2.0FileDecisions equal officialMax |Δprob|Median ms per request
NPUnpu42/420.005528.6
GPU FP32npu42/420.0002669.1
GPU FP16 with FP32 accumulationnpu42/420.002860.2
GPU FP16npu41/420.01640.9
NPUdefault s12813/420.999826.2

Each row is one process on 2026-10-04, started at thermal status 0. The time is the median over the 42 requests of embedding lookup, input write, run and readback, one pass each. The first load compiles the graph for the NPU on the phone: 32.4 s. A second process read LiteRT's compiled cache in 0.68 s and returned the same 42 decisions. On the GPU, FP16 with FP32 accumulation keeps all 42 decisions and is 13 % faster than FP32. Plain FP16 flips one near-tie, whose official margin is 0.0051.

The native runner of the Latency section (input write → run → readback) timed the npu file on 2026-09-30. The NPU took 21.9–22.0 ms after a reboot. Before the reboot, with the phone up for a week, it took 41.2 ms, and the GPU at FP32 took 70.3 ms. The cause of the slower NPU state was not established. These are one phone's numbers on two days, not a benchmark across devices.

The NPU needs Qualcomm's runtime inside the app: the dispatch library and JIT compiler plugin of the LiteRT 2.2.0 release, and the QAIRT 2.47 HTP libraries for Hexagon v81. In Kotlin, create the Environment with the app's native library directory as DispatchLibraryDir and CompilerPluginLibraryDir, and pass Accelerator.NPU with QualcommOptions(htpPerformanceMode = BURST); every NPU run above used BURST. The Android sample in android/ runs the GPU and CPU paths only.

Fidelity

Reference: the official gliner2 2.0.0 classify_text on the fp32 checkpoint (desktop CPU), revision 7ee5da4c. Test requests: the 21 classify_text examples of the model card and 340 rows of the public dev split of fastino/fast-decisions (20 per domain, 17 domains, chosen by encoded length so that every row fits s512). Each window is tested on the requests that fit it. A decision matches when every task's label (or label set) equals the official one. Correlation was never used as a gate.

ConfigurationRequestsDecisions equal officialMax |Δlogit|
Desktop LiteRT CPU, wfp16 graphs + float16 tables128 42 / s256 328 / s512 36142/42 · 328/328 · 361/3614.81e-3
Desktop LiteRT CPU, fp32 graph + float32 tables512 361361/3613.29e-5
Galaxy S26 GPU FP32, native CompiledModel runner, wfp16 + float16 table42 per window42/42 · 42/42 · 42/424.78e-3
Galaxy S26 GPU FP32, Android sample126 (request, window) pairs126/1264.78e-3
Galaxy S26 CPU (XNNPACK, 4 threads), Android sample126 pairs126/1264.82e-3

Desktop: ai-edge-litert 2.1.6 CompiledModel, macOS arm64, 4 threads. S26: LiteRT 2.2.0; every GPU graph ran fully delegated in one partition (1,780 of 1,780 operators) with no NaN. The 42 phone requests per window are the 21 model-card examples plus 21 dev-split rows (at s256 and s512, the longest rows that fit). In the Android sample the on-device token ids, marker positions and padding also equal the captured Python inputs for all 126 pairs.

The dev-split rows measure fidelity to the official model, not accuracy. On this length-selected subset the official model's decisions equal the dataset's gold labels on 164 of 340 rows (364 of 580 tasks). The publisher's benchmark (60.2 % exact match) uses held-out rows that are not in the public dataset; the two numbers are not comparable. Across the 1,700 public dev rows, the encoded length is ≤ 128 tokens for 92 rows, 129–256 for 1,410 and 257–512 for 198; none exceeds 512, and the largest label set has 28 labels.

Latency

Galaxy S26 (SM-S942Q, SM8850, Android 16), LiteRT 2.2.0, wfp16 graphs, float16 table inputs. Native CompiledModel runner, one process per job, 42 requests per window, 2 warm-up runs then 8 (s128) or 5 (s256, s512) timed runs per request. A timed run is input write → run → output readback. Every job started cool (GPU ≤ 50 °C, thermal status 0), screen on, USB powered.

WindowGPU FP32, opening request (ms)GPU FP32, back-to-back median / min (ms)Last request (ms)GPU clock limit at the end
s12866.369.6 / 65.9 (336 runs)69.91,200 MHz
s256174.9198.8 / 174.7 (210 runs)214.8902 MHz
s512575.2733.1 / 573.5 (210 runs)912.1646 MHz

Back-to-back runs heat the phone: the GPU clock limit steps down from 1,300 MHz and each request of s256 / s512 gets slower during the job. The opening-request column is the cool value; the median mixes it with the throttled end. For comparison, the CPU (XNNPACK, 4 threads) took 510.0 / 311.5 ms (median / min) at s256 and 1,666.0 / 863.1 ms at s512 under the same protocol; the s128 CPU value, 216.5 / 183.6 ms (fp32 graph), was measured right after a GPU job on a warm phone (thermal status 2).

The Android sample (debug build, same phone, app in the foreground, the example request at s128):

ScenarioEnd to end (ms)
One request right after a cold start (three starts; 5 warm-up passes of 0.38–0.39 s before Ready)77.9 · 72.4 · 75.5
20 requests, one every 2 s, GPU FP32 (median; min 80.5, max 99.3)93.0
20 requests, one every 2 s, CPU (median; min 167.7, max 182.8)171.9

End to end covers tokenize and embed, graph to readback and decode; the paced runs also include the hand-off to the LiteRT worker thread. In the paced GPU run the graph took a median 68.4 ms; the tokenize-and-embed phase rose from about 13 ms to about 28 ms after the tenth request, and the cause was not established. These are single-device numbers, not a benchmark across devices or thermal states.

Limits

  • English, as the source model. One text per call; no batching.
  • At most 512 encoded tokens (task schemas plus text) and 32 labels per request. Longer requests are rejected, never truncated; gliner2's long-text chunking (classify_text_long) is not ported.
  • GPU and NPU validation cover the Galaxy S26 (Adreno GPU, Hexagon v81 NPU) only. Only the s128 window has an NPU file; the default files give wrong decisions on the NPU.
  • No INT8 file is shipped. On GLiNER2.5 Small, the same family of graphs, dynamic-range INT8 FULLY_CONNECTED weights did not compile on the LiteRT 2.2.0 GPU; no INT8 variant of this model was built or tested.
  • Few-shot examples and the extraction tasks of gliner2 (entities, relations, structures) are not part of the graphs.

Prior art

The classification path of GLiNER2.5-Decide is also published as ONNX in onnx-community/GLiNER2.5-Decide-ONNX (Transformers.js, WebGPU) and nishparadox/gliner2.5-decide-onnx. A Core AI (Apple) conversion of the same path is published as coreai-community/GLiNER2.5-Decide-CoreAI.

Reproduce

conversion/ holds the scripts that produced every file here (graph definition, export with litert-torch 0.9.3, float16 weight storage with ai-edge-quantizer 0.8.0, embedding tables, the hero image) and the order they ran in; requirements-lock.txt pins the environment. The graph keeps the checkpoint's arithmetic: host-side embedding lookup, one-hot label routing as a matrix product, rank-4 attention, DeBERTa's logarithmic relative-position buckets baked as constants, a float attention mask, and the native LiteRT GELU. The NPU file comes from profile_ranges.py, npu_graph.py, export_npu.py and quantize_npu.py, in the order conversion/README.md gives.

License

Apache-2.0, the license of the checkpoint (LICENSE). The DeBERTa-v3 encoder is MIT (licenses/DeBERTa-MIT.txt); the gliner2 and transformers code the host runtime and the Android sample build on is Apache-2.0 (licenses/). NOTICE lists the sources and the changes. Upstream: the GLiNER2 repository and the paper arXiv:2507.18546.

Citation

@misc{zaratiana2025gliner2efficientmultitaskinformation,
      title={GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface},
      author={Urchade Zaratiana and Gil Pasternak and Oliver Boyd and George Hurn-Maloney and Ash Lewis},
      year={2025},
      eprint={2507.18546},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2507.18546},
}
android
deberta-v3
gliner2
litert
text-classification
tflite
zero-shot-classification

litert-community/GLiNER2.5-Decide-LiteRT

Model

GLiNER2.5-Decide for LiteRT — Android GPU FP32

11

6 commits

3 linked in READMEs

updated Oct 4, 2026

See the code

README

GLiNER2.5-Decide for LiteRT — Android GPU FP32

Run GLiNER2.5-Decide on a phone GPU: pass a text and any set of labels at call time (intent, routing, sentiment, yes/no gates, multi-label tags), get one decision per task from a single forward pass. The files are the official fastino/GLiNER2.5-Decide weights converted with LiteRT Torch. The validated Android configuration is LiteRT 2.2.0 with explicit GPU FP32 computation on a Samsung Galaxy S26 (SM-S942Q, SM8850, Android 16). In the Android sample on that phone, the decisions equal the official gliner2 fp32 result on all 126 tested (request, window) pairs; on desktop LiteRT CPU they equal it on all 361 test requests. Other Android GPU families have not been validated here.

Decisions and label probabilities for the example request

The image shows the request in host_assets/example.json (an invented name) and the probability of every label. The values are model output: the s128 graph returned them in the Android sample on the Galaxy S26 GPU (FP32), and examples/run_example.py prints the same values to the shown precision on desktop CPU with the s256 graph. It is a rendering of model output, not a screenshot.

Files

Three encoded-token windows are shipped. Use the smallest window that holds the request: the task schemas plus the text must fit in N encoded tokens, with at most 32 labels in total.

FileWindow NBytesRole
gliner25_decide_s128_wfp16.tflite128660,383,872default
gliner25_decide_s256_wfp16.tflite256710,715,520default
gliner25_decide_s512_wfp16.tflite512811,378,816default
gliner25_decide_s128_npu_wfp16.tflite128660,302,784Qualcomm NPU (see NPU)
fp32/gliner25_decide_s128_fp32.tflite1281,268,514,036float32 reference
fp32/gliner25_decide_s256_fp32.tflite2561,318,845,684float32 reference
fp32/gliner25_decide_s512_fp32.tflite5121,419,508,980float32 reference

wfp16 files store the 146 FULLY_CONNECTED weight tensors as float16 with a DEQUANTIZE to float32 (1,780 operators); activations and every other constant stay float32. The fp32 files are the same graphs with float32 weights (1,634 operators). The three default graphs, the float16 embedding table and tokenizer.json are 2,452,978,688 bytes together. The npu file is the s128 graph changed for float16 hardware (see NPU): 170 float16 FULLY_CONNECTED weights and 1,874 operators, with the inputs and outputs of graph_contract_s128.json.

Each graph is the DeBERTa-v3-large encoder and the classification head of the checkpoint. The host looks up the word embeddings before the graph and turns the logits into decisions after it.

File in host_assets/Purpose
word_embeddings_fp16.bin[128011,1024] float16 table, 262,166,528 B; default, upcast to float32 on lookup
word_embeddings_fp32.binthe same table in float32, 524,333,056 B; reference
tokenizer.json (8,333,952 B), tokenizer_config.json, special_tokens_map.json, config.json, encoder_config/config.jsonexact files of the pinned checkpoint
graph_contract_s{128,256,512}.jsonsignature names, shapes, bytes and sha256 of both storages
runtime/Python host: decide_inputs.py builds the inputs with gliner2 2.0.0, host_decide.py decides
example.jsonthe worked request with the official result and logits
Graph input / outputShapeContent
args_0 inputs_embeds[1,N,1024] float32table rows at the right-padded token ids (padding rows included)
args_1 attention_mask[1,N] float321 for encoded tokens, 0 for padding
args_2 label_routing[1,32,N] float32row j one-hot at the j-th [L] label marker; unused rows zero
output_0 logits[1,1,1,32] float32one logit per label slot, in request order

HOST_CONTRACT.md specifies the encoded sequence, the marker positions and the decision rules.

Minimal usage

Python — desktop CPU

Install requirements-lock.txt (Python 3.12) and run from the downloaded repository. The host runtime builds the inputs with the pinned gliner2 2.0.0 schema and tokenizer code, so no checkpoint download is needed. python examples/run_example.py runs the same request and compares it with the official result stored in example.json.

import sys
import numpy as np
from ai_edge_litert.compiled_model import (
    CompiledModel, CpuOptions, HardwareAccelerator, Options)

sys.path.insert(0, "host_assets/runtime")
from decide_inputs import DecideHost  # gliner2 2.0.0 schema + tokenizer
from host_decide import decide

host = DecideHost("host_assets")  # float16 table, upcast to float32 on lookup
text = ("Hi, this is Mirela Tovantis. The wireless headphones I ordered arrived on "
        "Tuesday with a cracked case, and the left earbud will not charge. Please "
        "send a replacement or refund me before the weekend.")
tasks = {
    "intent": ["refund_request", "replacement_request", "order_status",
               "technical_support", "other"],
    "urgency": ["low", "medium", "high"],
    "sentiment": ["positive", "neutral", "negative"],
    "topics": {"labels": ["shipping", "product_quality", "billing", "battery"],
               "multi_label": True, "cls_threshold": 0.4},
}
request = host.prepare_classification(text, tasks)  # smallest window: s128

model = CompiledModel.from_file(
    f"gliner25_decide_s{request.seq}_wfp16.tflite",
    options=Options(hardware_accelerators=HardwareAccelerator.CPU,
                    cpu_options=CpuOptions(num_threads=4)))
inputs, outputs = model.create_input_buffers(0), model.create_output_buffers(0)
for buffer, array in zip(inputs, request.ordered_inputs()):  # args_0..args_2
    buffer.write(array)
model.run_by_index(0, inputs, outputs)
logits = outputs[0].read(32, np.float32)
decisions, probabilities, _ = decide(logits, request.tasks)
print(decisions)
# {'intent': 'refund_request', 'urgency': 'high', 'sentiment': 'negative',
#  'topics': ['shipping', 'product_quality', 'battery']}

Kotlin — Android GPU with explicit FP32

Use implementation("com.google.ai.edge.litert:litert:2.2.0"), stage the model file in the app's private directory, and keep one Environment for the process. The three inputs are built as HOST_CONTRACT.md describes; the buffers follow the signature order args_0..args_2.

import com.google.ai.edge.litert.Accelerator
import com.google.ai.edge.litert.CompiledModel
import com.google.ai.edge.litert.Environment
import java.io.File
import kotlin.math.exp

/** Runs one request through the s128 graph on the GPU and returns the 32 logits. */
fun decideLogits(
  env: Environment,
  modelDir: File,
  embeds: FloatArray, // [1,128,1024], float16 table rows upcast to float32
  mask: FloatArray, // [1,128]
  routing: FloatArray, // [1,32,128]
): FloatArray {
  val options =
    CompiledModel.Options(Accelerator.GPU).apply {
      // Explicit FP32: the default GPU precision changed 28 of 42 decisions.
      gpuOptions = CompiledModel.GpuOptions(precision = CompiledModel.GpuOptions.Precision.FP32)
    }
  val path = File(modelDir, "gliner25_decide_s128_wfp16.tflite").path
  CompiledModel.create(path, options, env).use { model ->
    val inputs = model.createInputBuffers()
    val outputs = model.createOutputBuffers()
    try {
      inputs[0].writeFloat(embeds)
      inputs[1].writeFloat(mask)
      inputs[2].writeFloat(routing)
      model.run(inputs, outputs)
      return outputs[0].readFloat() // slots 0..K-1 hold the K labels in request order
    } finally {
      (inputs + outputs).forEach { it.close() }
    }
  }
}

/** Single-label task: softmax over its slice of the logits, then the lowest-index maximum. */
fun decideSingleLabel(logits: FloatArray, offset: Int, labels: List<String>): Pair<String, Double> {
  val part = logits.copyOfRange(offset, offset + labels.size)
  val max = part.max()
  val weights = part.map { exp((it - max).toDouble()) }
  val best = weights.indices.maxBy { weights[it] }
  return labels[best] to weights[best] / weights.sum()
}

The complete Kotlin host is in android/: word splitter, SentencePiece Unigram tokenizer, gliner2 schema and input construction, the memory-mapped float16 table, and a float32 port of the decision rules including multi-label thresholds (DecideDecoder.kt), inside a Jetpack Compose app. Type a text and one task per line, tap Classify, and read each task's decision with its probability. Build and install steps are in android/README.md.

GPU precision

Weight storage and GPU computation precision are separate settings. With the runtime's default GPU precision, the s128 graph (float32 weights) compiled and returned finite logits on the Galaxy S26, but only 14 of 42 decisions matched the official result. With GpuOptions(precision = FP32) all 42 matched. Use the explicit option above for every file.

NPU

gliner25_decide_s128_npu_wfp16.tflite runs on the Qualcomm NPU of the Galaxy S26 (Hexagon v81 HTP) and gives the official decision on all 42 phone requests. The default s128 file does not: on the same NPU it compiles and runs, but 13 of 42 decisions match.

The NPU computes in float16. Three changes keep the graph within float16's range and precision. Each is exact in float32: on the 42 requests, the changed graph's outputs equal the default graph's bit for bit, in torch float32 and on desktop LiteRT CPU.

  • LayerNorm pre-scale. 22 of the 49 LayerNorms (the feed-forward output norms of layers 0–21) compute on x·2⁻ᵏ with eps·4⁻ᵏ, k from 1 to 5. In float32 their input reaches |x| = 1,423, so the squared deviation reaches 2.0e6 in 15 of 24 layers, above the float16 maximum of 65,504.
  • Attention mask. Masked positions add −1e4 instead of float32's lowest value, which has no float16 representation.
  • Feed-forward split. The first FULLY_CONNECTED of each feed-forward block runs as two halves joined by CONCATENATION before the GELU. Without the split, the NPU compiled the block's FC → GELU → FC chain into a path that added 19 to 25 times the float16 rounding error per block, and one decision flipped.
Galaxy S26, LiteRT 2.2.0FileDecisions equal officialMax |Δprob|Median ms per request
NPUnpu42/420.005528.6
GPU FP32npu42/420.0002669.1
GPU FP16 with FP32 accumulationnpu42/420.002860.2
GPU FP16npu41/420.01640.9
NPUdefault s12813/420.999826.2

Each row is one process on 2026-10-04, started at thermal status 0. The time is the median over the 42 requests of embedding lookup, input write, run and readback, one pass each. The first load compiles the graph for the NPU on the phone: 32.4 s. A second process read LiteRT's compiled cache in 0.68 s and returned the same 42 decisions. On the GPU, FP16 with FP32 accumulation keeps all 42 decisions and is 13 % faster than FP32. Plain FP16 flips one near-tie, whose official margin is 0.0051.

The native runner of the Latency section (input write → run → readback) timed the npu file on 2026-09-30. The NPU took 21.9–22.0 ms after a reboot. Before the reboot, with the phone up for a week, it took 41.2 ms, and the GPU at FP32 took 70.3 ms. The cause of the slower NPU state was not established. These are one phone's numbers on two days, not a benchmark across devices.

The NPU needs Qualcomm's runtime inside the app: the dispatch library and JIT compiler plugin of the LiteRT 2.2.0 release, and the QAIRT 2.47 HTP libraries for Hexagon v81. In Kotlin, create the Environment with the app's native library directory as DispatchLibraryDir and CompilerPluginLibraryDir, and pass Accelerator.NPU with QualcommOptions(htpPerformanceMode = BURST); every NPU run above used BURST. The Android sample in android/ runs the GPU and CPU paths only.

Fidelity

Reference: the official gliner2 2.0.0 classify_text on the fp32 checkpoint (desktop CPU), revision 7ee5da4c. Test requests: the 21 classify_text examples of the model card and 340 rows of the public dev split of fastino/fast-decisions (20 per domain, 17 domains, chosen by encoded length so that every row fits s512). Each window is tested on the requests that fit it. A decision matches when every task's label (or label set) equals the official one. Correlation was never used as a gate.

ConfigurationRequestsDecisions equal officialMax |Δlogit|
Desktop LiteRT CPU, wfp16 graphs + float16 tables128 42 / s256 328 / s512 36142/42 · 328/328 · 361/3614.81e-3
Desktop LiteRT CPU, fp32 graph + float32 tables512 361361/3613.29e-5
Galaxy S26 GPU FP32, native CompiledModel runner, wfp16 + float16 table42 per window42/42 · 42/42 · 42/424.78e-3
Galaxy S26 GPU FP32, Android sample126 (request, window) pairs126/1264.78e-3
Galaxy S26 CPU (XNNPACK, 4 threads), Android sample126 pairs126/1264.82e-3

Desktop: ai-edge-litert 2.1.6 CompiledModel, macOS arm64, 4 threads. S26: LiteRT 2.2.0; every GPU graph ran fully delegated in one partition (1,780 of 1,780 operators) with no NaN. The 42 phone requests per window are the 21 model-card examples plus 21 dev-split rows (at s256 and s512, the longest rows that fit). In the Android sample the on-device token ids, marker positions and padding also equal the captured Python inputs for all 126 pairs.

The dev-split rows measure fidelity to the official model, not accuracy. On this length-selected subset the official model's decisions equal the dataset's gold labels on 164 of 340 rows (364 of 580 tasks). The publisher's benchmark (60.2 % exact match) uses held-out rows that are not in the public dataset; the two numbers are not comparable. Across the 1,700 public dev rows, the encoded length is ≤ 128 tokens for 92 rows, 129–256 for 1,410 and 257–512 for 198; none exceeds 512, and the largest label set has 28 labels.

Latency

Galaxy S26 (SM-S942Q, SM8850, Android 16), LiteRT 2.2.0, wfp16 graphs, float16 table inputs. Native CompiledModel runner, one process per job, 42 requests per window, 2 warm-up runs then 8 (s128) or 5 (s256, s512) timed runs per request. A timed run is input write → run → output readback. Every job started cool (GPU ≤ 50 °C, thermal status 0), screen on, USB powered.

WindowGPU FP32, opening request (ms)GPU FP32, back-to-back median / min (ms)Last request (ms)GPU clock limit at the end
s12866.369.6 / 65.9 (336 runs)69.91,200 MHz
s256174.9198.8 / 174.7 (210 runs)214.8902 MHz
s512575.2733.1 / 573.5 (210 runs)912.1646 MHz

Back-to-back runs heat the phone: the GPU clock limit steps down from 1,300 MHz and each request of s256 / s512 gets slower during the job. The opening-request column is the cool value; the median mixes it with the throttled end. For comparison, the CPU (XNNPACK, 4 threads) took 510.0 / 311.5 ms (median / min) at s256 and 1,666.0 / 863.1 ms at s512 under the same protocol; the s128 CPU value, 216.5 / 183.6 ms (fp32 graph), was measured right after a GPU job on a warm phone (thermal status 2).

The Android sample (debug build, same phone, app in the foreground, the example request at s128):

ScenarioEnd to end (ms)
One request right after a cold start (three starts; 5 warm-up passes of 0.38–0.39 s before Ready)77.9 · 72.4 · 75.5
20 requests, one every 2 s, GPU FP32 (median; min 80.5, max 99.3)93.0
20 requests, one every 2 s, CPU (median; min 167.7, max 182.8)171.9

End to end covers tokenize and embed, graph to readback and decode; the paced runs also include the hand-off to the LiteRT worker thread. In the paced GPU run the graph took a median 68.4 ms; the tokenize-and-embed phase rose from about 13 ms to about 28 ms after the tenth request, and the cause was not established. These are single-device numbers, not a benchmark across devices or thermal states.

Limits

  • English, as the source model. One text per call; no batching.
  • At most 512 encoded tokens (task schemas plus text) and 32 labels per request. Longer requests are rejected, never truncated; gliner2's long-text chunking (classify_text_long) is not ported.
  • GPU and NPU validation cover the Galaxy S26 (Adreno GPU, Hexagon v81 NPU) only. Only the s128 window has an NPU file; the default files give wrong decisions on the NPU.
  • No INT8 file is shipped. On GLiNER2.5 Small, the same family of graphs, dynamic-range INT8 FULLY_CONNECTED weights did not compile on the LiteRT 2.2.0 GPU; no INT8 variant of this model was built or tested.
  • Few-shot examples and the extraction tasks of gliner2 (entities, relations, structures) are not part of the graphs.

Prior art

The classification path of GLiNER2.5-Decide is also published as ONNX in onnx-community/GLiNER2.5-Decide-ONNX (Transformers.js, WebGPU) and nishparadox/gliner2.5-decide-onnx. A Core AI (Apple) conversion of the same path is published as coreai-community/GLiNER2.5-Decide-CoreAI.

Reproduce

conversion/ holds the scripts that produced every file here (graph definition, export with litert-torch 0.9.3, float16 weight storage with ai-edge-quantizer 0.8.0, embedding tables, the hero image) and the order they ran in; requirements-lock.txt pins the environment. The graph keeps the checkpoint's arithmetic: host-side embedding lookup, one-hot label routing as a matrix product, rank-4 attention, DeBERTa's logarithmic relative-position buckets baked as constants, a float attention mask, and the native LiteRT GELU. The NPU file comes from profile_ranges.py, npu_graph.py, export_npu.py and quantize_npu.py, in the order conversion/README.md gives.

License

Apache-2.0, the license of the checkpoint (LICENSE). The DeBERTa-v3 encoder is MIT (licenses/DeBERTa-MIT.txt); the gliner2 and transformers code the host runtime and the Android sample build on is Apache-2.0 (licenses/). NOTICE lists the sources and the changes. Upstream: the GLiNER2 repository and the paper arXiv:2507.18546.

Citation

@misc{zaratiana2025gliner2efficientmultitaskinformation,
      title={GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface},
      author={Urchade Zaratiana and Gil Pasternak and Oliver Boyd and George Hurn-Maloney and Ash Lewis},
      year={2025},
      eprint={2507.18546},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2507.18546},
}
android
deberta-v3
gliner2
litert
text-classification
tflite
zero-shot-classification