knowledgator/gliclass-edge-v3.0 is a zero-shot text classifier with 32.7M parameters, built on a ModernBERT encoder. It reads a text and up to 25 candidate labels as one sequence and scores every label in one forward pass. This repository holds it converted to LiteRT with litert-torch 0.9.3. The verified path on desktop is the LiteRT CompiledModel Python API on CPU and, for the 128-token graph, on the Metal GPU with explicit FP32 computation (ai-edge-litert 2.1.6, Apple M4 Max). On a phone, it is the Kotlin CompiledModel API on the Galaxy S26 GPU with explicit FP32 computation, or on its CPU (LiteRT 2.2.0). All checks ran on 2026-10-02.
On desktop CPU, the shipped graphs with the float16 embedding table gave the same answers as the official gliclass 0.1.20 pipeline (CPU FP32) on all 552 test requests, in single-label and in multi-label mode. No logit differed by more than 0.009. On the Galaxy S26 GPU with explicit FP32, the Android sample in android/sample/ also gave the official answers on all 552 requests, from text to labels. Its graph call took a median of 5.86 ms at 128 tokens and 8.23 ms at 256 tokens.

The ticket is invented. The numbers are what gliclass_edge_v3_s128_fp32.tflite with the float16 table returned for it on desktop CPU. The official pipeline gives the same values to 3 decimals.
Other formats:
FluidInference/gliclass-edge-apps-coreml holds Core ML packages of a fine-tune of this model. Its card says the weights were tuned further for its own application before conversion, so they are not the same weights.cnmoro/gliclass-edge-v3.0-onnx holds an ONNX export of this checkpoint and an int8-quantized copy.| File | Bytes | Role |
|---|---|---|
gliclass_edge_v3_s128_fp32.tflite | 53,703,480 | Graph for up to 128 tokens (labels, prompt and text). The examples below use it. |
gliclass_edge_v3_s256_fp32.tflite | 53,965,476 | Graph for up to 256 tokens. |
gliclass_edge_v3_s128_wfp16.tflite | 27,012,624 | The 128-token graph with float16 weights. |
gliclass_edge_v3_s256_wfp16.tflite | 27,274,624 | The 256-token graph with float16 weights. |
host_assets/tok_embeddings_fp16.bin | 38,684,160 | Token-embedding table [50370, 384], little-endian float16, for the host lookup. |
host_assets/tokenizer.json | 3,583,595 | The source repository's tokenizer.json, unchanged. |
gliclass_litert.py | Python host: builds the request string, tokenizes, looks up the table rows, runs the graph and returns the pipeline's output. | |
requirements.txt | numpy, tokenizers, ai-edge-litert==2.1.6 for the Python host. | |
examples/run_example.py | The Python example below. | |
conversion/ | Conversion, fixture, check and figure scripts, with the environment lock. | |
fixtures/ | 152 test requests with the official outputs, and the pins that rebuild the other 400. | |
android/sample/ | Android sample app (Compose) with a pure-Kotlin host that ran on the Galaxy S26: tokenizer, request builder, float16 table lookup, LiteRT call, softmax and sigmoid. | |
android/CardSnippet.kt | The Kotlin block below, in the package it was compiled in. | |
REPRODUCE.md | How to reproduce the files and the checks. | |
LICENSE, NOTICE, licenses/ | Apache License 2.0 text, attribution and the retained upstream licenses. | |
assets/hero.png | The figure above. | |
SHA256SUMS | SHA-256 checksums of the files. |
The two fp32 graphs and host_assets/ are 149,936,711 bytes. The wfp16 graphs are half the size. They store the 78 FULLY_CONNECTED weights in float16, each behind a DEQUANTIZE operator, and keep activations in float32. On the 552 test requests they change 2 answers, both near ties (see below).
Each graph holds the 10-layer encoder, the read-out at [CLS] and at every <<LABEL>> token, the two projectors and the scorer. It returns one logit per label slot. The host does the rest: it builds the request string, tokenizes, looks up each token's table row, builds the label routing and applies softmax or sigmoid.
Needs numpy, tokenizers and ai-edge-litert (tested with 2.1.6). PyTorch, transformers and the gliclass package are not needed. Run it from the directory that holds the downloaded files. examples/run_example.py is the same code.
from gliclass_litert import GliclassLiteRT
text = (
"The Zorvik X2 router from Tallowmere Networks keeps dropping the connection every evening "
"around 9 pm. I have restarted it and reset it to factory settings, but the status light "
"still blinks orange."
)
labels = ["connectivity problem", "hardware defect", "billing issue", "feature request",
"installation help"]
with GliclassLiteRT(".") as classifier:
for mode in ("single-label", "multi-label"):
result = classifier.classify(text, labels, classification_type=mode)
print(mode, [(r["label"], round(r["score"], 3)) for r in result])
On a Mac CPU it prints the following. The code rounds each score to 3 decimals.
single-label [('connectivity problem', 0.856)]
multi-label [('connectivity problem', 0.995), ('hardware defect', 0.95), ('billing issue', 0.886), ('feature request', 0.811)]
Single-label mode picks connectivity problem with softmax 0.856. Multi-label mode keeps every label whose sigmoid is at least 0.5, here four of the five. On the same request (61 tokens), the official pipeline on CPU FP32 returns the same labels and the same scores to 3 decimals.
classify() returns the pipeline's output: a list of {"label", "score"}. Pass a task description as prompt=; it goes right before the text. threshold= sets the multi-label threshold. For the desktop GPU, pass accelerator="gpu", which sets explicit FP32 (GpuOptions(enforce_f32=True)). The graph passed that way on the Metal GPU (482 of 482 requests). This host's own GPU path has not been run. To use the float16-weight graphs, pass their paths in the mapping form of the constructor, described in its docstring.
import android.util.Half
import com.google.ai.edge.litert.Accelerator
import com.google.ai.edge.litert.CompiledModel
import java.io.File
import java.io.RandomAccessFile
import java.nio.ByteOrder
import java.nio.channels.FileChannel
import kotlin.math.exp
/** One GLiClass-Edge window (128 or 256 tokens) on the GPU with explicit FP32. */
class GliclassGpu(dir: File, private val window: Int = 128) : AutoCloseable {
private val options = CompiledModel.Options(Accelerator.GPU).apply {
gpuOptions = CompiledModel.GpuOptions(precision = CompiledModel.GpuOptions.Precision.FP32)
}
private val model =
CompiledModel.create(File(dir, "gliclass_edge_v3_s${window}_fp32.tflite").absolutePath, options, null)
private val inputs = listOf("inputs_embeds", "attention_mask", "label_routing")
.associateWith { model.createInputBuffer(it, "serving_default") }
private val outputs = mapOf("logits" to model.createOutputBuffer("logits", "serving_default"))
private val table = RandomAccessFile(File(dir, "host_assets/tok_embeddings_fp16.bin"), "r").use {
it.channel.map(FileChannel.MapMode.READ_ONLY, 0, it.length()).order(ByteOrder.LITTLE_ENDIAN).asShortBuffer()
}
/** ids = [CLS] … [SEP] of the linearized request; labelPositions = index of each <<LABEL>> token. */
fun logits(ids: IntArray, labelPositions: IntArray): FloatArray {
require(ids.size <= window && labelPositions.size <= 25)
val embeds = FloatArray(window * 384)
for (p in 0 until window) {
val row = (if (p < ids.size) ids[p] else 50283) * 384 // 50283 = [PAD]
for (c in 0 until 384) embeds[p * 384 + c] = Half.toFloat(table.get(row + c))
}
val routing = FloatArray(25 * window)
labelPositions.forEachIndexed { k, position -> routing[k * window + position] = 1f }
inputs.getValue("inputs_embeds").writeFloat(embeds)
inputs.getValue("attention_mask").writeFloat(FloatArray(window) { if (it < ids.size) 1f else 0f })
inputs.getValue("label_routing").writeFloat(routing)
model.run(inputs, outputs, "serving_default")
return outputs.getValue("logits").readFloat().copyOf(labelPositions.size)
}
override fun close() {
(inputs.values + outputs.values).forEach { it.close() }
model.close()
}
}
/** Single-label: softmax over the labels, the top one wins. */
fun softmax(logits: FloatArray): List<Double> {
val top = logits.max()
val e = logits.map { exp((it - top).toDouble()) }
val sum = e.sum()
return e.map { it / sum }
}
/** Multi-label: every label whose sigmoid reaches the threshold. */
fun chosen(logits: FloatArray, threshold: Double = 0.5): List<Int> =
logits.indices.filter { 1 / (1 + exp(-logits[it].toDouble())) >= threshold }
Create GliclassGpu once per window and call logits from one worker thread. ids are the token IDs of the request string with [CLS] and [SEP]. labelPositions holds the index of each <<LABEL>> token. Both must come from a tokenizer that matches gliclass_litert.py. The block compiles with LiteRT 2.2.0 (AGP 8.9.1, Kotlin 2.2.21). It has not itself run on a phone; the sample below has.
The full Kotlin host that ran on the Galaxy S26 is in android/sample/, with its own README for download, build and install. Its GliclassTokenizer.kt turns text into IDs with tokenizer.json, without native code. GliclassInputs.kt builds the request and the inputs, and GliclassDecoder.kt reproduces the pipeline's float32 softmax and sigmoid.
gliclass_litert.py implements this contract. It follows ZeroShotClassificationPipeline of gliclass 0.1.20 with a uni-encoder model.
<<LABEL>>label 1<<LABEL>>label 2…<<SEP>>, then the prompt, then the text. Nothing separates the prompt from the text, not even a space.tokenizer.json: [CLS] at the start, [SEP] at the end, no truncation. The IDs are <<LABEL>> 50368, <<SEP>> 50369, [CLS] 50281, [SEP] 50282 and [PAD] 50283. The ByteLevel pre-tokenizer adds a leading space to each piece between added tokens, so the prompt starts with a space token. Text glued to the end of a prompt does not.[PAD] to the window S. inputs_embeds [1,S,384] holds the table row of every position, padding included, widened from float16 to float32. attention_mask [1,S] is 1 for real tokens and 0 for padding. In label_routing [1,25,S], row k is 1 at the k-th <<LABEL>> token and 0 elsewhere. Rows past the last label stay zero.serving_default signature. It returns logits [1,1,1,25], one logit per label slot. Slots 0 to n−1 belong to the n labels; ignore the rest.On all 552 test requests, gliclass_litert.py builds the same string, token IDs and <<LABEL>> positions as the official pipeline. With the shipped fp32 graphs and the float16 table on desktop CPU, it returns the same single-label and multi-label answers on all 552.
The reference is the official gliclass 0.1.20 pipeline on CPU FP32 (torch 2.12.1, transformers 5.17.0), one request per call. "Same top label" compares the single-label answer, the softmax top-1. "Same label set" compares the multi-label answer, the labels with sigmoid at or above 0.5.
The 552 test requests are all 144 items of SemIf authored144.jsonl (three options each; the item's question is the prompt), 200 ag_news test rows, 200 banking77 test rows, the source model card's two examples and six invented requests. banking77 has 77 intents and the graph holds 25 labels, so each banking77 request carries the gold intent and 24 others drawn with a fixed seed. 482 requests fit 128 tokens; the other 70 need 256. For 419 of the 552, the reference's top softmax score is below 0.9.
The desktop rows ran on an Apple M4 Max with ai-edge-litert 2.1.6 through the Python CompiledModel API. Other jobs shared the machine, so desktop times are informational. The phone rows ran on one Galaxy S26 (SM-S942Q, Android 16) with LiteRT 2.2.0 through the Kotlin CompiledModel API, on USB power, in the foreground, at thermal status 0 or 1 (each table says which). A phone median is the middle value of the sorted times; for an even count it is the upper of the two middle values.
| Where | Graph, table | Requests | Same top label | Same label set | Max logit difference | Median ms |
|---|---|---|---|---|---|---|
| CPU | S128 fp32, float16 table (shipped) | 482 | 482 | 482 | 0.0090 | 8.77 |
| CPU | S128 wfp16, float16 table (shipped) | 482 | 481 | 481 | 0.017 | 9.25 |
| CPU | S128 fp32 †, float32 table | 482 | 482 | 482 | 0.000036 | 21.21 |
| CPU | S256 fp32, float16 table (shipped) | 552 | 552 | 552 | 0.0090 | 11.04 |
| CPU | S256 wfp16, float16 table (shipped) | 552 | 551 | 551 | 0.017 | 11.35 |
| CPU, Python host from text to labels | S128 and S256 fp32, float16 table (shipped) | 552 | 552 | 552 | 0.0090 | |
| Metal GPU, explicit FP32 | S128 fp32, float16 table (shipped) | 482 | 482 | 482 | 0.0091 | 2.77 |
| Metal GPU, default precision | S128 fp32 and wfp16, float16 table (shipped) | 482 | 479 | 469 | 0.135 | 2.29 / 2.34 |
| Metal GPU, default precision | S128 fp32 †, float16 table | 482 | 258 | 133 | 10.0 | 2.26 |
† The same graph without the SafeLayerNorm rescale (see the fp16 section). With FP32 computation it returns bit-identical logits to the shipped graph: on all 482 S128 requests on desktop CPU, for the fp32 and the wfp16 form, and on the S26 GPU.
All logits were finite in every desktop row. On the shipped S128 fp32 row, the largest softmax difference is 0.0015 and the largest sigmoid difference 0.0018. The CPU timings come from 8-thread runs on a loaded machine.
A debug gate app fed each graph the reference token IDs, so these rows measure the graph and the table, not a tokenizer. Times are warm medians of writing the inputs, running and reading back, after 5 warm-up requests; the table lookup is not included. The explicit-FP32 and CPU rows ran at thermal status 0 with the battery at 37 to 38 °C. The default-precision and NPU rows ran earlier the same day at thermal status 1 with the battery at 40 to 41 °C.
| Accelerator | Graph, float16 table | Requests | Same top label | Same label set | Max logit difference | Compile ms | Warm median ms | Delegation |
|---|---|---|---|---|---|---|---|---|
| GPU, explicit FP32 | S128 fp32 (shipped) | 482 | 482 | 482 | 0.0091 | 548 | 5.80 | 670 of 670 nodes, 1 partition |
| GPU, explicit FP32 | S128 wfp16 (shipped) | 482 | 481 | 481 | 0.017 | 554 | 5.84 | 748 of 748, 1 partition |
| GPU, explicit FP32 | S256 fp32 (shipped) | 552 | 552 | 552 | 0.0091 | 574 | 8.32 | 670 of 670, 1 partition |
| GPU, explicit FP32 | S256 wfp16 (shipped) | 552 | 551 | 551 | 0.017 | 568 | 8.33 | 748 of 748, 1 partition |
| CPU, 4 threads | S128 fp32 (shipped) | 482 | 482 | 482 | 0.0090 | 132 | 6.77 | 670 of 670 XNNPACK |
| GPU, default precision | S128 fp32 and wfp16 (shipped) | 482 | 475 | 455 | 0.37 | 568 / 592 | 3.86 / 3.85 | 670 of 670 / 748 of 748, 1 partition |
| GPU, default precision | S128 fp32 † | 482 | 261 | 138 | 10.1 | 808 | 3.83 | 655 of 655, 1 partition |
| NPU (Qualcomm, JIT, BURST) | S128 wfp16 (shipped) | 482 | 480 | 462 | 0.24 | 978, JIT included | 2.51 | 748 of 748 ops, 1 partition |
In an earlier run at thermal status 1, the shipped S128 fp32 graph returned bit-identical logits to its † twin on all 482 requests (5.83 and 5.76 ms), and the shipped fp32 and wfp16 graphs returned bit-identical logits to each other at default precision. Both wfp16 rows with explicit FP32 differ on the same 2 requests as on desktop.
The sample in android/sample/ uses the shipped fp32 graphs and the float16 table, with its own Kotlin tokenizer.
<<LABEL>> positions and padding equal the official pipeline's on all 552. Single-label and multi-label answers match on all 552. Max logit difference 0.0091 on the GPU and 0.0090 on the CPU. Both windows run whole on the GPU delegate, in one partition.Not measured: Pixel phones, other phones, sustained load, the NPU at 256 tokens.
The official pipeline's single-label answer equals the dataset label on 67 of 144 SemIf items, 115 of 200 ag_news rows and 71 of 200 banking77 rows (25-label subsets). These are our own subsets, not the publisher's benchmark. The shipped graphs give the same answers on all 552, so they score the same. For task accuracy, see the benchmark on the source model card.
Run the GPU with explicit FP32 computation, or the CPU. At fp16 precision, LayerNorm overflows.
The residual stream grows large in the later layers. Over the 552 requests, a LayerNorm input reaches a distance of 2,064 from its row mean, and the row's sum of squared deviations reaches 6.4 million. The fp16 maximum is 65,504. The outputs stay finite, but the answers change. With unmodified LayerNorms, the Metal GPU at default precision keeps 258 of 482 top labels and 133 of 482 label sets. The S26 GPU at default precision keeps 261 and 138.
The shipped graphs compute 15 LayerNorms on x·2⁻ᵏ with eps·2⁻²ᵏ: k = 5 in layers 3 to 6, and k = 6 in layers 7 to 9 and the final norm. This SafeLayerNorm is exact in FP32: the logits stay bit-identical to the unmodified graph (the † rows). At fp16 precision it keeps most answers. The Metal GPU at default precision keeps 479 top labels and 469 label sets of 482. The S26 GPU at default precision keeps 475 and 455, at 3.86 ms. The S26 NPU runs the whole S128 wfp16 graph in one partition at 2.51 ms and keeps 480 and 462. Every answer that changes sits near the boundary: on the NPU, the changed top labels had reference gaps of at most 0.0043, and the changed sets a sigmoid at most 0.026 from 0.5.
Those fp16 paths are measured here, but they are not the verified mode, because they change answers. Explicit FP32 is GpuOptions(precision = FP32) in Kotlin and GpuOptions(enforce_f32=True) in Python.
The checkpoint is float32, so float16 storage also costs a little. The float16 table rounds values by up to 1.2e-4: with the float32 table the largest logit difference is 0.000036, with the float16 table 0.0090, and the answers are the same. In the wfp16 graphs, the converter has folded each LayerNorm scale into the next FULLY_CONNECTED weights (50 of the 78), and float16 rounds those products by up to 7.7e-4. That changes 2 answers of 552: banking77_015, whose top two logits differ by 0.0012, and banking77_064, whose sigmoid for one label sits 0.0002 below 0.5.
knowledgator/gliclass-edge-v3.0 at revision df03993a2ed98e5e4a0d2dd7efbbd105abe874cf. model.safetensors is 130,829,312 bytes, all F32, with 32,705,154 parameters, SHA-256 ed2600439be991eaab06d831b259e99057a69242acf3a26a61b053695b1b979c.jhu-clsp/ettin-encoder-32m (MIT) with 10 layers, hidden size 384, 6 heads and a GeGLU FFN of 576. Layers 0, 3, 6 and 9 attend globally; the others see 64 tokens on each side. The <<LABEL>> hidden states go through one projector and the [CLS] hidden state through another. An MLP scorer (768 to 256 to 128 to 1) turns each pair into a logit.BioMike/formal-logic-reasoning-gliclass-2k (its card states no license), knowledgator/gliclass-v3-logic-dataset (Apache-2.0) and tau/commonsense_qa (MIT).torch.matmul to rank-3 batch matmuls and the label read-out to a rank-2 one. A rank-4 marker, vendored from an earlier litert-community DeBERTa conversion, keeps all 21 BATCH_MATMUL operators at rank 4.b22f9bd2aedb1768089da8126dc6cd69572b61f21e39d6254119773b5e147a27. tokenizer.json is the source file, unchanged.fixtures/ holds the 152 requests whose text may be shared (SemIf, MIT; the source card's examples; the invented requests) with the official outputs. The ag_news rows (license "unknown" on its dataset card) and the banking77 rows (CC BY 4.0) are not copied. fixtures/ pins them by dataset revision and row index, and conversion/make_fixtures.py rebuilds them.License: the source model is licensed under Apache 2.0 by Knowledgator. These converted files, the conversion scripts, the Python host, the Kotlin block and the Android sample are released under the same license. The encoder is MIT. The license text is in LICENSE, attribution is in NOTICE, and the retained upstream licenses are in licenses/.
knowledgator/gliclass-edge-v3.0 is a zero-shot text classifier with 32.7M parameters, built on a ModernBERT encoder. It reads a text and up to 25 candidate labels as one sequence and scores every label in one forward pass. This repository holds it converted to LiteRT with litert-torch 0.9.3. The verified path on desktop is the LiteRT CompiledModel Python API on CPU and, for the 128-token graph, on the Metal GPU with explicit FP32 computation (ai-edge-litert 2.1.6, Apple M4 Max). On a phone, it is the Kotlin CompiledModel API on the Galaxy S26 GPU with explicit FP32 computation, or on its CPU (LiteRT 2.2.0). All checks ran on 2026-10-02.
On desktop CPU, the shipped graphs with the float16 embedding table gave the same answers as the official gliclass 0.1.20 pipeline (CPU FP32) on all 552 test requests, in single-label and in multi-label mode. No logit differed by more than 0.009. On the Galaxy S26 GPU with explicit FP32, the Android sample in android/sample/ also gave the official answers on all 552 requests, from text to labels. Its graph call took a median of 5.86 ms at 128 tokens and 8.23 ms at 256 tokens.

The ticket is invented. The numbers are what gliclass_edge_v3_s128_fp32.tflite with the float16 table returned for it on desktop CPU. The official pipeline gives the same values to 3 decimals.
Other formats:
FluidInference/gliclass-edge-apps-coreml holds Core ML packages of a fine-tune of this model. Its card says the weights were tuned further for its own application before conversion, so they are not the same weights.cnmoro/gliclass-edge-v3.0-onnx holds an ONNX export of this checkpoint and an int8-quantized copy.| File | Bytes | Role |
|---|---|---|
gliclass_edge_v3_s128_fp32.tflite | 53,703,480 | Graph for up to 128 tokens (labels, prompt and text). The examples below use it. |
gliclass_edge_v3_s256_fp32.tflite | 53,965,476 | Graph for up to 256 tokens. |
gliclass_edge_v3_s128_wfp16.tflite | 27,012,624 | The 128-token graph with float16 weights. |
gliclass_edge_v3_s256_wfp16.tflite | 27,274,624 | The 256-token graph with float16 weights. |
host_assets/tok_embeddings_fp16.bin | 38,684,160 | Token-embedding table [50370, 384], little-endian float16, for the host lookup. |
host_assets/tokenizer.json | 3,583,595 | The source repository's tokenizer.json, unchanged. |
gliclass_litert.py | Python host: builds the request string, tokenizes, looks up the table rows, runs the graph and returns the pipeline's output. | |
requirements.txt | numpy, tokenizers, ai-edge-litert==2.1.6 for the Python host. | |
examples/run_example.py | The Python example below. | |
conversion/ | Conversion, fixture, check and figure scripts, with the environment lock. | |
fixtures/ | 152 test requests with the official outputs, and the pins that rebuild the other 400. | |
android/sample/ | Android sample app (Compose) with a pure-Kotlin host that ran on the Galaxy S26: tokenizer, request builder, float16 table lookup, LiteRT call, softmax and sigmoid. | |
android/CardSnippet.kt | The Kotlin block below, in the package it was compiled in. | |
REPRODUCE.md | How to reproduce the files and the checks. | |
LICENSE, NOTICE, licenses/ | Apache License 2.0 text, attribution and the retained upstream licenses. | |
assets/hero.png | The figure above. | |
SHA256SUMS | SHA-256 checksums of the files. |
The two fp32 graphs and host_assets/ are 149,936,711 bytes. The wfp16 graphs are half the size. They store the 78 FULLY_CONNECTED weights in float16, each behind a DEQUANTIZE operator, and keep activations in float32. On the 552 test requests they change 2 answers, both near ties (see below).
Each graph holds the 10-layer encoder, the read-out at [CLS] and at every <<LABEL>> token, the two projectors and the scorer. It returns one logit per label slot. The host does the rest: it builds the request string, tokenizes, looks up each token's table row, builds the label routing and applies softmax or sigmoid.
Needs numpy, tokenizers and ai-edge-litert (tested with 2.1.6). PyTorch, transformers and the gliclass package are not needed. Run it from the directory that holds the downloaded files. examples/run_example.py is the same code.
from gliclass_litert import GliclassLiteRT
text = (
"The Zorvik X2 router from Tallowmere Networks keeps dropping the connection every evening "
"around 9 pm. I have restarted it and reset it to factory settings, but the status light "
"still blinks orange."
)
labels = ["connectivity problem", "hardware defect", "billing issue", "feature request",
"installation help"]
with GliclassLiteRT(".") as classifier:
for mode in ("single-label", "multi-label"):
result = classifier.classify(text, labels, classification_type=mode)
print(mode, [(r["label"], round(r["score"], 3)) for r in result])
On a Mac CPU it prints the following. The code rounds each score to 3 decimals.
single-label [('connectivity problem', 0.856)]
multi-label [('connectivity problem', 0.995), ('hardware defect', 0.95), ('billing issue', 0.886), ('feature request', 0.811)]
Single-label mode picks connectivity problem with softmax 0.856. Multi-label mode keeps every label whose sigmoid is at least 0.5, here four of the five. On the same request (61 tokens), the official pipeline on CPU FP32 returns the same labels and the same scores to 3 decimals.
classify() returns the pipeline's output: a list of {"label", "score"}. Pass a task description as prompt=; it goes right before the text. threshold= sets the multi-label threshold. For the desktop GPU, pass accelerator="gpu", which sets explicit FP32 (GpuOptions(enforce_f32=True)). The graph passed that way on the Metal GPU (482 of 482 requests). This host's own GPU path has not been run. To use the float16-weight graphs, pass their paths in the mapping form of the constructor, described in its docstring.
import android.util.Half
import com.google.ai.edge.litert.Accelerator
import com.google.ai.edge.litert.CompiledModel
import java.io.File
import java.io.RandomAccessFile
import java.nio.ByteOrder
import java.nio.channels.FileChannel
import kotlin.math.exp
/** One GLiClass-Edge window (128 or 256 tokens) on the GPU with explicit FP32. */
class GliclassGpu(dir: File, private val window: Int = 128) : AutoCloseable {
private val options = CompiledModel.Options(Accelerator.GPU).apply {
gpuOptions = CompiledModel.GpuOptions(precision = CompiledModel.GpuOptions.Precision.FP32)
}
private val model =
CompiledModel.create(File(dir, "gliclass_edge_v3_s${window}_fp32.tflite").absolutePath, options, null)
private val inputs = listOf("inputs_embeds", "attention_mask", "label_routing")
.associateWith { model.createInputBuffer(it, "serving_default") }
private val outputs = mapOf("logits" to model.createOutputBuffer("logits", "serving_default"))
private val table = RandomAccessFile(File(dir, "host_assets/tok_embeddings_fp16.bin"), "r").use {
it.channel.map(FileChannel.MapMode.READ_ONLY, 0, it.length()).order(ByteOrder.LITTLE_ENDIAN).asShortBuffer()
}
/** ids = [CLS] … [SEP] of the linearized request; labelPositions = index of each <<LABEL>> token. */
fun logits(ids: IntArray, labelPositions: IntArray): FloatArray {
require(ids.size <= window && labelPositions.size <= 25)
val embeds = FloatArray(window * 384)
for (p in 0 until window) {
val row = (if (p < ids.size) ids[p] else 50283) * 384 // 50283 = [PAD]
for (c in 0 until 384) embeds[p * 384 + c] = Half.toFloat(table.get(row + c))
}
val routing = FloatArray(25 * window)
labelPositions.forEachIndexed { k, position -> routing[k * window + position] = 1f }
inputs.getValue("inputs_embeds").writeFloat(embeds)
inputs.getValue("attention_mask").writeFloat(FloatArray(window) { if (it < ids.size) 1f else 0f })
inputs.getValue("label_routing").writeFloat(routing)
model.run(inputs, outputs, "serving_default")
return outputs.getValue("logits").readFloat().copyOf(labelPositions.size)
}
override fun close() {
(inputs.values + outputs.values).forEach { it.close() }
model.close()
}
}
/** Single-label: softmax over the labels, the top one wins. */
fun softmax(logits: FloatArray): List<Double> {
val top = logits.max()
val e = logits.map { exp((it - top).toDouble()) }
val sum = e.sum()
return e.map { it / sum }
}
/** Multi-label: every label whose sigmoid reaches the threshold. */
fun chosen(logits: FloatArray, threshold: Double = 0.5): List<Int> =
logits.indices.filter { 1 / (1 + exp(-logits[it].toDouble())) >= threshold }
Create GliclassGpu once per window and call logits from one worker thread. ids are the token IDs of the request string with [CLS] and [SEP]. labelPositions holds the index of each <<LABEL>> token. Both must come from a tokenizer that matches gliclass_litert.py. The block compiles with LiteRT 2.2.0 (AGP 8.9.1, Kotlin 2.2.21). It has not itself run on a phone; the sample below has.
The full Kotlin host that ran on the Galaxy S26 is in android/sample/, with its own README for download, build and install. Its GliclassTokenizer.kt turns text into IDs with tokenizer.json, without native code. GliclassInputs.kt builds the request and the inputs, and GliclassDecoder.kt reproduces the pipeline's float32 softmax and sigmoid.
gliclass_litert.py implements this contract. It follows ZeroShotClassificationPipeline of gliclass 0.1.20 with a uni-encoder model.
<<LABEL>>label 1<<LABEL>>label 2…<<SEP>>, then the prompt, then the text. Nothing separates the prompt from the text, not even a space.tokenizer.json: [CLS] at the start, [SEP] at the end, no truncation. The IDs are <<LABEL>> 50368, <<SEP>> 50369, [CLS] 50281, [SEP] 50282 and [PAD] 50283. The ByteLevel pre-tokenizer adds a leading space to each piece between added tokens, so the prompt starts with a space token. Text glued to the end of a prompt does not.[PAD] to the window S. inputs_embeds [1,S,384] holds the table row of every position, padding included, widened from float16 to float32. attention_mask [1,S] is 1 for real tokens and 0 for padding. In label_routing [1,25,S], row k is 1 at the k-th <<LABEL>> token and 0 elsewhere. Rows past the last label stay zero.serving_default signature. It returns logits [1,1,1,25], one logit per label slot. Slots 0 to n−1 belong to the n labels; ignore the rest.On all 552 test requests, gliclass_litert.py builds the same string, token IDs and <<LABEL>> positions as the official pipeline. With the shipped fp32 graphs and the float16 table on desktop CPU, it returns the same single-label and multi-label answers on all 552.
The reference is the official gliclass 0.1.20 pipeline on CPU FP32 (torch 2.12.1, transformers 5.17.0), one request per call. "Same top label" compares the single-label answer, the softmax top-1. "Same label set" compares the multi-label answer, the labels with sigmoid at or above 0.5.
The 552 test requests are all 144 items of SemIf authored144.jsonl (three options each; the item's question is the prompt), 200 ag_news test rows, 200 banking77 test rows, the source model card's two examples and six invented requests. banking77 has 77 intents and the graph holds 25 labels, so each banking77 request carries the gold intent and 24 others drawn with a fixed seed. 482 requests fit 128 tokens; the other 70 need 256. For 419 of the 552, the reference's top softmax score is below 0.9.
The desktop rows ran on an Apple M4 Max with ai-edge-litert 2.1.6 through the Python CompiledModel API. Other jobs shared the machine, so desktop times are informational. The phone rows ran on one Galaxy S26 (SM-S942Q, Android 16) with LiteRT 2.2.0 through the Kotlin CompiledModel API, on USB power, in the foreground, at thermal status 0 or 1 (each table says which). A phone median is the middle value of the sorted times; for an even count it is the upper of the two middle values.
| Where | Graph, table | Requests | Same top label | Same label set | Max logit difference | Median ms |
|---|---|---|---|---|---|---|
| CPU | S128 fp32, float16 table (shipped) | 482 | 482 | 482 | 0.0090 | 8.77 |
| CPU | S128 wfp16, float16 table (shipped) | 482 | 481 | 481 | 0.017 | 9.25 |
| CPU | S128 fp32 †, float32 table | 482 | 482 | 482 | 0.000036 | 21.21 |
| CPU | S256 fp32, float16 table (shipped) | 552 | 552 | 552 | 0.0090 | 11.04 |
| CPU | S256 wfp16, float16 table (shipped) | 552 | 551 | 551 | 0.017 | 11.35 |
| CPU, Python host from text to labels | S128 and S256 fp32, float16 table (shipped) | 552 | 552 | 552 | 0.0090 | |
| Metal GPU, explicit FP32 | S128 fp32, float16 table (shipped) | 482 | 482 | 482 | 0.0091 | 2.77 |
| Metal GPU, default precision | S128 fp32 and wfp16, float16 table (shipped) | 482 | 479 | 469 | 0.135 | 2.29 / 2.34 |
| Metal GPU, default precision | S128 fp32 †, float16 table | 482 | 258 | 133 | 10.0 | 2.26 |
† The same graph without the SafeLayerNorm rescale (see the fp16 section). With FP32 computation it returns bit-identical logits to the shipped graph: on all 482 S128 requests on desktop CPU, for the fp32 and the wfp16 form, and on the S26 GPU.
All logits were finite in every desktop row. On the shipped S128 fp32 row, the largest softmax difference is 0.0015 and the largest sigmoid difference 0.0018. The CPU timings come from 8-thread runs on a loaded machine.
A debug gate app fed each graph the reference token IDs, so these rows measure the graph and the table, not a tokenizer. Times are warm medians of writing the inputs, running and reading back, after 5 warm-up requests; the table lookup is not included. The explicit-FP32 and CPU rows ran at thermal status 0 with the battery at 37 to 38 °C. The default-precision and NPU rows ran earlier the same day at thermal status 1 with the battery at 40 to 41 °C.
| Accelerator | Graph, float16 table | Requests | Same top label | Same label set | Max logit difference | Compile ms | Warm median ms | Delegation |
|---|---|---|---|---|---|---|---|---|
| GPU, explicit FP32 | S128 fp32 (shipped) | 482 | 482 | 482 | 0.0091 | 548 | 5.80 | 670 of 670 nodes, 1 partition |
| GPU, explicit FP32 | S128 wfp16 (shipped) | 482 | 481 | 481 | 0.017 | 554 | 5.84 | 748 of 748, 1 partition |
| GPU, explicit FP32 | S256 fp32 (shipped) | 552 | 552 | 552 | 0.0091 | 574 | 8.32 | 670 of 670, 1 partition |
| GPU, explicit FP32 | S256 wfp16 (shipped) | 552 | 551 | 551 | 0.017 | 568 | 8.33 | 748 of 748, 1 partition |
| CPU, 4 threads | S128 fp32 (shipped) | 482 | 482 | 482 | 0.0090 | 132 | 6.77 | 670 of 670 XNNPACK |
| GPU, default precision | S128 fp32 and wfp16 (shipped) | 482 | 475 | 455 | 0.37 | 568 / 592 | 3.86 / 3.85 | 670 of 670 / 748 of 748, 1 partition |
| GPU, default precision | S128 fp32 † | 482 | 261 | 138 | 10.1 | 808 | 3.83 | 655 of 655, 1 partition |
| NPU (Qualcomm, JIT, BURST) | S128 wfp16 (shipped) | 482 | 480 | 462 | 0.24 | 978, JIT included | 2.51 | 748 of 748 ops, 1 partition |
In an earlier run at thermal status 1, the shipped S128 fp32 graph returned bit-identical logits to its † twin on all 482 requests (5.83 and 5.76 ms), and the shipped fp32 and wfp16 graphs returned bit-identical logits to each other at default precision. Both wfp16 rows with explicit FP32 differ on the same 2 requests as on desktop.
The sample in android/sample/ uses the shipped fp32 graphs and the float16 table, with its own Kotlin tokenizer.
<<LABEL>> positions and padding equal the official pipeline's on all 552. Single-label and multi-label answers match on all 552. Max logit difference 0.0091 on the GPU and 0.0090 on the CPU. Both windows run whole on the GPU delegate, in one partition.Not measured: Pixel phones, other phones, sustained load, the NPU at 256 tokens.
The official pipeline's single-label answer equals the dataset label on 67 of 144 SemIf items, 115 of 200 ag_news rows and 71 of 200 banking77 rows (25-label subsets). These are our own subsets, not the publisher's benchmark. The shipped graphs give the same answers on all 552, so they score the same. For task accuracy, see the benchmark on the source model card.
Run the GPU with explicit FP32 computation, or the CPU. At fp16 precision, LayerNorm overflows.
The residual stream grows large in the later layers. Over the 552 requests, a LayerNorm input reaches a distance of 2,064 from its row mean, and the row's sum of squared deviations reaches 6.4 million. The fp16 maximum is 65,504. The outputs stay finite, but the answers change. With unmodified LayerNorms, the Metal GPU at default precision keeps 258 of 482 top labels and 133 of 482 label sets. The S26 GPU at default precision keeps 261 and 138.
The shipped graphs compute 15 LayerNorms on x·2⁻ᵏ with eps·2⁻²ᵏ: k = 5 in layers 3 to 6, and k = 6 in layers 7 to 9 and the final norm. This SafeLayerNorm is exact in FP32: the logits stay bit-identical to the unmodified graph (the † rows). At fp16 precision it keeps most answers. The Metal GPU at default precision keeps 479 top labels and 469 label sets of 482. The S26 GPU at default precision keeps 475 and 455, at 3.86 ms. The S26 NPU runs the whole S128 wfp16 graph in one partition at 2.51 ms and keeps 480 and 462. Every answer that changes sits near the boundary: on the NPU, the changed top labels had reference gaps of at most 0.0043, and the changed sets a sigmoid at most 0.026 from 0.5.
Those fp16 paths are measured here, but they are not the verified mode, because they change answers. Explicit FP32 is GpuOptions(precision = FP32) in Kotlin and GpuOptions(enforce_f32=True) in Python.
The checkpoint is float32, so float16 storage also costs a little. The float16 table rounds values by up to 1.2e-4: with the float32 table the largest logit difference is 0.000036, with the float16 table 0.0090, and the answers are the same. In the wfp16 graphs, the converter has folded each LayerNorm scale into the next FULLY_CONNECTED weights (50 of the 78), and float16 rounds those products by up to 7.7e-4. That changes 2 answers of 552: banking77_015, whose top two logits differ by 0.0012, and banking77_064, whose sigmoid for one label sits 0.0002 below 0.5.
knowledgator/gliclass-edge-v3.0 at revision df03993a2ed98e5e4a0d2dd7efbbd105abe874cf. model.safetensors is 130,829,312 bytes, all F32, with 32,705,154 parameters, SHA-256 ed2600439be991eaab06d831b259e99057a69242acf3a26a61b053695b1b979c.jhu-clsp/ettin-encoder-32m (MIT) with 10 layers, hidden size 384, 6 heads and a GeGLU FFN of 576. Layers 0, 3, 6 and 9 attend globally; the others see 64 tokens on each side. The <<LABEL>> hidden states go through one projector and the [CLS] hidden state through another. An MLP scorer (768 to 256 to 128 to 1) turns each pair into a logit.BioMike/formal-logic-reasoning-gliclass-2k (its card states no license), knowledgator/gliclass-v3-logic-dataset (Apache-2.0) and tau/commonsense_qa (MIT).torch.matmul to rank-3 batch matmuls and the label read-out to a rank-2 one. A rank-4 marker, vendored from an earlier litert-community DeBERTa conversion, keeps all 21 BATCH_MATMUL operators at rank 4.b22f9bd2aedb1768089da8126dc6cd69572b61f21e39d6254119773b5e147a27. tokenizer.json is the source file, unchanged.fixtures/ holds the 152 requests whose text may be shared (SemIf, MIT; the source card's examples; the invented requests) with the official outputs. The ag_news rows (license "unknown" on its dataset card) and the banking77 rows (CC BY 4.0) are not copied. fixtures/ pins them by dataset revision and row index, and conversion/make_fixtures.py rebuilds them.License: the source model is licensed under Apache 2.0 by Knowledgator. These converted files, the conversion scripts, the Python host, the Kotlin block and the Android sample are released under the same license. The encoder is MIT. The license text is in LICENSE, attribution is in NOTICE, and the retained upstream licenses are in licenses/.