litert-community/GLiClass-Edge-v3.0-LiteRT

Model

GLiClass-Edge v3.0 for LiteRT

1

3 commits

3 linked in READMEs

updated Oct 3, 2026

See the code

README

GLiClass-Edge v3.0 for LiteRT

knowledgator/gliclass-edge-v3.0 is a zero-shot text classifier with 32.7M parameters, built on a ModernBERT encoder. It reads a text and up to 25 candidate labels as one sequence and scores every label in one forward pass. This repository holds it converted to LiteRT with litert-torch 0.9.3. The verified path on desktop is the LiteRT CompiledModel Python API on CPU and, for the 128-token graph, on the Metal GPU with explicit FP32 computation (ai-edge-litert 2.1.6, Apple M4 Max). On a phone, it is the Kotlin CompiledModel API on the Galaxy S26 GPU with explicit FP32 computation, or on its CPU (LiteRT 2.2.0). All checks ran on 2026-10-02.

On desktop CPU, the shipped graphs with the float16 embedding table gave the same answers as the official gliclass 0.1.20 pipeline (CPU FP32) on all 552 test requests, in single-label and in multi-label mode. No logit differed by more than 0.009. On the Galaxy S26 GPU with explicit FP32, the Android sample in android/sample/ also gave the official answers on all 552 requests, from text to labels. Its graph call took a median of 5.86 ms at 128 tokens and 8.23 ms at 256 tokens.

An invented support ticket on the left; the softmax of its five label scores and the labels with sigmoid at or above 0.5 on the right

The ticket is invented. The numbers are what gliclass_edge_v3_s128_fp32.tflite with the float16 table returned for it on desktop CPU. The official pipeline gives the same values to 3 decimals.

Other formats:

Files

FileBytesRole
gliclass_edge_v3_s128_fp32.tflite53,703,480Graph for up to 128 tokens (labels, prompt and text). The examples below use it.
gliclass_edge_v3_s256_fp32.tflite53,965,476Graph for up to 256 tokens.
gliclass_edge_v3_s128_wfp16.tflite27,012,624The 128-token graph with float16 weights.
gliclass_edge_v3_s256_wfp16.tflite27,274,624The 256-token graph with float16 weights.
host_assets/tok_embeddings_fp16.bin38,684,160Token-embedding table [50370, 384], little-endian float16, for the host lookup.
host_assets/tokenizer.json3,583,595The source repository's tokenizer.json, unchanged.
gliclass_litert.pyPython host: builds the request string, tokenizes, looks up the table rows, runs the graph and returns the pipeline's output.
requirements.txtnumpy, tokenizers, ai-edge-litert==2.1.6 for the Python host.
examples/run_example.pyThe Python example below.
conversion/Conversion, fixture, check and figure scripts, with the environment lock.
fixtures/152 test requests with the official outputs, and the pins that rebuild the other 400.
android/sample/Android sample app (Compose) with a pure-Kotlin host that ran on the Galaxy S26: tokenizer, request builder, float16 table lookup, LiteRT call, softmax and sigmoid.
android/CardSnippet.ktThe Kotlin block below, in the package it was compiled in.
REPRODUCE.mdHow to reproduce the files and the checks.
LICENSE, NOTICE, licenses/Apache License 2.0 text, attribution and the retained upstream licenses.
assets/hero.pngThe figure above.
SHA256SUMSSHA-256 checksums of the files.

The two fp32 graphs and host_assets/ are 149,936,711 bytes. The wfp16 graphs are half the size. They store the 78 FULLY_CONNECTED weights in float16, each behind a DEQUANTIZE operator, and keep activations in float32. On the 552 test requests they change 2 answers, both near ties (see below).

Each graph holds the 10-layer encoder, the read-out at [CLS] and at every <<LABEL>> token, the two projectors and the scorer. It returns one logit per label slot. The host does the rest: it builds the request string, tokenizes, looks up each token's table row, builds the label routing and applies softmax or sigmoid.

Minimal usage

Python: desktop CPU

Needs numpy, tokenizers and ai-edge-litert (tested with 2.1.6). PyTorch, transformers and the gliclass package are not needed. Run it from the directory that holds the downloaded files. examples/run_example.py is the same code.

from gliclass_litert import GliclassLiteRT

text = (
    "The Zorvik X2 router from Tallowmere Networks keeps dropping the connection every evening "
    "around 9 pm. I have restarted it and reset it to factory settings, but the status light "
    "still blinks orange."
)
labels = ["connectivity problem", "hardware defect", "billing issue", "feature request",
          "installation help"]

with GliclassLiteRT(".") as classifier:
    for mode in ("single-label", "multi-label"):
        result = classifier.classify(text, labels, classification_type=mode)
        print(mode, [(r["label"], round(r["score"], 3)) for r in result])

On a Mac CPU it prints the following. The code rounds each score to 3 decimals.

single-label [('connectivity problem', 0.856)]
multi-label [('connectivity problem', 0.995), ('hardware defect', 0.95), ('billing issue', 0.886), ('feature request', 0.811)]

Single-label mode picks connectivity problem with softmax 0.856. Multi-label mode keeps every label whose sigmoid is at least 0.5, here four of the five. On the same request (61 tokens), the official pipeline on CPU FP32 returns the same labels and the same scores to 3 decimals.

classify() returns the pipeline's output: a list of {"label", "score"}. Pass a task description as prompt=; it goes right before the text. threshold= sets the multi-label threshold. For the desktop GPU, pass accelerator="gpu", which sets explicit FP32 (GpuOptions(enforce_f32=True)). The graph passed that way on the Metal GPU (482 of 482 requests). This host's own GPU path has not been run. To use the float16-weight graphs, pass their paths in the mapping form of the constructor, described in its docstring.

Kotlin: Android GPU with explicit FP32

import android.util.Half
import com.google.ai.edge.litert.Accelerator
import com.google.ai.edge.litert.CompiledModel
import java.io.File
import java.io.RandomAccessFile
import java.nio.ByteOrder
import java.nio.channels.FileChannel
import kotlin.math.exp

/** One GLiClass-Edge window (128 or 256 tokens) on the GPU with explicit FP32. */
class GliclassGpu(dir: File, private val window: Int = 128) : AutoCloseable {
  private val options = CompiledModel.Options(Accelerator.GPU).apply {
    gpuOptions = CompiledModel.GpuOptions(precision = CompiledModel.GpuOptions.Precision.FP32)
  }
  private val model =
    CompiledModel.create(File(dir, "gliclass_edge_v3_s${window}_fp32.tflite").absolutePath, options, null)
  private val inputs = listOf("inputs_embeds", "attention_mask", "label_routing")
    .associateWith { model.createInputBuffer(it, "serving_default") }
  private val outputs = mapOf("logits" to model.createOutputBuffer("logits", "serving_default"))
  private val table = RandomAccessFile(File(dir, "host_assets/tok_embeddings_fp16.bin"), "r").use {
    it.channel.map(FileChannel.MapMode.READ_ONLY, 0, it.length()).order(ByteOrder.LITTLE_ENDIAN).asShortBuffer()
  }

  /** ids = [CLS] … [SEP] of the linearized request; labelPositions = index of each <<LABEL>> token. */
  fun logits(ids: IntArray, labelPositions: IntArray): FloatArray {
    require(ids.size <= window && labelPositions.size <= 25)
    val embeds = FloatArray(window * 384)
    for (p in 0 until window) {
      val row = (if (p < ids.size) ids[p] else 50283) * 384 // 50283 = [PAD]
      for (c in 0 until 384) embeds[p * 384 + c] = Half.toFloat(table.get(row + c))
    }
    val routing = FloatArray(25 * window)
    labelPositions.forEachIndexed { k, position -> routing[k * window + position] = 1f }
    inputs.getValue("inputs_embeds").writeFloat(embeds)
    inputs.getValue("attention_mask").writeFloat(FloatArray(window) { if (it < ids.size) 1f else 0f })
    inputs.getValue("label_routing").writeFloat(routing)
    model.run(inputs, outputs, "serving_default")
    return outputs.getValue("logits").readFloat().copyOf(labelPositions.size)
  }

  override fun close() {
    (inputs.values + outputs.values).forEach { it.close() }
    model.close()
  }
}

/** Single-label: softmax over the labels, the top one wins. */
fun softmax(logits: FloatArray): List<Double> {
  val top = logits.max()
  val e = logits.map { exp((it - top).toDouble()) }
  val sum = e.sum()
  return e.map { it / sum }
}

/** Multi-label: every label whose sigmoid reaches the threshold. */
fun chosen(logits: FloatArray, threshold: Double = 0.5): List<Int> =
  logits.indices.filter { 1 / (1 + exp(-logits[it].toDouble())) >= threshold }

Create GliclassGpu once per window and call logits from one worker thread. ids are the token IDs of the request string with [CLS] and [SEP]. labelPositions holds the index of each <<LABEL>> token. Both must come from a tokenizer that matches gliclass_litert.py. The block compiles with LiteRT 2.2.0 (AGP 8.9.1, Kotlin 2.2.21). It has not itself run on a phone; the sample below has.

The full Kotlin host that ran on the Galaxy S26 is in android/sample/, with its own README for download, build and install. Its GliclassTokenizer.kt turns text into IDs with tokenizer.json, without native code. GliclassInputs.kt builds the request and the inputs, and GliclassDecoder.kt reproduces the pipeline's float32 softmax and sigmoid.

Host contract

gliclass_litert.py implements this contract. It follows ZeroShotClassificationPipeline of gliclass 0.1.20 with a uni-encoder model.

  1. Build the request string: <<LABEL>>label 1<<LABEL>>label 2…<<SEP>>, then the prompt, then the text. Nothing separates the prompt from the text, not even a space.
  2. Tokenize the string with tokenizer.json: [CLS] at the start, [SEP] at the end, no truncation. The IDs are <<LABEL>> 50368, <<SEP>> 50369, [CLS] 50281, [SEP] 50282 and [PAD] 50283. The ByteLevel pre-tokenizer adds a leading space to each piece between added tokens, so the prompt starts with a space token. Text glued to the end of a prompt does not.
  3. Pick the smallest window, 128 or 256 tokens, that holds the IDs. A request with more than 256 tokens or more than 25 labels raises an error. Nothing is truncated.
  4. Pad the IDs with [PAD] to the window S. inputs_embeds [1,S,384] holds the table row of every position, padding included, widened from float16 to float32. attention_mask [1,S] is 1 for real tokens and 0 for padding. In label_routing [1,25,S], row k is 1 at the k-th <<LABEL>> token and 0 elsewhere. Rows past the last label stay zero.
  5. Run the serving_default signature. It returns logits [1,1,1,25], one logit per label slot. Slots 0 to n−1 belong to the n labels; ignore the rest.
  6. Single-label: softmax over the n logits, and the top label wins (the earlier label on a tie). Multi-label: sigmoid of each logit, and every label at or above the threshold (0.5 by default) is returned, in input order. The list can be empty.

On all 552 test requests, gliclass_litert.py builds the same string, token IDs and <<LABEL>> positions as the official pipeline. With the shipped fp32 graphs and the float16 table on desktop CPU, it returns the same single-label and multi-label answers on all 552.

Measured quality and performance

The reference is the official gliclass 0.1.20 pipeline on CPU FP32 (torch 2.12.1, transformers 5.17.0), one request per call. "Same top label" compares the single-label answer, the softmax top-1. "Same label set" compares the multi-label answer, the labels with sigmoid at or above 0.5.

The 552 test requests are all 144 items of SemIf authored144.jsonl (three options each; the item's question is the prompt), 200 ag_news test rows, 200 banking77 test rows, the source model card's two examples and six invented requests. banking77 has 77 intents and the graph holds 25 labels, so each banking77 request carries the gold intent and 24 others drawn with a fixed seed. 482 requests fit 128 tokens; the other 70 need 256. For 419 of the 552, the reference's top softmax score is below 0.9.

The desktop rows ran on an Apple M4 Max with ai-edge-litert 2.1.6 through the Python CompiledModel API. Other jobs shared the machine, so desktop times are informational. The phone rows ran on one Galaxy S26 (SM-S942Q, Android 16) with LiteRT 2.2.0 through the Kotlin CompiledModel API, on USB power, in the foreground, at thermal status 0 or 1 (each table says which). A phone median is the middle value of the sorted times; for an even count it is the upper of the two middle values.

Desktop

WhereGraph, tableRequestsSame top labelSame label setMax logit differenceMedian ms
CPUS128 fp32, float16 table (shipped)4824824820.00908.77
CPUS128 wfp16, float16 table (shipped)4824814810.0179.25
CPUS128 fp32 †, float32 table4824824820.00003621.21
CPUS256 fp32, float16 table (shipped)5525525520.009011.04
CPUS256 wfp16, float16 table (shipped)5525515510.01711.35
CPU, Python host from text to labelsS128 and S256 fp32, float16 table (shipped)5525525520.0090
Metal GPU, explicit FP32S128 fp32, float16 table (shipped)4824824820.00912.77
Metal GPU, default precisionS128 fp32 and wfp16, float16 table (shipped)4824794690.1352.29 / 2.34
Metal GPU, default precisionS128 fp32 †, float16 table48225813310.02.26

† The same graph without the SafeLayerNorm rescale (see the fp16 section). With FP32 computation it returns bit-identical logits to the shipped graph: on all 482 S128 requests on desktop CPU, for the fp32 and the wfp16 form, and on the S26 GPU.

All logits were finite in every desktop row. On the shipped S128 fp32 row, the largest softmax difference is 0.0015 and the largest sigmoid difference 0.0018. The CPU timings come from 8-thread runs on a loaded machine.

Galaxy S26, gate app

A debug gate app fed each graph the reference token IDs, so these rows measure the graph and the table, not a tokenizer. Times are warm medians of writing the inputs, running and reading back, after 5 warm-up requests; the table lookup is not included. The explicit-FP32 and CPU rows ran at thermal status 0 with the battery at 37 to 38 °C. The default-precision and NPU rows ran earlier the same day at thermal status 1 with the battery at 40 to 41 °C.

AcceleratorGraph, float16 tableRequestsSame top labelSame label setMax logit differenceCompile msWarm median msDelegation
GPU, explicit FP32S128 fp32 (shipped)4824824820.00915485.80670 of 670 nodes, 1 partition
GPU, explicit FP32S128 wfp16 (shipped)4824814810.0175545.84748 of 748, 1 partition
GPU, explicit FP32S256 fp32 (shipped)5525525520.00915748.32670 of 670, 1 partition
GPU, explicit FP32S256 wfp16 (shipped)5525515510.0175688.33748 of 748, 1 partition
CPU, 4 threadsS128 fp32 (shipped)4824824820.00901326.77670 of 670 XNNPACK
GPU, default precisionS128 fp32 and wfp16 (shipped)4824754550.37568 / 5923.86 / 3.85670 of 670 / 748 of 748, 1 partition
GPU, default precisionS128 fp32 †48226113810.18083.83655 of 655, 1 partition
NPU (Qualcomm, JIT, BURST)S128 wfp16 (shipped)4824804620.24978, JIT included2.51748 of 748 ops, 1 partition

In an earlier run at thermal status 1, the shipped S128 fp32 graph returned bit-identical logits to its † twin on all 482 requests (5.83 and 5.76 ms), and the shipped fp32 and wfp16 graphs returned bit-identical logits to each other at default precision. Both wfp16 rows with explicit FP32 differ on the same 2 requests as on desktop.

Galaxy S26, Android sample from text to labels

The sample in android/sample/ uses the shipped fp32 graphs and the float16 table, with its own Kotlin tokenizer.

  • Gate, debug build, 552 requests (482 at 128 tokens, 70 at 256): with the GPU at explicit FP32 and with the CPU at 4 threads, the on-device token IDs, <<LABEL>> positions and padding equal the official pipeline's on all 552. Single-label and multi-label answers match on all 552. Max logit difference 0.0091 on the GPU and 0.0090 on the CPU. Both windows run whole on the GPU delegate, in one partition.
  • Graph call (input write to output readback), median over the gate requests: GPU 5.86 ms at 128 tokens and 8.23 ms at 256; CPU 7.83 ms and 15.65 ms.
  • Cold start, benchmark build, GPU, the prefilled request (61 tokens, 5 labels): the app was ready 0.59 s after the process started. The request that followed took 8.05, 8.38 and 8.20 ms end to end in three cold starts (graph about 6.1 ms, tokenize and embed 1.9 to 2.1 ms).
  • Paced, benchmark build, 20 requests, one every 2 s: median end to end 13.15 ms on the GPU (graph 6.92 ms, tokenize and embed 5.90 ms) and 22.36 ms on the CPU.

Not measured: Pixel phones, other phones, sustained load, the NPU at 256 tokens.

Agreement with the dataset labels

The official pipeline's single-label answer equals the dataset label on 67 of 144 SemIf items, 115 of 200 ag_news rows and 71 of 200 banking77 rows (25-label subsets). These are our own subsets, not the publisher's benchmark. The shipped graphs give the same answers on all 552, so they score the same. For task accuracy, see the benchmark on the source model card.

fp16, the GPU default precision and the NPU

Run the GPU with explicit FP32 computation, or the CPU. At fp16 precision, LayerNorm overflows.

The residual stream grows large in the later layers. Over the 552 requests, a LayerNorm input reaches a distance of 2,064 from its row mean, and the row's sum of squared deviations reaches 6.4 million. The fp16 maximum is 65,504. The outputs stay finite, but the answers change. With unmodified LayerNorms, the Metal GPU at default precision keeps 258 of 482 top labels and 133 of 482 label sets. The S26 GPU at default precision keeps 261 and 138.

The shipped graphs compute 15 LayerNorms on x·2⁻ᵏ with eps·2⁻²ᵏ: k = 5 in layers 3 to 6, and k = 6 in layers 7 to 9 and the final norm. This SafeLayerNorm is exact in FP32: the logits stay bit-identical to the unmodified graph (the † rows). At fp16 precision it keeps most answers. The Metal GPU at default precision keeps 479 top labels and 469 label sets of 482. The S26 GPU at default precision keeps 475 and 455, at 3.86 ms. The S26 NPU runs the whole S128 wfp16 graph in one partition at 2.51 ms and keeps 480 and 462. Every answer that changes sits near the boundary: on the NPU, the changed top labels had reference gaps of at most 0.0043, and the changed sets a sigmoid at most 0.026 from 0.5.

Those fp16 paths are measured here, but they are not the verified mode, because they change answers. Explicit FP32 is GpuOptions(precision = FP32) in Kotlin and GpuOptions(enforce_f32=True) in Python.

The checkpoint is float32, so float16 storage also costs a little. The float16 table rounds values by up to 1.2e-4: with the float32 table the largest logit difference is 0.000036, with the float16 table 0.0090, and the answers are the same. In the wfp16 graphs, the converter has folded each LayerNorm scale into the next FULLY_CONNECTED weights (50 of the 78), and float16 rounds those products by up to 7.7e-4. That changes 2 answers of 552: banking77_015, whose top two logits differ by 0.0012, and banking77_064, whose sigmoid for one label sits 0.0002 below 0.5.

Limits

  • The agreement numbers measure how closely the conversion follows the official pipeline, not task accuracy. The conversion reproduces the official answers, including the wrong ones.
  • English text only.
  • At most 25 labels per request, and at most 256 tokens of labels, prompt and text together. A longer request raises an error. The official pipeline would truncate at 1,024 tokens.
  • The wfp16 graphs change 2 near-tie answers of the 552 test requests. The fp32 graphs change none.
  • The GPU at default precision and the NPU change some answers near the boundary (S26: 475 and 480 top labels of 482).
  • Few-shot examples and hierarchical labels of the pipeline are not ported.
  • One desktop and one phone, the Galaxy S26, were tested. Pixel phones, other phones and sustained load are not measured.

Provenance, conversion and license

  • Source: knowledgator/gliclass-edge-v3.0 at revision df03993a2ed98e5e4a0d2dd7efbbd105abe874cf. model.safetensors is 130,829,312 bytes, all F32, with 32,705,154 parameters, SHA-256 ed2600439be991eaab06d831b259e99057a69242acf3a26a61b053695b1b979c.
  • Model: a ModernBERT encoder from jhu-clsp/ettin-encoder-32m (MIT) with 10 layers, hidden size 384, 6 heads and a GeGLU FFN of 576. Layers 0, 3, 6 and 9 attend globally; the others see 64 tokens on each side. The <<LABEL>> hidden states go through one projector and the [CLS] hidden state through another. An MLP scorer (768 to 256 to 128 to 1) turns each pair into a logit.
  • Training data named on the source card: BioMike/formal-logic-reasoning-gliclass-2k (its card states no license), knowledgator/gliclass-v3-logic-dataset (Apache-2.0) and tau/commonsense_qa (MIT).
  • Conversion: litert-torch 0.9.3 with torch 2.12.1 and transformers 5.17.0, at fixed shapes. The graph covers the classification path: the host looks up the token embeddings, and the graph takes them with the attention mask and the label routing. The encoder is rewritten for fixed shapes: the fused QKV and FFN input projections are split into separate layers, rotary positions become one constant table, the sliding window becomes a constant band, the mask uses −1e4, and GELU stays exact.
  • litert-torch lowered the attention torch.matmul to rank-3 batch matmuls and the label read-out to a rank-2 one. A rank-4 marker, vendored from an earlier litert-community DeBERTa conversion, keeps all 21 BATCH_MATMUL operators at rank 4.
  • Each fp32 graph has 670 operators, and each wfp16 graph 748. None is GATHER, GATHER_ND, CAST, SELECT_V2, BROADCAST_TO or MAXIMUM, and no BATCH_MATMUL has a constant left operand. No tensor is int64, and none is above rank 4.
  • Float16 weights: ai-edge-quantizer 0.8.0, FLOAT_CASTING, weight-only.
  • Rewrite check: in PyTorch FP32, the rewritten graph gives the same answers as the official pipeline on all 552 requests, and its logits are within 4.4e-5.
  • Table: the checkpoint's token embeddings rounded to the nearest float16, with no inf or NaN. SHA-256 b22f9bd2aedb1768089da8126dc6cd69572b61f21e39d6254119773b5e147a27. tokenizer.json is the source file, unchanged.
  • Test data: fixtures/ holds the 152 requests whose text may be shared (SemIf, MIT; the source card's examples; the invented requests) with the official outputs. The ag_news rows (license "unknown" on its dataset card) and the banking77 rows (CC BY 4.0) are not copied. fixtures/ pins them by dataset revision and row index, and conversion/make_fixtures.py rebuilds them.
  • Verification: the LiteRT CompiledModel Python API on desktop CPU and Metal, and the Kotlin CompiledModel API on the Galaxy S26. Every check is the same top label and the same label set, plus the largest logit difference, against the official pipeline.

License: the source model is licensed under Apache 2.0 by Knowledgator. These converted files, the conversion scripts, the Python host, the Kotlin block and the Android sample are released under the same license. The encoder is MIT. The license text is in LICENSE, attribution is in NOTICE, and the retained upstream licenses are in licenses/.

gliclass
litert
modernbert
text-classification
tflite
zero-shot-classification

litert-community/GLiClass-Edge-v3.0-LiteRT

Model

GLiClass-Edge v3.0 for LiteRT

1

3 commits

3 linked in READMEs

updated Oct 3, 2026

See the code

README

GLiClass-Edge v3.0 for LiteRT

knowledgator/gliclass-edge-v3.0 is a zero-shot text classifier with 32.7M parameters, built on a ModernBERT encoder. It reads a text and up to 25 candidate labels as one sequence and scores every label in one forward pass. This repository holds it converted to LiteRT with litert-torch 0.9.3. The verified path on desktop is the LiteRT CompiledModel Python API on CPU and, for the 128-token graph, on the Metal GPU with explicit FP32 computation (ai-edge-litert 2.1.6, Apple M4 Max). On a phone, it is the Kotlin CompiledModel API on the Galaxy S26 GPU with explicit FP32 computation, or on its CPU (LiteRT 2.2.0). All checks ran on 2026-10-02.

On desktop CPU, the shipped graphs with the float16 embedding table gave the same answers as the official gliclass 0.1.20 pipeline (CPU FP32) on all 552 test requests, in single-label and in multi-label mode. No logit differed by more than 0.009. On the Galaxy S26 GPU with explicit FP32, the Android sample in android/sample/ also gave the official answers on all 552 requests, from text to labels. Its graph call took a median of 5.86 ms at 128 tokens and 8.23 ms at 256 tokens.

An invented support ticket on the left; the softmax of its five label scores and the labels with sigmoid at or above 0.5 on the right

The ticket is invented. The numbers are what gliclass_edge_v3_s128_fp32.tflite with the float16 table returned for it on desktop CPU. The official pipeline gives the same values to 3 decimals.

Other formats:

Files

FileBytesRole
gliclass_edge_v3_s128_fp32.tflite53,703,480Graph for up to 128 tokens (labels, prompt and text). The examples below use it.
gliclass_edge_v3_s256_fp32.tflite53,965,476Graph for up to 256 tokens.
gliclass_edge_v3_s128_wfp16.tflite27,012,624The 128-token graph with float16 weights.
gliclass_edge_v3_s256_wfp16.tflite27,274,624The 256-token graph with float16 weights.
host_assets/tok_embeddings_fp16.bin38,684,160Token-embedding table [50370, 384], little-endian float16, for the host lookup.
host_assets/tokenizer.json3,583,595The source repository's tokenizer.json, unchanged.
gliclass_litert.pyPython host: builds the request string, tokenizes, looks up the table rows, runs the graph and returns the pipeline's output.
requirements.txtnumpy, tokenizers, ai-edge-litert==2.1.6 for the Python host.
examples/run_example.pyThe Python example below.
conversion/Conversion, fixture, check and figure scripts, with the environment lock.
fixtures/152 test requests with the official outputs, and the pins that rebuild the other 400.
android/sample/Android sample app (Compose) with a pure-Kotlin host that ran on the Galaxy S26: tokenizer, request builder, float16 table lookup, LiteRT call, softmax and sigmoid.
android/CardSnippet.ktThe Kotlin block below, in the package it was compiled in.
REPRODUCE.mdHow to reproduce the files and the checks.
LICENSE, NOTICE, licenses/Apache License 2.0 text, attribution and the retained upstream licenses.
assets/hero.pngThe figure above.
SHA256SUMSSHA-256 checksums of the files.

The two fp32 graphs and host_assets/ are 149,936,711 bytes. The wfp16 graphs are half the size. They store the 78 FULLY_CONNECTED weights in float16, each behind a DEQUANTIZE operator, and keep activations in float32. On the 552 test requests they change 2 answers, both near ties (see below).

Each graph holds the 10-layer encoder, the read-out at [CLS] and at every <<LABEL>> token, the two projectors and the scorer. It returns one logit per label slot. The host does the rest: it builds the request string, tokenizes, looks up each token's table row, builds the label routing and applies softmax or sigmoid.

Minimal usage

Python: desktop CPU

Needs numpy, tokenizers and ai-edge-litert (tested with 2.1.6). PyTorch, transformers and the gliclass package are not needed. Run it from the directory that holds the downloaded files. examples/run_example.py is the same code.

from gliclass_litert import GliclassLiteRT

text = (
    "The Zorvik X2 router from Tallowmere Networks keeps dropping the connection every evening "
    "around 9 pm. I have restarted it and reset it to factory settings, but the status light "
    "still blinks orange."
)
labels = ["connectivity problem", "hardware defect", "billing issue", "feature request",
          "installation help"]

with GliclassLiteRT(".") as classifier:
    for mode in ("single-label", "multi-label"):
        result = classifier.classify(text, labels, classification_type=mode)
        print(mode, [(r["label"], round(r["score"], 3)) for r in result])

On a Mac CPU it prints the following. The code rounds each score to 3 decimals.

single-label [('connectivity problem', 0.856)]
multi-label [('connectivity problem', 0.995), ('hardware defect', 0.95), ('billing issue', 0.886), ('feature request', 0.811)]

Single-label mode picks connectivity problem with softmax 0.856. Multi-label mode keeps every label whose sigmoid is at least 0.5, here four of the five. On the same request (61 tokens), the official pipeline on CPU FP32 returns the same labels and the same scores to 3 decimals.

classify() returns the pipeline's output: a list of {"label", "score"}. Pass a task description as prompt=; it goes right before the text. threshold= sets the multi-label threshold. For the desktop GPU, pass accelerator="gpu", which sets explicit FP32 (GpuOptions(enforce_f32=True)). The graph passed that way on the Metal GPU (482 of 482 requests). This host's own GPU path has not been run. To use the float16-weight graphs, pass their paths in the mapping form of the constructor, described in its docstring.

Kotlin: Android GPU with explicit FP32

import android.util.Half
import com.google.ai.edge.litert.Accelerator
import com.google.ai.edge.litert.CompiledModel
import java.io.File
import java.io.RandomAccessFile
import java.nio.ByteOrder
import java.nio.channels.FileChannel
import kotlin.math.exp

/** One GLiClass-Edge window (128 or 256 tokens) on the GPU with explicit FP32. */
class GliclassGpu(dir: File, private val window: Int = 128) : AutoCloseable {
  private val options = CompiledModel.Options(Accelerator.GPU).apply {
    gpuOptions = CompiledModel.GpuOptions(precision = CompiledModel.GpuOptions.Precision.FP32)
  }
  private val model =
    CompiledModel.create(File(dir, "gliclass_edge_v3_s${window}_fp32.tflite").absolutePath, options, null)
  private val inputs = listOf("inputs_embeds", "attention_mask", "label_routing")
    .associateWith { model.createInputBuffer(it, "serving_default") }
  private val outputs = mapOf("logits" to model.createOutputBuffer("logits", "serving_default"))
  private val table = RandomAccessFile(File(dir, "host_assets/tok_embeddings_fp16.bin"), "r").use {
    it.channel.map(FileChannel.MapMode.READ_ONLY, 0, it.length()).order(ByteOrder.LITTLE_ENDIAN).asShortBuffer()
  }

  /** ids = [CLS] … [SEP] of the linearized request; labelPositions = index of each <<LABEL>> token. */
  fun logits(ids: IntArray, labelPositions: IntArray): FloatArray {
    require(ids.size <= window && labelPositions.size <= 25)
    val embeds = FloatArray(window * 384)
    for (p in 0 until window) {
      val row = (if (p < ids.size) ids[p] else 50283) * 384 // 50283 = [PAD]
      for (c in 0 until 384) embeds[p * 384 + c] = Half.toFloat(table.get(row + c))
    }
    val routing = FloatArray(25 * window)
    labelPositions.forEachIndexed { k, position -> routing[k * window + position] = 1f }
    inputs.getValue("inputs_embeds").writeFloat(embeds)
    inputs.getValue("attention_mask").writeFloat(FloatArray(window) { if (it < ids.size) 1f else 0f })
    inputs.getValue("label_routing").writeFloat(routing)
    model.run(inputs, outputs, "serving_default")
    return outputs.getValue("logits").readFloat().copyOf(labelPositions.size)
  }

  override fun close() {
    (inputs.values + outputs.values).forEach { it.close() }
    model.close()
  }
}

/** Single-label: softmax over the labels, the top one wins. */
fun softmax(logits: FloatArray): List<Double> {
  val top = logits.max()
  val e = logits.map { exp((it - top).toDouble()) }
  val sum = e.sum()
  return e.map { it / sum }
}

/** Multi-label: every label whose sigmoid reaches the threshold. */
fun chosen(logits: FloatArray, threshold: Double = 0.5): List<Int> =
  logits.indices.filter { 1 / (1 + exp(-logits[it].toDouble())) >= threshold }

Create GliclassGpu once per window and call logits from one worker thread. ids are the token IDs of the request string with [CLS] and [SEP]. labelPositions holds the index of each <<LABEL>> token. Both must come from a tokenizer that matches gliclass_litert.py. The block compiles with LiteRT 2.2.0 (AGP 8.9.1, Kotlin 2.2.21). It has not itself run on a phone; the sample below has.

The full Kotlin host that ran on the Galaxy S26 is in android/sample/, with its own README for download, build and install. Its GliclassTokenizer.kt turns text into IDs with tokenizer.json, without native code. GliclassInputs.kt builds the request and the inputs, and GliclassDecoder.kt reproduces the pipeline's float32 softmax and sigmoid.

Host contract

gliclass_litert.py implements this contract. It follows ZeroShotClassificationPipeline of gliclass 0.1.20 with a uni-encoder model.

  1. Build the request string: <<LABEL>>label 1<<LABEL>>label 2…<<SEP>>, then the prompt, then the text. Nothing separates the prompt from the text, not even a space.
  2. Tokenize the string with tokenizer.json: [CLS] at the start, [SEP] at the end, no truncation. The IDs are <<LABEL>> 50368, <<SEP>> 50369, [CLS] 50281, [SEP] 50282 and [PAD] 50283. The ByteLevel pre-tokenizer adds a leading space to each piece between added tokens, so the prompt starts with a space token. Text glued to the end of a prompt does not.
  3. Pick the smallest window, 128 or 256 tokens, that holds the IDs. A request with more than 256 tokens or more than 25 labels raises an error. Nothing is truncated.
  4. Pad the IDs with [PAD] to the window S. inputs_embeds [1,S,384] holds the table row of every position, padding included, widened from float16 to float32. attention_mask [1,S] is 1 for real tokens and 0 for padding. In label_routing [1,25,S], row k is 1 at the k-th <<LABEL>> token and 0 elsewhere. Rows past the last label stay zero.
  5. Run the serving_default signature. It returns logits [1,1,1,25], one logit per label slot. Slots 0 to n−1 belong to the n labels; ignore the rest.
  6. Single-label: softmax over the n logits, and the top label wins (the earlier label on a tie). Multi-label: sigmoid of each logit, and every label at or above the threshold (0.5 by default) is returned, in input order. The list can be empty.

On all 552 test requests, gliclass_litert.py builds the same string, token IDs and <<LABEL>> positions as the official pipeline. With the shipped fp32 graphs and the float16 table on desktop CPU, it returns the same single-label and multi-label answers on all 552.

Measured quality and performance

The reference is the official gliclass 0.1.20 pipeline on CPU FP32 (torch 2.12.1, transformers 5.17.0), one request per call. "Same top label" compares the single-label answer, the softmax top-1. "Same label set" compares the multi-label answer, the labels with sigmoid at or above 0.5.

The 552 test requests are all 144 items of SemIf authored144.jsonl (three options each; the item's question is the prompt), 200 ag_news test rows, 200 banking77 test rows, the source model card's two examples and six invented requests. banking77 has 77 intents and the graph holds 25 labels, so each banking77 request carries the gold intent and 24 others drawn with a fixed seed. 482 requests fit 128 tokens; the other 70 need 256. For 419 of the 552, the reference's top softmax score is below 0.9.

The desktop rows ran on an Apple M4 Max with ai-edge-litert 2.1.6 through the Python CompiledModel API. Other jobs shared the machine, so desktop times are informational. The phone rows ran on one Galaxy S26 (SM-S942Q, Android 16) with LiteRT 2.2.0 through the Kotlin CompiledModel API, on USB power, in the foreground, at thermal status 0 or 1 (each table says which). A phone median is the middle value of the sorted times; for an even count it is the upper of the two middle values.

Desktop

WhereGraph, tableRequestsSame top labelSame label setMax logit differenceMedian ms
CPUS128 fp32, float16 table (shipped)4824824820.00908.77
CPUS128 wfp16, float16 table (shipped)4824814810.0179.25
CPUS128 fp32 †, float32 table4824824820.00003621.21
CPUS256 fp32, float16 table (shipped)5525525520.009011.04
CPUS256 wfp16, float16 table (shipped)5525515510.01711.35
CPU, Python host from text to labelsS128 and S256 fp32, float16 table (shipped)5525525520.0090
Metal GPU, explicit FP32S128 fp32, float16 table (shipped)4824824820.00912.77
Metal GPU, default precisionS128 fp32 and wfp16, float16 table (shipped)4824794690.1352.29 / 2.34
Metal GPU, default precisionS128 fp32 †, float16 table48225813310.02.26

† The same graph without the SafeLayerNorm rescale (see the fp16 section). With FP32 computation it returns bit-identical logits to the shipped graph: on all 482 S128 requests on desktop CPU, for the fp32 and the wfp16 form, and on the S26 GPU.

All logits were finite in every desktop row. On the shipped S128 fp32 row, the largest softmax difference is 0.0015 and the largest sigmoid difference 0.0018. The CPU timings come from 8-thread runs on a loaded machine.

Galaxy S26, gate app

A debug gate app fed each graph the reference token IDs, so these rows measure the graph and the table, not a tokenizer. Times are warm medians of writing the inputs, running and reading back, after 5 warm-up requests; the table lookup is not included. The explicit-FP32 and CPU rows ran at thermal status 0 with the battery at 37 to 38 °C. The default-precision and NPU rows ran earlier the same day at thermal status 1 with the battery at 40 to 41 °C.

AcceleratorGraph, float16 tableRequestsSame top labelSame label setMax logit differenceCompile msWarm median msDelegation
GPU, explicit FP32S128 fp32 (shipped)4824824820.00915485.80670 of 670 nodes, 1 partition
GPU, explicit FP32S128 wfp16 (shipped)4824814810.0175545.84748 of 748, 1 partition
GPU, explicit FP32S256 fp32 (shipped)5525525520.00915748.32670 of 670, 1 partition
GPU, explicit FP32S256 wfp16 (shipped)5525515510.0175688.33748 of 748, 1 partition
CPU, 4 threadsS128 fp32 (shipped)4824824820.00901326.77670 of 670 XNNPACK
GPU, default precisionS128 fp32 and wfp16 (shipped)4824754550.37568 / 5923.86 / 3.85670 of 670 / 748 of 748, 1 partition
GPU, default precisionS128 fp32 †48226113810.18083.83655 of 655, 1 partition
NPU (Qualcomm, JIT, BURST)S128 wfp16 (shipped)4824804620.24978, JIT included2.51748 of 748 ops, 1 partition

In an earlier run at thermal status 1, the shipped S128 fp32 graph returned bit-identical logits to its † twin on all 482 requests (5.83 and 5.76 ms), and the shipped fp32 and wfp16 graphs returned bit-identical logits to each other at default precision. Both wfp16 rows with explicit FP32 differ on the same 2 requests as on desktop.

Galaxy S26, Android sample from text to labels

The sample in android/sample/ uses the shipped fp32 graphs and the float16 table, with its own Kotlin tokenizer.

  • Gate, debug build, 552 requests (482 at 128 tokens, 70 at 256): with the GPU at explicit FP32 and with the CPU at 4 threads, the on-device token IDs, <<LABEL>> positions and padding equal the official pipeline's on all 552. Single-label and multi-label answers match on all 552. Max logit difference 0.0091 on the GPU and 0.0090 on the CPU. Both windows run whole on the GPU delegate, in one partition.
  • Graph call (input write to output readback), median over the gate requests: GPU 5.86 ms at 128 tokens and 8.23 ms at 256; CPU 7.83 ms and 15.65 ms.
  • Cold start, benchmark build, GPU, the prefilled request (61 tokens, 5 labels): the app was ready 0.59 s after the process started. The request that followed took 8.05, 8.38 and 8.20 ms end to end in three cold starts (graph about 6.1 ms, tokenize and embed 1.9 to 2.1 ms).
  • Paced, benchmark build, 20 requests, one every 2 s: median end to end 13.15 ms on the GPU (graph 6.92 ms, tokenize and embed 5.90 ms) and 22.36 ms on the CPU.

Not measured: Pixel phones, other phones, sustained load, the NPU at 256 tokens.

Agreement with the dataset labels

The official pipeline's single-label answer equals the dataset label on 67 of 144 SemIf items, 115 of 200 ag_news rows and 71 of 200 banking77 rows (25-label subsets). These are our own subsets, not the publisher's benchmark. The shipped graphs give the same answers on all 552, so they score the same. For task accuracy, see the benchmark on the source model card.

fp16, the GPU default precision and the NPU

Run the GPU with explicit FP32 computation, or the CPU. At fp16 precision, LayerNorm overflows.

The residual stream grows large in the later layers. Over the 552 requests, a LayerNorm input reaches a distance of 2,064 from its row mean, and the row's sum of squared deviations reaches 6.4 million. The fp16 maximum is 65,504. The outputs stay finite, but the answers change. With unmodified LayerNorms, the Metal GPU at default precision keeps 258 of 482 top labels and 133 of 482 label sets. The S26 GPU at default precision keeps 261 and 138.

The shipped graphs compute 15 LayerNorms on x·2⁻ᵏ with eps·2⁻²ᵏ: k = 5 in layers 3 to 6, and k = 6 in layers 7 to 9 and the final norm. This SafeLayerNorm is exact in FP32: the logits stay bit-identical to the unmodified graph (the † rows). At fp16 precision it keeps most answers. The Metal GPU at default precision keeps 479 top labels and 469 label sets of 482. The S26 GPU at default precision keeps 475 and 455, at 3.86 ms. The S26 NPU runs the whole S128 wfp16 graph in one partition at 2.51 ms and keeps 480 and 462. Every answer that changes sits near the boundary: on the NPU, the changed top labels had reference gaps of at most 0.0043, and the changed sets a sigmoid at most 0.026 from 0.5.

Those fp16 paths are measured here, but they are not the verified mode, because they change answers. Explicit FP32 is GpuOptions(precision = FP32) in Kotlin and GpuOptions(enforce_f32=True) in Python.

The checkpoint is float32, so float16 storage also costs a little. The float16 table rounds values by up to 1.2e-4: with the float32 table the largest logit difference is 0.000036, with the float16 table 0.0090, and the answers are the same. In the wfp16 graphs, the converter has folded each LayerNorm scale into the next FULLY_CONNECTED weights (50 of the 78), and float16 rounds those products by up to 7.7e-4. That changes 2 answers of 552: banking77_015, whose top two logits differ by 0.0012, and banking77_064, whose sigmoid for one label sits 0.0002 below 0.5.

Limits

  • The agreement numbers measure how closely the conversion follows the official pipeline, not task accuracy. The conversion reproduces the official answers, including the wrong ones.
  • English text only.
  • At most 25 labels per request, and at most 256 tokens of labels, prompt and text together. A longer request raises an error. The official pipeline would truncate at 1,024 tokens.
  • The wfp16 graphs change 2 near-tie answers of the 552 test requests. The fp32 graphs change none.
  • The GPU at default precision and the NPU change some answers near the boundary (S26: 475 and 480 top labels of 482).
  • Few-shot examples and hierarchical labels of the pipeline are not ported.
  • One desktop and one phone, the Galaxy S26, were tested. Pixel phones, other phones and sustained load are not measured.

Provenance, conversion and license

  • Source: knowledgator/gliclass-edge-v3.0 at revision df03993a2ed98e5e4a0d2dd7efbbd105abe874cf. model.safetensors is 130,829,312 bytes, all F32, with 32,705,154 parameters, SHA-256 ed2600439be991eaab06d831b259e99057a69242acf3a26a61b053695b1b979c.
  • Model: a ModernBERT encoder from jhu-clsp/ettin-encoder-32m (MIT) with 10 layers, hidden size 384, 6 heads and a GeGLU FFN of 576. Layers 0, 3, 6 and 9 attend globally; the others see 64 tokens on each side. The <<LABEL>> hidden states go through one projector and the [CLS] hidden state through another. An MLP scorer (768 to 256 to 128 to 1) turns each pair into a logit.
  • Training data named on the source card: BioMike/formal-logic-reasoning-gliclass-2k (its card states no license), knowledgator/gliclass-v3-logic-dataset (Apache-2.0) and tau/commonsense_qa (MIT).
  • Conversion: litert-torch 0.9.3 with torch 2.12.1 and transformers 5.17.0, at fixed shapes. The graph covers the classification path: the host looks up the token embeddings, and the graph takes them with the attention mask and the label routing. The encoder is rewritten for fixed shapes: the fused QKV and FFN input projections are split into separate layers, rotary positions become one constant table, the sliding window becomes a constant band, the mask uses −1e4, and GELU stays exact.
  • litert-torch lowered the attention torch.matmul to rank-3 batch matmuls and the label read-out to a rank-2 one. A rank-4 marker, vendored from an earlier litert-community DeBERTa conversion, keeps all 21 BATCH_MATMUL operators at rank 4.
  • Each fp32 graph has 670 operators, and each wfp16 graph 748. None is GATHER, GATHER_ND, CAST, SELECT_V2, BROADCAST_TO or MAXIMUM, and no BATCH_MATMUL has a constant left operand. No tensor is int64, and none is above rank 4.
  • Float16 weights: ai-edge-quantizer 0.8.0, FLOAT_CASTING, weight-only.
  • Rewrite check: in PyTorch FP32, the rewritten graph gives the same answers as the official pipeline on all 552 requests, and its logits are within 4.4e-5.
  • Table: the checkpoint's token embeddings rounded to the nearest float16, with no inf or NaN. SHA-256 b22f9bd2aedb1768089da8126dc6cd69572b61f21e39d6254119773b5e147a27. tokenizer.json is the source file, unchanged.
  • Test data: fixtures/ holds the 152 requests whose text may be shared (SemIf, MIT; the source card's examples; the invented requests) with the official outputs. The ag_news rows (license "unknown" on its dataset card) and the banking77 rows (CC BY 4.0) are not copied. fixtures/ pins them by dataset revision and row index, and conversion/make_fixtures.py rebuilds them.
  • Verification: the LiteRT CompiledModel Python API on desktop CPU and Metal, and the Kotlin CompiledModel API on the Galaxy S26. Every check is the same top label and the same label set, plus the largest logit difference, against the official pipeline.

License: the source model is licensed under Apache 2.0 by Knowledgator. These converted files, the conversion scripts, the Python host, the Kotlin block and the Android sample are released under the same license. The encoder is MIT. The license text is in LICENSE, attribution is in NOTICE, and the retained upstream licenses are in licenses/.

gliclass
litert
modernbert
text-classification
tflite
zero-shot-classification