litert-community/Laya-English-LiteRT

Model

Laya English for LiteRT — Android GPU and NPU

1

4 commits

3 linked in READMEs

updated Sep 29, 2026

See the code

README

Laya English for LiteRT — Android GPU and NPU

Run the English checkpoint of convaiinnovations/laya and its typed-decisions/ fine-tune on a phone GPU or NPU with LiteRT. Laya reads a text or a JSON state and answers questions you define at request time: pick one of several options, score on an ordinal scale, or give a yes/no probability. Each question is one forward pass of the main graph plus one pass of a small act-head graph. Validated with LiteRT 2.2.0 CompiledModel on a Samsung Galaxy S26 (SM-S942Q, Android 16): at 256 tokens, 66 ms of graph time per question for the English graph and 73 ms for the typed-decisions graph on the Hexagon NPU, 123 ms and 126 ms on the GPU with explicit FP32 computation. On 140 English and 100 typed-decisions question rows the answers match the official laya 0.3.4 fp32 CPU implementation: the same argmax on every choice and score question, and a maximum probability difference of 0.0008 on the GPU and 0.0072 on the NPU. Other Android GPUs and NPUs have not been validated here.

Which of the three Laya packages to take. All three were tested on a Galaxy S26 GPU with LiteRT 2.2.0 at explicit FP32; the times below are graph time per question at a 256-token window, and the first two packages add the app's table lookup (about 15–20 ms):

  • Laya-Multilingual-LiteRT: the multilingual checkpoint (mmBERT-base; the publisher lists 100+ languages; validated in English and Japanese). 0.68 GB on the phone, 51 ms per question. Has an Android sample app with a Kotlin tokenizer (android/; also zero_shot_classification in litert-samples). Take this one for non-English or mixed text; it answers English too, close to the English checkpoint (Convai's XNLI English: 0.843 against 0.860).
  • Laya-English-LiteRT: the English checkpoint (ModernBERT-large) and its typed-decisions/ fine-tune for four workflows (customer service, invoice processing, security incidents, agent-trace observability). 0.85 GB each, 123 ms per question (126 ms for the fine-tune); 66 ms (73 ms) on the NPU. No app of its own; the Kotlin GPU function on its card takes token ids from the app's own tokenizer. Take this one for English-only text when the higher English score matters, and its fine-tune for those four workflows.
  • laya-LiteRT: the same two checkpoints (not the fine-tune) in the token-id form: token ids in, embedding table inside the graph (the two packages above take embedding rows that the app looks up from a table file). Its GPU files keep fp32 weights: 1.3 GB (multilingual) or 1.7 GB (English), 54 ms (multilingual) or 127 ms (English) per question with no lookup to add. No app of its own; its card links an Android SDK with Kotlin tokenizers for both checkpoints. Take this one to feed token ids and skip the table lookup in the app.

This repository is the English package.

An invented support message and the five answers the English graph returned

The message is invented. The answers and probabilities are the values the English S256 wfp16 graph returned through laya_host.py on a desktop CPU, with the checkpoint's own temperatures. This package has no app of its own.

Files and supported configuration

English files are at the repository root; typed-decisions/ holds the fine-tune, the same layout as the source repository at revision 1c5edc17a7acd8701df6fc341c0d179f1c62c982. SHA256SUMS lists every other file.

English

FileWindow NBytesRole
laya_en_s256_embeds_wfp16.tflite256739,765,104recommended; validated on the S26 GPU, NPU and CPU
laya_en_s256_embeds_fp32.tflite2561,478,482,924fp32 reference; validated on the S26 GPU
laya_en_s512_embeds_wfp16.tflite512740,813,680longer texts; validated on the S26 GPU and NPU
laya_en_s512_embeds_fp32.tflite5121,479,531,500fp32 reference; validated on the S26 GPU
laya_en_act_head_fp32.tfliteany1,057,960act head, shared by every English main graph
token_embeddings_en_fp16.bin, token_embeddings_en_fp16.json103,153,664[50368,1024] float16 token table for the host lookup
tokenizer.json, tokenizer_config.json3,583,536exact files from the pinned checkpoint
laya_en_config.json745the checkpoint's rl_agent_config.json: temperatures and builder budgets
laya_host.py22,933Python host for both checkpoints: prompt builder, embedding lookup, decoder
fixtures/gate_rows_en_s256.json517,382the 140 validation rows with the reference answers

typed-decisions/

FileWindow NBytesRole
laya_td_s256_embeds_wfp16.tflite256739,765,104recommended; validated on the S26 GPU and NPU
laya_td_s256_embeds_fp32.tflite2561,478,482,924fp32 reference; validated on the S26 GPU
laya_td_s512_embeds_wfp16.tflite512740,813,680longer texts; validated on the S26 GPU and NPU
laya_td_s512_embeds_fp32.tflite5121,479,531,500fp32 reference; validated on the S26 GPU
laya_td_act_head_fp32.tfliteany1,057,960act head, shared by every typed-decisions main graph
token_embeddings_td_fp16.bin, token_embeddings_td_fp16.json103,153,664[50368,1024] float16 token table for the host lookup
tokenizer.json, tokenizer_config.json3,583,565exact files from the pinned checkpoint
laya_td_config.json847the checkpoint's rl_agent_config.json
fixtures/gate_rows_td_s256.json450,460the 100 validation rows with the reference answers

The recommended English phone set is the S256 wfp16 graph, the act head, the token table pair, the tokenizer pair and the config file: 847,582,366 bytes. Each checkpoint keeps its own table, tokenizer pair and config; the two tables differ.

wfp16 files store the 123 FULLY_CONNECTED weight tensors as float16 with a DEQUANTIZE to float32. Every activation and every other constant stays float32. The float16 tables reproduce the checkpoints' float32 tables exactly (maximum difference 0 over all 51,576,832 values of each). Padded positions gather the real PAD row (id 50283).

The main graph holds the ModernBERT-large encoder, Laya's two typed head layers and the option scorer, applied at every position. The host does the rest: tokenize, build the prompt with one [MASK] marker per option, look up the embedding rows, read the logits at the marker positions, and apply the checkpoint temperature and softmax. The act head is a separate small graph because its input depends on those host-side probabilities.

The graphs take embedding rows, not token ids. With the token table inside the graph, the float16-weight file does not compile on LiteRT 2.2.0 CompiledModel GPU: the GPU delegate rejects the EMBEDDING_LOOKUP that reads a DEQUANTIZE-fed table (Empty quantization params), and CompiledModel needs every operator on the GPU. The fp32 file with the table inside does compile; litert-community/laya-LiteRT ships that form. With the lookup on the host, the wfp16 graphs compile as one GPU partition (2279 of 2279 operators) at 0.74 GB.

The graphs are safe to compute in fp16 (revision of 2026-09-29). The NPU computes in fp16, and ModernBERT-large carries values up to 3.3e4 in its residual stream from layer 20 on, whose squares overflow fp16 inside LayerNorm. Each large LayerNorm therefore computes on its input scaled by a power of two, with epsilon scaled by the square, and the attention masks use −1e4 instead of −1e9 (which is −inf in fp16 and turns a zero mask weight into NaN). Both rewrites give identical results in fp32: the graphs before and after agree bit for bit in PyTorch and give the same answers on every validation row.

On the GPU, use explicit FP32 computation, as in the Kotlin block below; it is the only GPU precision validated here. On the NPU, the same wfp16 file compiles on the phone through LiteRT's Qualcomm plugin (JIT): 31–53 s on the first launch, cached afterwards. The app must package the Qualcomm runtime libraries listed in Laya-Multilingual-LiteRT's android/README.md. INT8 files were not built. The two attention layer types use different RoPE tables (theta 160000 for full attention, 10000 for sliding attention); the conversion keeps them distinct.

Minimal usage

Python — complete pipeline, desktop CPU

Needs numpy, transformers (tokenizer only) and ai-edge-litert; neither torch nor the laya package is required. Set USE_TF=0. Put these eight files in one directory: laya_host.py, laya_en_s256_embeds_wfp16.tflite, laya_en_act_head_fp32.tflite, token_embeddings_en_fp16.bin, token_embeddings_en_fp16.json, tokenizer.json, tokenizer_config.json, laya_en_config.json.

import json
from pathlib import Path
from laya_host import LayaHost

assets = Path(".")
questions = {
    "intent": {
        "type": "choice",
        "instructions": "What does the customer need?",
        "criteria": {
            "refund": "return a duplicate payment",
            "help": "technical help",
        },
    }
}
with LayaHost.from_directory(
    assets, checkpoint="en", window=256, storage="wfp16"
) as host:
    result = host.predict(
        "I was charged twice for one order. Please return the extra payment.",
        questions,
    )
print(json.dumps(result, indent=2))

Run from a directory that held only those eight files, this exact code printed:

"choice": "refund", "probabilities": {"refund": 0.9792, "help": 0.0208},
"confidence": 0.854, "action": {"act_probability": 1.0}, "usage": {"input_tokens": 39, ...}

checkpoint="td" with the repository root selects the fine-tune (the host reads typed-decisions/); window=512 selects the S512 graph without changing the checkpoint's prompt-builder budgets; storage="fp32" selects the fp32 file.

Kotlin — Android GPU with explicit FP32

The ModernBERT ByteLevel BPE tokenizer is not ported to Kotlin in this package: the token ids and marker positions must come from the app's own tokenizer and prompt builder (HOST_CONTRACT.md specifies both). The multilingual sample's Kotlin tokenizer covers only that checkpoint. This function does the rest: it memory-maps the float16 table, gathers the rows for the given ids (real PAD rows for the padding), runs the main graph with GPU FP32 computation and returns the temperature-scaled option probabilities. dir is the checkpoint directory. Reuse the mapping, the model and the buffers in an application; the act-head features and the answer dictionaries are in HOST_CONTRACT.md.

import android.util.Half
import com.google.ai.edge.litert.Accelerator
import com.google.ai.edge.litert.CompiledModel
import com.google.ai.edge.litert.Environment
import com.google.ai.edge.litert.TensorBuffer
import java.io.File
import java.io.FileInputStream
import java.nio.ByteOrder
import java.nio.channels.FileChannel
import org.json.JSONObject
import kotlin.math.exp

// Call on a worker thread. ids and markers follow HOST_CONTRACT.md's builder.
fun scoreRow(dir: File, ids: IntArray, markers: IntArray, qtype: Int,
             tag: String = "en", window: Int = 256): FloatArray {
    require(tag in setOf("en", "td") && window in setOf(256, 512))
    require(ids.isNotEmpty() && ids.size <= window && qtype in 0..2)
    require(markers.isNotEmpty() && markers.all { it in ids.indices })
    val meta = JSONObject(File(dir, "token_embeddings_${tag}_fp16.json").readText())
    val shape = meta.getJSONArray("shape")
    val vocab = shape.getInt(0)
    val width = shape.getInt(1)
    val pad = meta.getInt("pad_id")
    require(meta.getString("dtype") == "float16" && meta.getString("byte_order") == "little")
    require(meta.getString("layout") == "row-major")
    require(pad in 0 until vocab && ids.all { it in 0 until vocab })
    val tableFile = File(dir, meta.getString("filename"))
    require(tableFile.length() == vocab.toLong() * width * 2)
    require(tableFile.length() == meta.getLong("size_bytes"))
    val table = FileInputStream(tableFile).channel.use { channel ->
        channel.map(FileChannel.MapMode.READ_ONLY, 0, tableFile.length())
            .order(ByteOrder.LITTLE_ENDIAN)
    }
    val embeds = FloatArray(window * width) { index ->
        val token = ids.getOrElse(index / width) { pad }
        Half.toFloat(table.getShort((token * width + index % width) * 2))
    }
    val mask = FloatArray(window) { if (it < ids.size) 1f else 0f }
    val type = FloatArray(3).also { it[qtype] = 1f }
    val config = JSONObject(File(dir, "laya_${tag}_config.json").readText())
    val count = markers.size
    val bucket = when { count <= 2 -> "2"; count <= 5 -> "3-5"; count <= 10 -> "6-10"; else -> "11+" }
    val key = "${listOf("choice", "score", "noul")[qtype]}:$bucket"
    val temperature = config.getJSONObject("temperature_by_options")
        .optDouble(key, config.getJSONArray("temperature").getDouble(qtype))
        .toFloat().coerceAtLeast(1e-3f)
    val options = CompiledModel.Options(Accelerator.GPU).apply {
        gpuOptions = CompiledModel.GpuOptions(
            precision = CompiledModel.GpuOptions.Precision.FP32)
    }
    Environment.create().use { environment ->
        CompiledModel.create(File(dir, "laya_${tag}_s${window}_embeds_wfp16.tflite").absolutePath,
                             options, environment).use { model ->
            val inputs = linkedMapOf<String, TensorBuffer>()
            val outputs = linkedMapOf<String, TensorBuffer>()
            try {
                listOf("inputs_embeds", "attention_mask", "qtype_onehot").forEach {
                    inputs[it] = model.createInputBuffer(it, "serving_default")
                }
                listOf("token_logits", "pooled_cls").forEach {
                    outputs[it] = model.createOutputBuffer(it, "serving_default")
                }
                inputs.getValue("inputs_embeds").writeFloat(embeds)
                inputs.getValue("attention_mask").writeFloat(mask)
                inputs.getValue("qtype_onehot").writeFloat(type)
                model.run(inputs, outputs, "serving_default")
                val logits = outputs.getValue("token_logits").readFloat()
                val pooled = outputs.getValue("pooled_cls").readFloat()
                require(logits.all { it.isFinite() } && pooled.all { it.isFinite() })
                val scores = FloatArray(count) { logits[markers[it]] / temperature }
                val peak = scores.max()
                val weights = FloatArray(count) { exp((scores[it] - peak).toDouble()).toFloat() }
                val total = weights.sum()
                return FloatArray(count) { weights[it] / total }
            } finally {
                (inputs.values + outputs.values).forEach { it.close() }
            }
        }
    }
}

For the NPU, create the Environment with DispatchLibraryDir and CompilerPluginLibraryDir set to context.applicationInfo.nativeLibraryDir, use CompiledModel.Options(Accelerator.NPU) with QualcommOptions(htpPerformanceMode = BURST) for the main graph, and keep the act head on the CPU.

Validation

Reference: the official laya 0.3.4 predict on CPU fp32 at checkpoint revision 1c5edc17a7acd8701df6fc341c0d179f1c62c982, every question of a fixture in one call, the original builder budgets (English 512 / 192, typed-decisions 1024 / 256) and each checkpoint's own config temperatures. The fixtures are invented: 46 English fixtures (209 question rows) and 40 typed-decisions fixtures (200 rows, ten per upstream workflow). A row is evaluated at a window only when its original sequence fits it.

All 1,248 row executions on the desktop CPU and all 1,248 on the S26 GPU were finite. Argmax is identical on every choice and score row. Maximum |Δp| covers every option probability, the yes/no probability and act_probability against the reference four-decimal dictionaries; the limits are 1e-3 for fp32 and 1e-2 for wfp16.

CheckpointNStorageDesktop CPU rowsArgmaxMax ΔpS26 GPU rowsArgmaxMax Δp
English256wfp1614059/590.00063333253914059/590.000631127167
English256fp3214059/596.15482807e-0514059/594.9820447e-05
English512wfp1620987/870.00076218571720987/870.00076248374
English512fp3220987/876.15482807e-0520987/875.03234029e-05
typed-decisions256wfp1610065/650.00049436836210065/650.000495828676
typed-decisions256fp3210065/655.11034489e-0510065/655.15206814e-05
typed-decisions512wfp16175112/1120.000494368362175112/1120.000495828676
typed-decisions512fp32175112/1125.11034489e-05175112/1125.15206814e-05

Desktop CPU: ai-edge-litert 2.1.6 CompiledModel, four threads, Apple M4 Max, macOS 27.0. Phone: the validated configuration above. The phone runs consumed the captured ids and marker positions of the same rows and their raw outputs were decoded with the same Python arithmetic; the phone validation therefore covers the graphs, the table lookup and the act head, not a Kotlin tokenizer. The English S256 wfp16 graph also passed on the S26 CPU (XNNPACK, four threads; 140 rows, 59/59, max Δp 0.000633).

Every main-graph run and every act-head run compiled to one GPU partition. Verbatim from the PID-filtered logcat, one line per graph kind:

wfp16 main: Replacing 2223 out of 2223 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 0 (main).
fp32 main:  Replacing 2100 out of 2100 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 0 (main).
act:        Replacing 4 out of 4 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 0 (main).

No unsupported-operation line appeared. The GPU rows above were measured on 2026-09-22 with the graphs before the fp16 rewrites, which give the same fp32 results. On 2026-09-29 the rewritten S256 wfp16 graphs gave on the S26 GPU FP32 59/59 with max Δp 0.00063 (English, 2279/2279 operators) and 65/65 with 0.00049 (typed-decisions), and on the desktop CPU the same values as the table to the fourth decimal.

On the S26 NPU (Hexagon v81, JIT, the whole main graph in one DispatchDelegate node, the act head on the CPU), 2026-09-29, rewritten wfp16 graphs:

CheckpointNNPU rowsArgmaxMax Δp
English25614059/590.0072
English51220987/870.0072
typed-decisions25610065/650.0034
typed-decisions512175112/1120.0034

The S256 validation rows, with state, questions, ids, markers and reference answers, are in fixtures/gate_rows_en_s256.json (140 rows) and typed-decisions/fixtures/gate_rows_td_s256.json (100 rows).

Galaxy S26 timings

Measured on 2026-09-22 on the held phone, screen on, Android build S942QOPS1AZF2_SJP1AZF2, LiteRT 2.2.0, GPU FP32, one process per run. Main+act covers the input writes, run() and the output readback of both graphs; the host table lookup is separate. Cold is the first row of the run; warm is the median over the remaining rows. The eight GPU runs were executed back to back and the battery temperature rose from 34.5 °C to 44.1 °C over the sequence; the S512 runs and the later fp32 runs were measured warm, which is why their medians sit above their cold rows. These are observations from one device and one thermal sequence, not a latency benchmark.

CheckpointNStorageCompile msCold main+act msWarm median [min, max] msLookup median ms
English256wfp162711121.2122.9 [117.8, 126.8]20.2
English256fp323491124.4164.0 [122.5, 209.0]23.7
English512wfp162481289.6615.6 [322.7, 636.4]28.3
English512fp323315296.6616.1 [295.2, 630.1]31.5
typed-decisions256wfp162528122.2125.9 [117.8, 207.9]19.7
typed-decisions256fp324081238.1266.6 [223.4, 291.0]24.4
typed-decisions512wfp164664631.4618.5 [581.7, 632.9]28.9
typed-decisions512fp325823632.1622.6 [604.6, 681.7]30.2

For comparison on the same phone, the English S256 wfp16 graph on the CPU (XNNPACK, four threads) took 582.7 ms warm median [336.2, 588.1] per question.

NPU, 2026-09-29, rewritten wfp16 graphs, same definitions, battery 33–36 °C before each run. In every NPU run the slowest row was within 7 % of the fastest, while the English S256 GPU FP32 run of the same afternoon moved between 120 and 253 ms as the phone warmed.

CheckpointNFirst-launch compile sWarm median [min, max] ms
English25630.965.7 [64.3, 68.6]
English51250.0167.2 [165.7, 171.9]
typed-decisions25632.472.8 [71.4, 75.2]
typed-decisions51252.6173.7 [170.3, 176.1]

Limits

  • Questions with more than 20 options are outside what was validated.
  • The NPU was validated on one Snapdragon 8 Elite Gen 5 phone (Hexagon v81). The first NPU launch compiles for 31–53 s; later launches load the cached result.
  • There is no S1024 graph. A row longer than 512 tokens after the checkpoint's own builder needs a shorter input; the graph window does not change the builder budget.
  • act_probability was 1.0 on every reference row of both checkpoints. The act graph and its formula are unchanged; this saturation says nothing about escalation quality.
  • The typed-decisions fixtures use the four upstream workflow question-id sets (agent trace observability, customer service, invoice processing, security incidents) with invented wording; laya 0.3.4 defines the id sets, not full question presets.
  • Numerical agreement with the official implementation is not task accuracy for a new schema. Validate the question wording for the application.

License and attribution

Laya is by Convai Innovations, built on ModernBERT-large; the checkpoints and the host logic are Apache-2.0, as are ModernBERT, Transformers and tokenizers. NOTICE and licenses/ hold the retained texts. The tokenizer files and the two config files are exact bytes from the pinned checkpoint directories (the configs were renamed; their contents are unchanged). laya_host.py adapts the Laya 0.3.4 prompt builder and decoder; the conversion adds fixed windows, explicit attention and RoPE, a separate act graph and the host token lookup, and since 2026-09-29 the fp16-safe rewrites described above, which are exact in fp32.

android
laya
litert
modernbert
text-classification
tflite
zero-shot-classification

litert-community/Laya-English-LiteRT

Model

Laya English for LiteRT — Android GPU and NPU

1

4 commits

3 linked in READMEs

updated Sep 29, 2026

See the code

README

Laya English for LiteRT — Android GPU and NPU

Run the English checkpoint of convaiinnovations/laya and its typed-decisions/ fine-tune on a phone GPU or NPU with LiteRT. Laya reads a text or a JSON state and answers questions you define at request time: pick one of several options, score on an ordinal scale, or give a yes/no probability. Each question is one forward pass of the main graph plus one pass of a small act-head graph. Validated with LiteRT 2.2.0 CompiledModel on a Samsung Galaxy S26 (SM-S942Q, Android 16): at 256 tokens, 66 ms of graph time per question for the English graph and 73 ms for the typed-decisions graph on the Hexagon NPU, 123 ms and 126 ms on the GPU with explicit FP32 computation. On 140 English and 100 typed-decisions question rows the answers match the official laya 0.3.4 fp32 CPU implementation: the same argmax on every choice and score question, and a maximum probability difference of 0.0008 on the GPU and 0.0072 on the NPU. Other Android GPUs and NPUs have not been validated here.

Which of the three Laya packages to take. All three were tested on a Galaxy S26 GPU with LiteRT 2.2.0 at explicit FP32; the times below are graph time per question at a 256-token window, and the first two packages add the app's table lookup (about 15–20 ms):

  • Laya-Multilingual-LiteRT: the multilingual checkpoint (mmBERT-base; the publisher lists 100+ languages; validated in English and Japanese). 0.68 GB on the phone, 51 ms per question. Has an Android sample app with a Kotlin tokenizer (android/; also zero_shot_classification in litert-samples). Take this one for non-English or mixed text; it answers English too, close to the English checkpoint (Convai's XNLI English: 0.843 against 0.860).
  • Laya-English-LiteRT: the English checkpoint (ModernBERT-large) and its typed-decisions/ fine-tune for four workflows (customer service, invoice processing, security incidents, agent-trace observability). 0.85 GB each, 123 ms per question (126 ms for the fine-tune); 66 ms (73 ms) on the NPU. No app of its own; the Kotlin GPU function on its card takes token ids from the app's own tokenizer. Take this one for English-only text when the higher English score matters, and its fine-tune for those four workflows.
  • laya-LiteRT: the same two checkpoints (not the fine-tune) in the token-id form: token ids in, embedding table inside the graph (the two packages above take embedding rows that the app looks up from a table file). Its GPU files keep fp32 weights: 1.3 GB (multilingual) or 1.7 GB (English), 54 ms (multilingual) or 127 ms (English) per question with no lookup to add. No app of its own; its card links an Android SDK with Kotlin tokenizers for both checkpoints. Take this one to feed token ids and skip the table lookup in the app.

This repository is the English package.

An invented support message and the five answers the English graph returned

The message is invented. The answers and probabilities are the values the English S256 wfp16 graph returned through laya_host.py on a desktop CPU, with the checkpoint's own temperatures. This package has no app of its own.

Files and supported configuration

English files are at the repository root; typed-decisions/ holds the fine-tune, the same layout as the source repository at revision 1c5edc17a7acd8701df6fc341c0d179f1c62c982. SHA256SUMS lists every other file.

English

FileWindow NBytesRole
laya_en_s256_embeds_wfp16.tflite256739,765,104recommended; validated on the S26 GPU, NPU and CPU
laya_en_s256_embeds_fp32.tflite2561,478,482,924fp32 reference; validated on the S26 GPU
laya_en_s512_embeds_wfp16.tflite512740,813,680longer texts; validated on the S26 GPU and NPU
laya_en_s512_embeds_fp32.tflite5121,479,531,500fp32 reference; validated on the S26 GPU
laya_en_act_head_fp32.tfliteany1,057,960act head, shared by every English main graph
token_embeddings_en_fp16.bin, token_embeddings_en_fp16.json103,153,664[50368,1024] float16 token table for the host lookup
tokenizer.json, tokenizer_config.json3,583,536exact files from the pinned checkpoint
laya_en_config.json745the checkpoint's rl_agent_config.json: temperatures and builder budgets
laya_host.py22,933Python host for both checkpoints: prompt builder, embedding lookup, decoder
fixtures/gate_rows_en_s256.json517,382the 140 validation rows with the reference answers

typed-decisions/

FileWindow NBytesRole
laya_td_s256_embeds_wfp16.tflite256739,765,104recommended; validated on the S26 GPU and NPU
laya_td_s256_embeds_fp32.tflite2561,478,482,924fp32 reference; validated on the S26 GPU
laya_td_s512_embeds_wfp16.tflite512740,813,680longer texts; validated on the S26 GPU and NPU
laya_td_s512_embeds_fp32.tflite5121,479,531,500fp32 reference; validated on the S26 GPU
laya_td_act_head_fp32.tfliteany1,057,960act head, shared by every typed-decisions main graph
token_embeddings_td_fp16.bin, token_embeddings_td_fp16.json103,153,664[50368,1024] float16 token table for the host lookup
tokenizer.json, tokenizer_config.json3,583,565exact files from the pinned checkpoint
laya_td_config.json847the checkpoint's rl_agent_config.json
fixtures/gate_rows_td_s256.json450,460the 100 validation rows with the reference answers

The recommended English phone set is the S256 wfp16 graph, the act head, the token table pair, the tokenizer pair and the config file: 847,582,366 bytes. Each checkpoint keeps its own table, tokenizer pair and config; the two tables differ.

wfp16 files store the 123 FULLY_CONNECTED weight tensors as float16 with a DEQUANTIZE to float32. Every activation and every other constant stays float32. The float16 tables reproduce the checkpoints' float32 tables exactly (maximum difference 0 over all 51,576,832 values of each). Padded positions gather the real PAD row (id 50283).

The main graph holds the ModernBERT-large encoder, Laya's two typed head layers and the option scorer, applied at every position. The host does the rest: tokenize, build the prompt with one [MASK] marker per option, look up the embedding rows, read the logits at the marker positions, and apply the checkpoint temperature and softmax. The act head is a separate small graph because its input depends on those host-side probabilities.

The graphs take embedding rows, not token ids. With the token table inside the graph, the float16-weight file does not compile on LiteRT 2.2.0 CompiledModel GPU: the GPU delegate rejects the EMBEDDING_LOOKUP that reads a DEQUANTIZE-fed table (Empty quantization params), and CompiledModel needs every operator on the GPU. The fp32 file with the table inside does compile; litert-community/laya-LiteRT ships that form. With the lookup on the host, the wfp16 graphs compile as one GPU partition (2279 of 2279 operators) at 0.74 GB.

The graphs are safe to compute in fp16 (revision of 2026-09-29). The NPU computes in fp16, and ModernBERT-large carries values up to 3.3e4 in its residual stream from layer 20 on, whose squares overflow fp16 inside LayerNorm. Each large LayerNorm therefore computes on its input scaled by a power of two, with epsilon scaled by the square, and the attention masks use −1e4 instead of −1e9 (which is −inf in fp16 and turns a zero mask weight into NaN). Both rewrites give identical results in fp32: the graphs before and after agree bit for bit in PyTorch and give the same answers on every validation row.

On the GPU, use explicit FP32 computation, as in the Kotlin block below; it is the only GPU precision validated here. On the NPU, the same wfp16 file compiles on the phone through LiteRT's Qualcomm plugin (JIT): 31–53 s on the first launch, cached afterwards. The app must package the Qualcomm runtime libraries listed in Laya-Multilingual-LiteRT's android/README.md. INT8 files were not built. The two attention layer types use different RoPE tables (theta 160000 for full attention, 10000 for sliding attention); the conversion keeps them distinct.

Minimal usage

Python — complete pipeline, desktop CPU

Needs numpy, transformers (tokenizer only) and ai-edge-litert; neither torch nor the laya package is required. Set USE_TF=0. Put these eight files in one directory: laya_host.py, laya_en_s256_embeds_wfp16.tflite, laya_en_act_head_fp32.tflite, token_embeddings_en_fp16.bin, token_embeddings_en_fp16.json, tokenizer.json, tokenizer_config.json, laya_en_config.json.

import json
from pathlib import Path
from laya_host import LayaHost

assets = Path(".")
questions = {
    "intent": {
        "type": "choice",
        "instructions": "What does the customer need?",
        "criteria": {
            "refund": "return a duplicate payment",
            "help": "technical help",
        },
    }
}
with LayaHost.from_directory(
    assets, checkpoint="en", window=256, storage="wfp16"
) as host:
    result = host.predict(
        "I was charged twice for one order. Please return the extra payment.",
        questions,
    )
print(json.dumps(result, indent=2))

Run from a directory that held only those eight files, this exact code printed:

"choice": "refund", "probabilities": {"refund": 0.9792, "help": 0.0208},
"confidence": 0.854, "action": {"act_probability": 1.0}, "usage": {"input_tokens": 39, ...}

checkpoint="td" with the repository root selects the fine-tune (the host reads typed-decisions/); window=512 selects the S512 graph without changing the checkpoint's prompt-builder budgets; storage="fp32" selects the fp32 file.

Kotlin — Android GPU with explicit FP32

The ModernBERT ByteLevel BPE tokenizer is not ported to Kotlin in this package: the token ids and marker positions must come from the app's own tokenizer and prompt builder (HOST_CONTRACT.md specifies both). The multilingual sample's Kotlin tokenizer covers only that checkpoint. This function does the rest: it memory-maps the float16 table, gathers the rows for the given ids (real PAD rows for the padding), runs the main graph with GPU FP32 computation and returns the temperature-scaled option probabilities. dir is the checkpoint directory. Reuse the mapping, the model and the buffers in an application; the act-head features and the answer dictionaries are in HOST_CONTRACT.md.

import android.util.Half
import com.google.ai.edge.litert.Accelerator
import com.google.ai.edge.litert.CompiledModel
import com.google.ai.edge.litert.Environment
import com.google.ai.edge.litert.TensorBuffer
import java.io.File
import java.io.FileInputStream
import java.nio.ByteOrder
import java.nio.channels.FileChannel
import org.json.JSONObject
import kotlin.math.exp

// Call on a worker thread. ids and markers follow HOST_CONTRACT.md's builder.
fun scoreRow(dir: File, ids: IntArray, markers: IntArray, qtype: Int,
             tag: String = "en", window: Int = 256): FloatArray {
    require(tag in setOf("en", "td") && window in setOf(256, 512))
    require(ids.isNotEmpty() && ids.size <= window && qtype in 0..2)
    require(markers.isNotEmpty() && markers.all { it in ids.indices })
    val meta = JSONObject(File(dir, "token_embeddings_${tag}_fp16.json").readText())
    val shape = meta.getJSONArray("shape")
    val vocab = shape.getInt(0)
    val width = shape.getInt(1)
    val pad = meta.getInt("pad_id")
    require(meta.getString("dtype") == "float16" && meta.getString("byte_order") == "little")
    require(meta.getString("layout") == "row-major")
    require(pad in 0 until vocab && ids.all { it in 0 until vocab })
    val tableFile = File(dir, meta.getString("filename"))
    require(tableFile.length() == vocab.toLong() * width * 2)
    require(tableFile.length() == meta.getLong("size_bytes"))
    val table = FileInputStream(tableFile).channel.use { channel ->
        channel.map(FileChannel.MapMode.READ_ONLY, 0, tableFile.length())
            .order(ByteOrder.LITTLE_ENDIAN)
    }
    val embeds = FloatArray(window * width) { index ->
        val token = ids.getOrElse(index / width) { pad }
        Half.toFloat(table.getShort((token * width + index % width) * 2))
    }
    val mask = FloatArray(window) { if (it < ids.size) 1f else 0f }
    val type = FloatArray(3).also { it[qtype] = 1f }
    val config = JSONObject(File(dir, "laya_${tag}_config.json").readText())
    val count = markers.size
    val bucket = when { count <= 2 -> "2"; count <= 5 -> "3-5"; count <= 10 -> "6-10"; else -> "11+" }
    val key = "${listOf("choice", "score", "noul")[qtype]}:$bucket"
    val temperature = config.getJSONObject("temperature_by_options")
        .optDouble(key, config.getJSONArray("temperature").getDouble(qtype))
        .toFloat().coerceAtLeast(1e-3f)
    val options = CompiledModel.Options(Accelerator.GPU).apply {
        gpuOptions = CompiledModel.GpuOptions(
            precision = CompiledModel.GpuOptions.Precision.FP32)
    }
    Environment.create().use { environment ->
        CompiledModel.create(File(dir, "laya_${tag}_s${window}_embeds_wfp16.tflite").absolutePath,
                             options, environment).use { model ->
            val inputs = linkedMapOf<String, TensorBuffer>()
            val outputs = linkedMapOf<String, TensorBuffer>()
            try {
                listOf("inputs_embeds", "attention_mask", "qtype_onehot").forEach {
                    inputs[it] = model.createInputBuffer(it, "serving_default")
                }
                listOf("token_logits", "pooled_cls").forEach {
                    outputs[it] = model.createOutputBuffer(it, "serving_default")
                }
                inputs.getValue("inputs_embeds").writeFloat(embeds)
                inputs.getValue("attention_mask").writeFloat(mask)
                inputs.getValue("qtype_onehot").writeFloat(type)
                model.run(inputs, outputs, "serving_default")
                val logits = outputs.getValue("token_logits").readFloat()
                val pooled = outputs.getValue("pooled_cls").readFloat()
                require(logits.all { it.isFinite() } && pooled.all { it.isFinite() })
                val scores = FloatArray(count) { logits[markers[it]] / temperature }
                val peak = scores.max()
                val weights = FloatArray(count) { exp((scores[it] - peak).toDouble()).toFloat() }
                val total = weights.sum()
                return FloatArray(count) { weights[it] / total }
            } finally {
                (inputs.values + outputs.values).forEach { it.close() }
            }
        }
    }
}

For the NPU, create the Environment with DispatchLibraryDir and CompilerPluginLibraryDir set to context.applicationInfo.nativeLibraryDir, use CompiledModel.Options(Accelerator.NPU) with QualcommOptions(htpPerformanceMode = BURST) for the main graph, and keep the act head on the CPU.

Validation

Reference: the official laya 0.3.4 predict on CPU fp32 at checkpoint revision 1c5edc17a7acd8701df6fc341c0d179f1c62c982, every question of a fixture in one call, the original builder budgets (English 512 / 192, typed-decisions 1024 / 256) and each checkpoint's own config temperatures. The fixtures are invented: 46 English fixtures (209 question rows) and 40 typed-decisions fixtures (200 rows, ten per upstream workflow). A row is evaluated at a window only when its original sequence fits it.

All 1,248 row executions on the desktop CPU and all 1,248 on the S26 GPU were finite. Argmax is identical on every choice and score row. Maximum |Δp| covers every option probability, the yes/no probability and act_probability against the reference four-decimal dictionaries; the limits are 1e-3 for fp32 and 1e-2 for wfp16.

CheckpointNStorageDesktop CPU rowsArgmaxMax ΔpS26 GPU rowsArgmaxMax Δp
English256wfp1614059/590.00063333253914059/590.000631127167
English256fp3214059/596.15482807e-0514059/594.9820447e-05
English512wfp1620987/870.00076218571720987/870.00076248374
English512fp3220987/876.15482807e-0520987/875.03234029e-05
typed-decisions256wfp1610065/650.00049436836210065/650.000495828676
typed-decisions256fp3210065/655.11034489e-0510065/655.15206814e-05
typed-decisions512wfp16175112/1120.000494368362175112/1120.000495828676
typed-decisions512fp32175112/1125.11034489e-05175112/1125.15206814e-05

Desktop CPU: ai-edge-litert 2.1.6 CompiledModel, four threads, Apple M4 Max, macOS 27.0. Phone: the validated configuration above. The phone runs consumed the captured ids and marker positions of the same rows and their raw outputs were decoded with the same Python arithmetic; the phone validation therefore covers the graphs, the table lookup and the act head, not a Kotlin tokenizer. The English S256 wfp16 graph also passed on the S26 CPU (XNNPACK, four threads; 140 rows, 59/59, max Δp 0.000633).

Every main-graph run and every act-head run compiled to one GPU partition. Verbatim from the PID-filtered logcat, one line per graph kind:

wfp16 main: Replacing 2223 out of 2223 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 0 (main).
fp32 main:  Replacing 2100 out of 2100 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 0 (main).
act:        Replacing 4 out of 4 node(s) with delegate (LITERT_CL) node, yielding 1 partitions for subgraph 0 (main).

No unsupported-operation line appeared. The GPU rows above were measured on 2026-09-22 with the graphs before the fp16 rewrites, which give the same fp32 results. On 2026-09-29 the rewritten S256 wfp16 graphs gave on the S26 GPU FP32 59/59 with max Δp 0.00063 (English, 2279/2279 operators) and 65/65 with 0.00049 (typed-decisions), and on the desktop CPU the same values as the table to the fourth decimal.

On the S26 NPU (Hexagon v81, JIT, the whole main graph in one DispatchDelegate node, the act head on the CPU), 2026-09-29, rewritten wfp16 graphs:

CheckpointNNPU rowsArgmaxMax Δp
English25614059/590.0072
English51220987/870.0072
typed-decisions25610065/650.0034
typed-decisions512175112/1120.0034

The S256 validation rows, with state, questions, ids, markers and reference answers, are in fixtures/gate_rows_en_s256.json (140 rows) and typed-decisions/fixtures/gate_rows_td_s256.json (100 rows).

Galaxy S26 timings

Measured on 2026-09-22 on the held phone, screen on, Android build S942QOPS1AZF2_SJP1AZF2, LiteRT 2.2.0, GPU FP32, one process per run. Main+act covers the input writes, run() and the output readback of both graphs; the host table lookup is separate. Cold is the first row of the run; warm is the median over the remaining rows. The eight GPU runs were executed back to back and the battery temperature rose from 34.5 °C to 44.1 °C over the sequence; the S512 runs and the later fp32 runs were measured warm, which is why their medians sit above their cold rows. These are observations from one device and one thermal sequence, not a latency benchmark.

CheckpointNStorageCompile msCold main+act msWarm median [min, max] msLookup median ms
English256wfp162711121.2122.9 [117.8, 126.8]20.2
English256fp323491124.4164.0 [122.5, 209.0]23.7
English512wfp162481289.6615.6 [322.7, 636.4]28.3
English512fp323315296.6616.1 [295.2, 630.1]31.5
typed-decisions256wfp162528122.2125.9 [117.8, 207.9]19.7
typed-decisions256fp324081238.1266.6 [223.4, 291.0]24.4
typed-decisions512wfp164664631.4618.5 [581.7, 632.9]28.9
typed-decisions512fp325823632.1622.6 [604.6, 681.7]30.2

For comparison on the same phone, the English S256 wfp16 graph on the CPU (XNNPACK, four threads) took 582.7 ms warm median [336.2, 588.1] per question.

NPU, 2026-09-29, rewritten wfp16 graphs, same definitions, battery 33–36 °C before each run. In every NPU run the slowest row was within 7 % of the fastest, while the English S256 GPU FP32 run of the same afternoon moved between 120 and 253 ms as the phone warmed.

CheckpointNFirst-launch compile sWarm median [min, max] ms
English25630.965.7 [64.3, 68.6]
English51250.0167.2 [165.7, 171.9]
typed-decisions25632.472.8 [71.4, 75.2]
typed-decisions51252.6173.7 [170.3, 176.1]

Limits

  • Questions with more than 20 options are outside what was validated.
  • The NPU was validated on one Snapdragon 8 Elite Gen 5 phone (Hexagon v81). The first NPU launch compiles for 31–53 s; later launches load the cached result.
  • There is no S1024 graph. A row longer than 512 tokens after the checkpoint's own builder needs a shorter input; the graph window does not change the builder budget.
  • act_probability was 1.0 on every reference row of both checkpoints. The act graph and its formula are unchanged; this saturation says nothing about escalation quality.
  • The typed-decisions fixtures use the four upstream workflow question-id sets (agent trace observability, customer service, invoice processing, security incidents) with invented wording; laya 0.3.4 defines the id sets, not full question presets.
  • Numerical agreement with the official implementation is not task accuracy for a new schema. Validate the question wording for the application.

License and attribution

Laya is by Convai Innovations, built on ModernBERT-large; the checkpoints and the host logic are Apache-2.0, as are ModernBERT, Transformers and tokenizers. NOTICE and licenses/ hold the retained texts. The tokenizer files and the two config files are exact bytes from the pinned checkpoint directories (the configs were renamed; their contents are unchanged). laya_host.py adapts the Laya 0.3.4 prompt builder and decoder; the conversion adds fixed windows, explicit attention and RoPE, a separate act graph and the host token lookup, and since 2026-09-29 the fp16-safe rewrites described above, which are exact in fp32.

android
laya
litert
modernbert
text-classification
tflite
zero-shot-classification