litert-community/Open-Decision-DeBERTa-v3-Large-LiteRT

Model

Open Decision DeBERTa-v3-Large for LiteRT

0

4 commits

3 linked in READMEs

updated Oct 3, 2026

See the code

README

Open Decision DeBERTa-v3-Large for LiteRT

com-kotobalabs/open-jev-deberta-v3-large is a decision model built on DeBERTa-v3-large. It reads one state and several typed questions (choice, score, noul) and returns a calibrated probability distribution for each question from one forward pass. This repository holds it converted to LiteRT with litert-torch 0.9.3. The verified path on desktop is the LiteRT CompiledModel Python API on CPU and on the Apple Metal GPU with explicit FP32 computation (ai-edge-litert 2.1.6, Apple M4 Max, 2026-10-01). On a phone, it is the Kotlin CompiledModel API on the Galaxy S26 GPU with explicit FP32 computation (LiteRT 2.2.0, 2026-10-01 and 2026-10-02). On desktop CPU, the shipped graph (512-token window, float16 weights, float16 word table) and the author's implementation (CPU FP32) picked the same option on all 4,327 questions of 1,809 requests, and no probability differed by more than 0.0016 at temperature 1.05.

On the Galaxy S26 GPU, the same graph picked the same option on all 1,302 questions of 500 requests, with a maximum probability difference of 0.00083 and a median of 697.5 ms per request. The 256-token graph also picked the same option on all 1,074 questions of 419 requests, at a median of 407.6 ms. The Kotlin host in android/sample/ ran the fixture gate on the S26: its ids and spans equal the author's Collator on all 160 requests, and all 420 answers match the author's implementation.

An invented support ticket on the left and the three answers the S512 wfp16 graph returned for it on the right

The ticket is invented. The bars are the probabilities that deberta_v3_large_decision_s512_wfp16.tflite with the float16 table returned for it on desktop CPU.

Other formats: onnx-community/open-jev-deberta-v3-large-ONNX, an ONNX port for Transformers.js with fp32, fp16, q4 and q4f16 weights.

Files

FileBytesRole
deberta_v3_large_decision_s512_wfp16.tflite813,443,728Graph for a 512-token window, with float16 weights. The examples below use it. Tested on desktop CPU (1,809 requests), on the Metal GPU with explicit FP32 (100 requests) and on the Galaxy S26 GPU with explicit FP32 (500 requests).
deberta_v3_large_decision_s256_wfp16.tflite712,780,432Graph for a 256-token window, with float16 weights. Tested on desktop CPU (the 1,281 requests that fit 256 tokens), on the Metal GPU with explicit FP32 (100 requests) and on the Galaxy S26 GPU with explicit FP32 (419 requests).
word_embeddings_fp16.bin262,348,800Word table [128100, 1024], little-endian float16, for the host lookup.
tokenizer.json8,657,170The source repository's tokenizer.json, unchanged.
decision_litert.pyPython host: encodes the request, looks up the table rows, builds the routing inputs, runs the graph and reads out the answers.
conversion/Conversion, fixture, check and figure scripts, with requirements-lock.txt.
android/sample/Android sample app (Compose) with the pure-Kotlin host that ran on the Galaxy S26: tokenizer, sequence builder, float16 table lookup, LiteRT call and read-out, plus JVM parity tests and a debug fixture gate.
android/CardSnippet.ktThe Kotlin block below, in the package it was compiled in. The full sample is in android/sample/.
REPRODUCE.mdHow to reproduce the conversion and the checks.
LICENSE, NOTICEApache License 2.0 text and attribution.
assets/hero.pngThe figure above.
SHA256SUMSSHA-256 checksums of the files.

In the wfp16 graphs only the weights are float16. Each of the 146 FULLY_CONNECTED weights feeds a DEQUANTIZE operator, and activations stay float32. The 48 relative-position constants of the batch matmuls also stay float32.

The measurements below also use three float32 files that are not in this repository: the S256 and S512 fp32 graphs (1,323,008,720 and 1,423,672,016 bytes) and the float32 table (524,697,600 bytes, a bit-exact copy of the checkpoint's word embeddings).

Every graph holds the 24-layer encoder, the span means and the scoring head, and returns one logit per option slot. The host does the rest. It tokenizes, looks up each token's table row, builds the two routing inputs, applies softmax at temperature 1.05 and reads out the answers.

Minimal usage

Python: desktop CPU

Needs numpy, tokenizers and ai-edge-litert (tested with 2.1.6). Torch, transformers and the author's typed_decisions package are not needed. Run it from the directory that holds the downloaded files and decision_litert.py.

from decision_litert import DecisionModel

model = DecisionModel(
    "deberta_v3_large_decision_s512_wfp16.tflite",
    "word_embeddings_fp16.bin",
    "tokenizer.json",
)
state = (
    "Order 8841 was supposed to arrive on Monday and the tracking page still "
    "says 'label created'. I need it before Friday for a gift."
)
questions = [
    {
        "type": "choice",
        "instructions": "Which team should handle this message?",
        "options": ["billing", "technical support", "shipping", "account access", "sales"],
    },
    {
        "type": "score",
        "instructions": "How frustrated is the customer?",
        "options": ["calm", "mildly annoyed", "frustrated", "angry"],
    },
    {"type": "noul", "instructions": "The customer asks for a refund."},
]
team, frustration, refund = model.decide(state, questions)
print(team["choice"], round(team["confidence"], 3))
print(round(frustration["score"], 3), round(frustration["confidence"], 3))
print(round(refund["noul"], 3))
model.close()

On a Mac CPU it prints the following. The code rounds each value to 3 decimals.

shipping 0.785
1.861 0.642
0.124

shipping wins with probability 0.785. The score answer is the expected zero-based level, so 1.861 falls between "mildly annoyed" and "frustrated". Its confidence, 0.642, is the probability of "frustrated". The refund question gives p(yes) = 0.124. On the same request (80 tokens), the author's decide() on CPU FP32 gives the same three answers, and no probability differs by more than 6.6e-5.

decide() returns one dict per question, and the choice and score answers also carry probabilities, one per option. To run on the desktop GPU, pass accelerator="gpu". It sets explicit FP32 (GpuOptions(enforce_f32=True)), the Metal setting measured below.

Kotlin: Android GPU with explicit FP32

import android.util.Half
import com.google.ai.edge.litert.Accelerator
import com.google.ai.edge.litert.CompiledModel
import java.io.File
import java.io.RandomAccessFile
import java.nio.ByteOrder
import java.nio.channels.FileChannel
import kotlin.math.exp

/** One request = token ids plus, per option, the token range of its text and of its question's text. */
class Span(val start: Int, val end: Int)

class DecisionGpu(dir: File, private val seq: Int = 512, private val slots: Int = 128) : AutoCloseable {
  private val options = CompiledModel.Options(Accelerator.GPU).apply {
    gpuOptions = CompiledModel.GpuOptions(precision = CompiledModel.GpuOptions.Precision.FP32)
  }
  private val model =
    CompiledModel.create(File(dir, "deberta_v3_large_decision_s${seq}_wfp16.tflite").absolutePath, options, null)
  private val inputs = listOf("inputs_embeds", "attention_mask", "q_routing", "o_routing")
    .associateWith { model.createInputBuffer(it, "serving_default") }
  private val outputs = mapOf("logits" to model.createOutputBuffer("logits", "serving_default"))
  private val table = RandomAccessFile(File(dir, "word_embeddings_fp16.bin"), "r").use {
    it.channel.map(FileChannel.MapMode.READ_ONLY, 0, it.length()).order(ByteOrder.LITTLE_ENDIAN).asShortBuffer()
  }
  private val embeds = FloatArray(seq * 1024)

  /** ids from a tokenizer that matches decision_litert.py; one (question span, option span) pair per option, in order. */
  fun logits(ids: IntArray, pairs: List<Pair<Span, Span>>): FloatArray {
    require(ids.size <= seq && pairs.size <= slots)
    for (p in 0 until seq) {
      val row = (if (p < ids.size) ids[p] else 0) * 1024
      for (c in 0 until 1024) embeds[p * 1024 + c] = Half.toFloat(table.get(row + c))
    }
    val qRouting = FloatArray(slots * seq)
    val oRouting = FloatArray(slots * seq)
    pairs.forEachIndexed { j, (q, o) ->
      for (t in q.start until q.end) qRouting[j * seq + t] = 1f / (q.end - q.start)
      for (t in o.start until o.end) oRouting[j * seq + t] = 1f / (o.end - o.start)
    }
    inputs.getValue("inputs_embeds").writeFloat(embeds)
    inputs.getValue("attention_mask").writeFloat(FloatArray(seq) { if (it < ids.size) 1f else 0f })
    inputs.getValue("q_routing").writeFloat(qRouting)
    inputs.getValue("o_routing").writeFloat(oRouting)
    model.run(inputs, outputs, "serving_default")
    return outputs.getValue("logits").readFloat().copyOf(pairs.size)
  }

  /** Probabilities of one question's options at the shipped temperature 1.05. */
  fun probabilities(logits: FloatArray, from: Int, count: Int): List<Double> {
    val z = (from until from + count).map { logits[it] / 1.05 }
    val top = z.max()
    val e = z.map { exp(it - top) }
    val sum = e.sum()
    return e.map { it / sum }
  }

  override fun close() {
    (inputs.values + outputs.values).forEach { it.close() }
    model.close()
  }
}

Create DecisionGpu once and call logits from one worker thread. logits takes the ids of one request and one (question span, option span) pair per option, in request order, and returns one logit per option. The ids and spans must come from a tokenizer that matches decision_litert.py. The block compiles with LiteRT 2.2.0 (AGP 8.9.1, Kotlin 2.2.21). It has not itself run on a device; the sample host below has.

The full Kotlin host that ran on the Galaxy S26 is in android/sample/, with its own README for download, build and install. Its DecisionModel.kt makes the same LiteRT calls as the block above. Its DecisionTokenizer.kt turns text into ids with tokenizer.json, and its DecisionInputs.kt builds the sequence and the spans.

Host contract

decision_litert.py implements this contract. It follows the Collator in the author's typed_decisions/encoder.py and the read-out in typed_decisions/schema.py.

  1. Tokenize each text with tokenizer.json and no added special tokens, then assemble the ids in this order: [CLS] [STATE] state [Q] instructions [OPT] option [OPT] option ... [Q] ... [SEP]. The state keeps its first 256 tokens. The marker ids are [STATE] 128001, [Q] 128002 and [OPT] 128003. [CLS] is 1, [SEP] is 2 and [PAD] is 0. A noul question always has the options ["no", "yes"]. Record the token span of each question's text and of each option's text as half-open ranges, markers excluded.
  2. Check the limits. A choice question takes 2 to 255 options, and a score question 2 to 10 ordered levels. The ids must fit the graph's window (256 or 512 tokens), and all options of the request must fit the 128 option slots. A request that does not fit raises an error instead of being truncated. Only the state is cut, at 256 tokens.
  3. Pad the ids on the right with id 0 to the window S, and fill the four float32 inputs. inputs_embeds [1,S,1024] holds the table row of every position, padding included. attention_mask [1,S] is 1 for real tokens and 0 for padding. In q_routing [1,128,S], row j is 1/len over the text tokens of option j's question. In o_routing [1,128,S], row j is 1/len over option j's text tokens. Rows past the request's last option stay zero.
  4. Run the serving_default signature. It returns logits [1,1,1,128], one logit per option slot in request order. Ignore the slots past the last option. Split the logits by question and apply softmax at temperature 1.05 within each question's options.
  5. Read the answer. choice is the option with the highest probability, and confidence is that probability. score is Σ i·pᵢ, the expected zero-based level, which may fall between levels. It also carries confidence. noul is p(yes).

On all 1,809 fixture requests, decision_litert.py produces the same ids and text spans as the author's Collator. Its decide() also gives the same answers as the author's decide() on 209 requests: the README example, the 8 invented tickets and the first 200 requests from data-fam. With the S512 wfp16 graph and the float16 table on desktop CPU, probabilities differ by at most 0.0016 and score values by at most 0.00066. With the S256 wfp16 graph, probabilities differ by at most 0.00081.

The tokenizers library on tokenizer.json gives the same ids as the author's AutoTokenizer on all 2,368 fixture texts. The SentencePiece spm.model gives the same ids on 2,367 of them, so build the ids from tokenizer.json.

Measured quality and performance

The reference is the author's implementation: the checkpoint's own typed_decisions package on CPU FP32 (torch 2.12.1, transformers 4.57.6, eager attention, 8 threads), one request per forward pass. Probabilities are compared at temperature 1.05. "Same argmax" means the same winning option for a question.

The fixtures are 1,809 requests with 4,327 questions. They are the first 1,500 states of the author's public data-fam/test.jsonl, the first 300 states of data-fam/ood-test.jsonl, the author's README example and 8 invented tickets with 4 questions each. Both test files come from the author's kotoba-lang/typed-decisions repository at revision e65215db7e554a31f557936f0bd1d8a0a78b6af5. A boundary question is one where the reference's top probability is below 0.9. There are 1,830 of them among the 4,327 questions, and 1,602 among the 2,796 questions that fit 256 tokens.

The desktop rows ran on 2026-10-01 on one desktop: Apple M4 Max (16 cores, 128 GB), macOS 27.0, ai-edge-litert 2.1.6 through the Python CompiledModel API. The Galaxy S26 rows ran on 2026-10-01 and 2026-10-02 on one Galaxy S26 (SM-S942Q, Android 16) with LiteRT 2.2.0 through the Kotlin CompiledModel API. They used a debug gate app, except the row marked as the sample in android/sample/. Times are informational. Each CPU row states the load on the machine. Each phone row states the thermal status and battery temperature where they were recorded, and the compile time. Phone times are warm medians of writing the inputs, running and reading back, with [min, max] where recorded.

WhereGraph, word tableRequestsQuestions (boundary)Same argmaxMax probability differenceQuestions over 0.01Median ms per request
CPU, 8 threads, 3 jobs in parallelS256 fp32, float32 table1,2812,796 (1,602)2,796 (1,602)1.1e-50152.8
CPU, 8 threads, 3 jobs in parallelS256 wfp16, float16 table1,2812,796 (1,602)2,796 (1,602)0.000810305.8
CPU, 8 threads, 2 to 3 jobs in parallelS256 fp32, float16 table1,2812,796 (1,602)2,796 (1,602)0.000240135.9
CPU, 8 threads, 2 to 3 jobs in parallelS256 wfp16, float32 table1,2812,796 (1,602)2,796 (1,602)0.000800350.2
CPU, 8 threads, 3 jobs in parallelS512 fp32, float32 table1,8094,327 (1,830)4,327 (1,830)1.1e-50291.1
CPU, 8 threads, 2 to 3 jobs in parallelS512 wfp16, float16 table (shipped)1,8094,327 (1,830)4,327 (1,830)0.00160531.8
CPU, 8 threads, 2 jobs in parallelS512 fp32, float16 table1,8094,327 (1,830)4,327 (1,830)0.000240214.2
CPU, 8 threads, 1 to 2 jobs in parallelS512 wfp16, float32 table1,8094,327 (1,830)4,327 (1,830)0.00160467.8
CPU, idle machine, 20-request probeS512 fp32, float32 table2048481.7e-60203.6
Metal GPU, explicit FP32S512 fp32, float32 table1002222223.5e-6068.4
Metal GPU, explicit FP32S512 wfp16, float16 table (shipped)100222 (92)222 (92)0.0016068.2
Metal GPU, explicit FP32S256 wfp16, float16 table1002002000.00081027.6
Metal GPU, default precision (fp16 kernels)S512 wfp16, float16 table1002220: every output non-finite55.3
Metal GPU, default precision (fp16 kernels)S256 wfp16, float16 table1002000: every output non-finite
Metal GPU, default precision (fp16 kernels)S512 fp32, float32 table1002220: every output non-finite55.9
Galaxy S26 GPU, explicit FP32; thermal status 0 to 3, battery 31.0 to 44.9 °C; compile 3.69 sS512 wfp16, float16 table (shipped)5001,302 (754)1,302 (754)0.000830697.5 [575.5, 1,201.7]
Galaxy S26 GPU, explicit FP32; thermal status 2, battery 42.4 to 43.0 °C; compile 4.97 sS256 wfp16, float16 table (shipped)4191,074 (712)1,074 (712)0.000740407.6 [375.6, 484.7]
Galaxy S26 GPU, explicit FP32; thermal status 2 on the debug-gate runs before and after; the sample in android/sample/ with its Kotlin tokenizerShipped S256 (143 requests) and S512 (17 requests), float16 table160420420; ids and spans equal to the author's Collator on 160 of 160 requests0.000690301 (S256 299, S512 620)
Galaxy S26 GPU, default precision (fp16); thermal status 2, battery 43.0 to 42.3 °C; compile 5.12 sS512 wfp16 with a -1e4 mask (experiment), float16 table100274 (189)216 (142): finite but wrong0.6662481,011.9 [933.5, 1,064.3]
Galaxy S26 NPU (Hexagon, JIT, BURST mode)S512 wfp16, float16 table (shipped)nonecompile failed: the HTP prepare ran out of host memory 94 s after the compile started
Galaxy S26 NPU (Hexagon, JIT, BURST mode); thermal status 1 to 2, battery 40.8 to 41.9 °C; JIT compile 186.9 sS256 wfp16, float16 table (shipped)100276 (192)95 (54): finite but wrong0.9986265129.7 [124.2, 140.3]
Galaxy S26 NPU (Hexagon, JIT, BURST mode); thermal status 2, battery 40.9 to 42.4 °C; JIT compile 200.8 sS256 wfp16 with a -1e4 mask (experiment), float16 table100270 (187)102 (65); 6 more questions non-finite0.999259128.3
Galaxy S26 NPU (Hexagon, JIT, BURST mode); JIT compile 180.7 sS256 wfp16 with a -1e4 mask and SafeLayerNorm k=2 (experiment), float16 table100268112; 8 more questions non-finite0.998258111.8
Galaxy S26 NPU (Hexagon, JIT, BURST mode); JIT compile 171.8 sS256 wfp16 with a -1e4 mask and SafeLayerNorm k=3 (experiment), float16 table100276 (192)274 (190), all finite0.02547110.0 [106.9, 117.4]
Galaxy S26 NPU (Hexagon, JIT, BURST mode); JIT compile 182.5 sS256 wfp16 with a -1e4 mask and SafeLayerNorm k=4 (experiment), float16 table100276 (192)273 (189), all finite0.027911108.7 [102.6, 112.8]
Galaxy S26 CPU; Pixel phones: GPU, NPU, CPUanynot measured

S512 and S256 name the window size. The S256 CPU rows cover the 1,281 requests that fit 256 tokens. On the shipped S512 CPU row, every question kind matches: choice 1,802 of 1,802, score 716 of 716 and noul 1,809 of 1,809. The largest change in a score answer, the expected level, is 0.0009. The CPU probe compiled in 0.5 s, and the shipped S512 graph compiled in 3.25 s on Metal with explicit FP32.

The S26 S512 row with explicit FP32 uses 500 of the fixture requests: all 300 from ood-test.jsonl, the README example, the 8 invented tickets and 191 random requests from test.jsonl. The S26 S256 row covers the 419 of them that fit 256 tokens. On both rows the largest change in a score answer is 0.00079. The sample row uses 160 fixture requests: the README example, the 8 invented tickets, 100 from ood-test.jsonl and 51 from test.jsonl.

With explicit FP32 on the S26 GPU, each shipped graph ran all 1,785 operators in one LITERT_CL partition. The initial call took 620.9 ms at S512 and 428.1 ms at S256. The Kotlin float16 table lookup is not in these times: it took a median of 37.5 ms per S512 request and 21.7 ms per S256 request in the debug gate app. On the NPU, the shipped S256 graph and the k=3 graph each ran whole as one DispatchDelegate node. Where a row has non-finite outputs, its question count covers only the requests whose outputs were finite.

At the default Metal precision, every output is non-finite on all three graphs tried. On the S26 NPU, the shipped S256 graph gives finite but wrong answers. The next section explains both.

Accuracy on the author's public test file comes from the author's implementation at temperature 1.05, scored with the author's metrics code. On the first 1,500 states of data-fam/test.jsonl (3,474 questions), accuracy is 0.8584. By domain it is 0.9531 for banking77, 0.7560 for sst5 and 0.8752 for boolq, and score questions reach 0.5726. On the first 300 states of data-fam/ood-test.jsonl (818 questions), accuracy is 0.7017, with 0.4312 on score questions. The metric is argmax accuracy, and on desktop CPU the shipped graph picks the same option on every one of these questions, so it scores the same.

The author reports 0.854 in-domain and 0.690 out of domain. Those are the author's numbers, on a test set of 1,500 states and 3,508 questions (4,012 out-of-domain questions) that is not in the author's repository. The repository's data-fam/test.jsonl is a later generation, with 489 sst5 states against 506 and 513 boolq states against 496. Our numbers are on that public file, not on the author's test set.

Not measured:

  • Pixel phones: GPU, NPU and CPU.
  • The Galaxy S26 CPU.
  • Sustained runs.

No NPU graph ships in this repository. The NPU rows above are not a verified path.

fp16, the GPU default precision and the NPU

Use the GPU with explicit FP32 computation. In fp16 computation, two parts of the graph fail.

One is the attention mask constant. The graph uses torch.finfo(float32).min, -3.4e38, the stock transformers DeBERTa value. In fp16 this constant becomes -inf. At every real token (1 - mask) is 0, and 0 × -inf is NaN. At the default Metal precision, which runs fp16 kernels, every output is non-finite, for the fp32 graph and the wfp16 graphs alike. A torch fp16 run of the graph agrees: the earliest non-finite module is layers.0.layer.attention.output.dense, and 0 logits are finite. On the S26 NPU, the shipped S256 graph returns finite outputs, but only 95 of 276 answers match the author's implementation.

The other is LayerNorm. Apart from the mask, the fp32 values are small. On the 413-token / 89-option request, the residual stream peaks at 28.5, the attention scores at 32.4 and the FFN intermediate at 142. But LayerNorm sums squares over 1024 channels, and 1024 × 28.5² is 8.3e5, above the fp16 maximum of 65,504. A -1e4 mask alone keeps fp32 results bit-identical on 6 probe requests at S512. On Metal at default precision the outputs then become finite, but only 185 of 222 answers match. The S26 GPU at default precision matches 216 of 274, and the S26 NPU 102 of 270, with 6 more questions non-finite.

The fix has two parts, and it is exact in fp32: the -1e4 mask, and SafeLayerNorm with a shift of k=3, which computes all 49 LayerNorms on x·2⁻³ with eps·2⁻⁶. The graph stays bit-identical in fp32 on 5 probe requests. With both parts, Metal at default precision matches all 200 answers of 100 S256 requests, with a maximum probability difference of 0.0152. The S26 NPU matches 274 of 276 answers, with a maximum probability difference of 0.0254, at a median of 110.0 ms per 256-token request. The S26 GPU with explicit FP32 takes 407.6 ms for the same window. The two answers that differ are near-ties: the reference's top two probabilities differ by 0.0065 and 0.040. A shift of k=2 still overflows, with 8 questions non-finite. A shift of k=4 is no better, at 273 of 276.

This release ships only the two graphs above, unchanged: neither rewrite is applied to them. Run them on the GPU with explicit FP32. The k=3 NPU graph is not included, because two of its answers differ from the author's implementation. The S512 graph cannot be compiled for the NPU on the S26: the HTP prepare ran out of host memory 94 s after the compile started.

Explicit FP32 is the GPU setting verified here: GpuOptions(enforce_f32=True) on Metal and GpuOptions(precision = FP32) on the S26.

Limits

  • The agreement numbers measure how closely the conversion follows the author's implementation, not task accuracy. The conversion reproduces the author's answers, including the wrong ones.
  • English only, per the author. Training covers three public domains: banking77 support messages, SST-5 review sentences and BoolQ passages. The author says anything else is out of distribution and should be measured before use.
  • The window is 512 tokens, and the author's contract cuts the state to 256 tokens. Of the 1,809 fixture requests, 1,281 fit the S256 graph and all fit the S512 graph. The longest has 485 tokens.
  • A request has 128 option slots, shared by all its questions. The author's contract allows up to 255 options in one choice question, and such a request does not fit this graph. The fixtures need at most 89.
  • A request that does not fit the window or the option slots raises an error. Only the state is cut.
  • In the author's words, the model "reads the question only partly". The author reports a gap between in-domain and out-of-domain questions, and "unseen ordered scales are the weakest". On the public test file the gap is 0.8584 against 0.7017, and score questions out of domain reach 0.4312.
  • Confidence comes from a temperature fitted on the author's validation split. The author reports an in-domain ECE of 0.022 and advises re-calibrating on your own data for out-of-domain questions.
  • The model does not generate text. It chooses among the options you give it.
  • One desktop and one phone, the Galaxy S26, were tested. The Kotlin tokenizer is in android/sample/.
  • The phone GPU warms under load. Over the 500-request S512 run, the thermal status rose from 0 to 3, and later requests took longer than early ones. Sustained throughput is lower than the early requests show.

Provenance, conversion and license

  • Source: com-kotobalabs/open-jev-deberta-v3-large at revision 188ee67a5c93122b916e5acd5bdb0cb3623e380a. model.safetensors is 1,736,094,384 bytes, all F32, with 434,012,160 parameters, SHA-256 3f1d5bc3b6d3dc412ea2c499446fcbd212242d79baf16bc0eda155b4ebbac806. head.safetensors is 12,591,412 bytes, SHA-256 f101be67c5808a810abbcede9ebe702f7e4af4aa6f7298bb243842ce9e3a295a.
  • Model: a DebertaV2Model encoder with 24 layers, hidden size 1024, 16 heads, FFN size 4096, relative attention with 256 position buckets, no absolute position embeddings and a vocabulary of 128,100. For each option, the head applies a 3072-to-1024 linear layer, GELU and a 1024-to-1 linear layer to [mean of the question's text tokens; mean of the option's text tokens; their product]. Softmax runs within each question's options at temperature 1.05, which the author fitted on the validation split.
  • Training data, per the author: public gold labels only, from banking77 (CC-BY-4.0), SST-5 (the author names no license) and BoolQ (CC-BY-SA-3.0). The base model is microsoft/deberta-v3-large, which the author lists as MIT.
  • The author describes the model as an "independent reproduction of the shape" of a hosted decision API. It is "not affiliated" with that API's maker and uses none of its data or code.
  • Conversion: litert-torch 0.9.3 with torch 2.12.1 and transformers 4.57.6, at fixed shapes. The encoder layers are a DeBERTa-v2 layer rewritten for fixed shapes, in conversion/shaped_layer.py (SHA-256 prefix 1adada39), the same file used for the GLiNER2.5 conversion. It pre-expands the log-bucket relative positions as constants and uses rank-4 batch matmuls, a float mask and native GELU.
  • The author's span pooling becomes the two routing inputs, q_routing and o_routing. The question and option means are then matrix products inside the graph.
  • Each graph has 1,639 operators, or 1,785 in the wfp16 form. None is GATHER, GATHER_ND, CAST, SELECT_V2, BROADCAST_TO or MAXIMUM, and no BATCH_MATMUL has a constant left operand. No tensor is int64, and none is above rank 4.
  • Float16 weights: ai-edge-quantizer 0.8.0, FLOAT_CASTING, weight-only.
  • fp16 experiments: conversion/graph.py carries the two optional fp16 rewrites used in the experiments, the finite mask constant and SafeLayerNorm. conversion/export.py turns them on with --mask-neg and --ln-shift. The shipped graphs use neither.
  • Rewrite check: in PyTorch fp32, the rewritten graph is bit-identical to the stock DebertaV2Model path. The maximum difference is 0.0 on 5 probe requests at S256 and 6 at S512, including the 485-token / 89-option request. The rewritten graph's logits are within 9.5e-7 of the author's implementation.
  • Word table: the checkpoint's float32 word embeddings rounded to the nearest float16, with no inf or NaN. SHA-256 d1b86bce8ffa0aae67e7d99db5824e67b5241337d0b5132bf25f22bf9dd72b45. tokenizer.json has SHA-256 cd119378b0160677b7a1e561ba29ada83918c1d420b326516a585382e83d9d39.
  • Verification: the LiteRT CompiledModel Python API on desktop CPU and Metal, and the Kotlin CompiledModel API on the Galaxy S26. Every check is the same argmax plus the absolute probability difference against the author's implementation.

License: the author's model artifacts are licensed under Apache 2.0. These converted files, the conversion scripts, the Python host and the Kotlin block are released under the same license. DeBERTa-v3-large is MIT. The source model is published by Mithril (formerly Kotoba Cloud), operated by Kotoba Labs Inc. The license text is in LICENSE, and attribution is in NOTICE.

calibrated
deberta-v2
decision-model
litert
text-classification
tflite
typed-decisions

litert-community/Open-Decision-DeBERTa-v3-Large-LiteRT

Model

Open Decision DeBERTa-v3-Large for LiteRT

0

4 commits

3 linked in READMEs

updated Oct 3, 2026

See the code

README

Open Decision DeBERTa-v3-Large for LiteRT

com-kotobalabs/open-jev-deberta-v3-large is a decision model built on DeBERTa-v3-large. It reads one state and several typed questions (choice, score, noul) and returns a calibrated probability distribution for each question from one forward pass. This repository holds it converted to LiteRT with litert-torch 0.9.3. The verified path on desktop is the LiteRT CompiledModel Python API on CPU and on the Apple Metal GPU with explicit FP32 computation (ai-edge-litert 2.1.6, Apple M4 Max, 2026-10-01). On a phone, it is the Kotlin CompiledModel API on the Galaxy S26 GPU with explicit FP32 computation (LiteRT 2.2.0, 2026-10-01 and 2026-10-02). On desktop CPU, the shipped graph (512-token window, float16 weights, float16 word table) and the author's implementation (CPU FP32) picked the same option on all 4,327 questions of 1,809 requests, and no probability differed by more than 0.0016 at temperature 1.05.

On the Galaxy S26 GPU, the same graph picked the same option on all 1,302 questions of 500 requests, with a maximum probability difference of 0.00083 and a median of 697.5 ms per request. The 256-token graph also picked the same option on all 1,074 questions of 419 requests, at a median of 407.6 ms. The Kotlin host in android/sample/ ran the fixture gate on the S26: its ids and spans equal the author's Collator on all 160 requests, and all 420 answers match the author's implementation.

An invented support ticket on the left and the three answers the S512 wfp16 graph returned for it on the right

The ticket is invented. The bars are the probabilities that deberta_v3_large_decision_s512_wfp16.tflite with the float16 table returned for it on desktop CPU.

Other formats: onnx-community/open-jev-deberta-v3-large-ONNX, an ONNX port for Transformers.js with fp32, fp16, q4 and q4f16 weights.

Files

FileBytesRole
deberta_v3_large_decision_s512_wfp16.tflite813,443,728Graph for a 512-token window, with float16 weights. The examples below use it. Tested on desktop CPU (1,809 requests), on the Metal GPU with explicit FP32 (100 requests) and on the Galaxy S26 GPU with explicit FP32 (500 requests).
deberta_v3_large_decision_s256_wfp16.tflite712,780,432Graph for a 256-token window, with float16 weights. Tested on desktop CPU (the 1,281 requests that fit 256 tokens), on the Metal GPU with explicit FP32 (100 requests) and on the Galaxy S26 GPU with explicit FP32 (419 requests).
word_embeddings_fp16.bin262,348,800Word table [128100, 1024], little-endian float16, for the host lookup.
tokenizer.json8,657,170The source repository's tokenizer.json, unchanged.
decision_litert.pyPython host: encodes the request, looks up the table rows, builds the routing inputs, runs the graph and reads out the answers.
conversion/Conversion, fixture, check and figure scripts, with requirements-lock.txt.
android/sample/Android sample app (Compose) with the pure-Kotlin host that ran on the Galaxy S26: tokenizer, sequence builder, float16 table lookup, LiteRT call and read-out, plus JVM parity tests and a debug fixture gate.
android/CardSnippet.ktThe Kotlin block below, in the package it was compiled in. The full sample is in android/sample/.
REPRODUCE.mdHow to reproduce the conversion and the checks.
LICENSE, NOTICEApache License 2.0 text and attribution.
assets/hero.pngThe figure above.
SHA256SUMSSHA-256 checksums of the files.

In the wfp16 graphs only the weights are float16. Each of the 146 FULLY_CONNECTED weights feeds a DEQUANTIZE operator, and activations stay float32. The 48 relative-position constants of the batch matmuls also stay float32.

The measurements below also use three float32 files that are not in this repository: the S256 and S512 fp32 graphs (1,323,008,720 and 1,423,672,016 bytes) and the float32 table (524,697,600 bytes, a bit-exact copy of the checkpoint's word embeddings).

Every graph holds the 24-layer encoder, the span means and the scoring head, and returns one logit per option slot. The host does the rest. It tokenizes, looks up each token's table row, builds the two routing inputs, applies softmax at temperature 1.05 and reads out the answers.

Minimal usage

Python: desktop CPU

Needs numpy, tokenizers and ai-edge-litert (tested with 2.1.6). Torch, transformers and the author's typed_decisions package are not needed. Run it from the directory that holds the downloaded files and decision_litert.py.

from decision_litert import DecisionModel

model = DecisionModel(
    "deberta_v3_large_decision_s512_wfp16.tflite",
    "word_embeddings_fp16.bin",
    "tokenizer.json",
)
state = (
    "Order 8841 was supposed to arrive on Monday and the tracking page still "
    "says 'label created'. I need it before Friday for a gift."
)
questions = [
    {
        "type": "choice",
        "instructions": "Which team should handle this message?",
        "options": ["billing", "technical support", "shipping", "account access", "sales"],
    },
    {
        "type": "score",
        "instructions": "How frustrated is the customer?",
        "options": ["calm", "mildly annoyed", "frustrated", "angry"],
    },
    {"type": "noul", "instructions": "The customer asks for a refund."},
]
team, frustration, refund = model.decide(state, questions)
print(team["choice"], round(team["confidence"], 3))
print(round(frustration["score"], 3), round(frustration["confidence"], 3))
print(round(refund["noul"], 3))
model.close()

On a Mac CPU it prints the following. The code rounds each value to 3 decimals.

shipping 0.785
1.861 0.642
0.124

shipping wins with probability 0.785. The score answer is the expected zero-based level, so 1.861 falls between "mildly annoyed" and "frustrated". Its confidence, 0.642, is the probability of "frustrated". The refund question gives p(yes) = 0.124. On the same request (80 tokens), the author's decide() on CPU FP32 gives the same three answers, and no probability differs by more than 6.6e-5.

decide() returns one dict per question, and the choice and score answers also carry probabilities, one per option. To run on the desktop GPU, pass accelerator="gpu". It sets explicit FP32 (GpuOptions(enforce_f32=True)), the Metal setting measured below.

Kotlin: Android GPU with explicit FP32

import android.util.Half
import com.google.ai.edge.litert.Accelerator
import com.google.ai.edge.litert.CompiledModel
import java.io.File
import java.io.RandomAccessFile
import java.nio.ByteOrder
import java.nio.channels.FileChannel
import kotlin.math.exp

/** One request = token ids plus, per option, the token range of its text and of its question's text. */
class Span(val start: Int, val end: Int)

class DecisionGpu(dir: File, private val seq: Int = 512, private val slots: Int = 128) : AutoCloseable {
  private val options = CompiledModel.Options(Accelerator.GPU).apply {
    gpuOptions = CompiledModel.GpuOptions(precision = CompiledModel.GpuOptions.Precision.FP32)
  }
  private val model =
    CompiledModel.create(File(dir, "deberta_v3_large_decision_s${seq}_wfp16.tflite").absolutePath, options, null)
  private val inputs = listOf("inputs_embeds", "attention_mask", "q_routing", "o_routing")
    .associateWith { model.createInputBuffer(it, "serving_default") }
  private val outputs = mapOf("logits" to model.createOutputBuffer("logits", "serving_default"))
  private val table = RandomAccessFile(File(dir, "word_embeddings_fp16.bin"), "r").use {
    it.channel.map(FileChannel.MapMode.READ_ONLY, 0, it.length()).order(ByteOrder.LITTLE_ENDIAN).asShortBuffer()
  }
  private val embeds = FloatArray(seq * 1024)

  /** ids from a tokenizer that matches decision_litert.py; one (question span, option span) pair per option, in order. */
  fun logits(ids: IntArray, pairs: List<Pair<Span, Span>>): FloatArray {
    require(ids.size <= seq && pairs.size <= slots)
    for (p in 0 until seq) {
      val row = (if (p < ids.size) ids[p] else 0) * 1024
      for (c in 0 until 1024) embeds[p * 1024 + c] = Half.toFloat(table.get(row + c))
    }
    val qRouting = FloatArray(slots * seq)
    val oRouting = FloatArray(slots * seq)
    pairs.forEachIndexed { j, (q, o) ->
      for (t in q.start until q.end) qRouting[j * seq + t] = 1f / (q.end - q.start)
      for (t in o.start until o.end) oRouting[j * seq + t] = 1f / (o.end - o.start)
    }
    inputs.getValue("inputs_embeds").writeFloat(embeds)
    inputs.getValue("attention_mask").writeFloat(FloatArray(seq) { if (it < ids.size) 1f else 0f })
    inputs.getValue("q_routing").writeFloat(qRouting)
    inputs.getValue("o_routing").writeFloat(oRouting)
    model.run(inputs, outputs, "serving_default")
    return outputs.getValue("logits").readFloat().copyOf(pairs.size)
  }

  /** Probabilities of one question's options at the shipped temperature 1.05. */
  fun probabilities(logits: FloatArray, from: Int, count: Int): List<Double> {
    val z = (from until from + count).map { logits[it] / 1.05 }
    val top = z.max()
    val e = z.map { exp(it - top) }
    val sum = e.sum()
    return e.map { it / sum }
  }

  override fun close() {
    (inputs.values + outputs.values).forEach { it.close() }
    model.close()
  }
}

Create DecisionGpu once and call logits from one worker thread. logits takes the ids of one request and one (question span, option span) pair per option, in request order, and returns one logit per option. The ids and spans must come from a tokenizer that matches decision_litert.py. The block compiles with LiteRT 2.2.0 (AGP 8.9.1, Kotlin 2.2.21). It has not itself run on a device; the sample host below has.

The full Kotlin host that ran on the Galaxy S26 is in android/sample/, with its own README for download, build and install. Its DecisionModel.kt makes the same LiteRT calls as the block above. Its DecisionTokenizer.kt turns text into ids with tokenizer.json, and its DecisionInputs.kt builds the sequence and the spans.

Host contract

decision_litert.py implements this contract. It follows the Collator in the author's typed_decisions/encoder.py and the read-out in typed_decisions/schema.py.

  1. Tokenize each text with tokenizer.json and no added special tokens, then assemble the ids in this order: [CLS] [STATE] state [Q] instructions [OPT] option [OPT] option ... [Q] ... [SEP]. The state keeps its first 256 tokens. The marker ids are [STATE] 128001, [Q] 128002 and [OPT] 128003. [CLS] is 1, [SEP] is 2 and [PAD] is 0. A noul question always has the options ["no", "yes"]. Record the token span of each question's text and of each option's text as half-open ranges, markers excluded.
  2. Check the limits. A choice question takes 2 to 255 options, and a score question 2 to 10 ordered levels. The ids must fit the graph's window (256 or 512 tokens), and all options of the request must fit the 128 option slots. A request that does not fit raises an error instead of being truncated. Only the state is cut, at 256 tokens.
  3. Pad the ids on the right with id 0 to the window S, and fill the four float32 inputs. inputs_embeds [1,S,1024] holds the table row of every position, padding included. attention_mask [1,S] is 1 for real tokens and 0 for padding. In q_routing [1,128,S], row j is 1/len over the text tokens of option j's question. In o_routing [1,128,S], row j is 1/len over option j's text tokens. Rows past the request's last option stay zero.
  4. Run the serving_default signature. It returns logits [1,1,1,128], one logit per option slot in request order. Ignore the slots past the last option. Split the logits by question and apply softmax at temperature 1.05 within each question's options.
  5. Read the answer. choice is the option with the highest probability, and confidence is that probability. score is Σ i·pᵢ, the expected zero-based level, which may fall between levels. It also carries confidence. noul is p(yes).

On all 1,809 fixture requests, decision_litert.py produces the same ids and text spans as the author's Collator. Its decide() also gives the same answers as the author's decide() on 209 requests: the README example, the 8 invented tickets and the first 200 requests from data-fam. With the S512 wfp16 graph and the float16 table on desktop CPU, probabilities differ by at most 0.0016 and score values by at most 0.00066. With the S256 wfp16 graph, probabilities differ by at most 0.00081.

The tokenizers library on tokenizer.json gives the same ids as the author's AutoTokenizer on all 2,368 fixture texts. The SentencePiece spm.model gives the same ids on 2,367 of them, so build the ids from tokenizer.json.

Measured quality and performance

The reference is the author's implementation: the checkpoint's own typed_decisions package on CPU FP32 (torch 2.12.1, transformers 4.57.6, eager attention, 8 threads), one request per forward pass. Probabilities are compared at temperature 1.05. "Same argmax" means the same winning option for a question.

The fixtures are 1,809 requests with 4,327 questions. They are the first 1,500 states of the author's public data-fam/test.jsonl, the first 300 states of data-fam/ood-test.jsonl, the author's README example and 8 invented tickets with 4 questions each. Both test files come from the author's kotoba-lang/typed-decisions repository at revision e65215db7e554a31f557936f0bd1d8a0a78b6af5. A boundary question is one where the reference's top probability is below 0.9. There are 1,830 of them among the 4,327 questions, and 1,602 among the 2,796 questions that fit 256 tokens.

The desktop rows ran on 2026-10-01 on one desktop: Apple M4 Max (16 cores, 128 GB), macOS 27.0, ai-edge-litert 2.1.6 through the Python CompiledModel API. The Galaxy S26 rows ran on 2026-10-01 and 2026-10-02 on one Galaxy S26 (SM-S942Q, Android 16) with LiteRT 2.2.0 through the Kotlin CompiledModel API. They used a debug gate app, except the row marked as the sample in android/sample/. Times are informational. Each CPU row states the load on the machine. Each phone row states the thermal status and battery temperature where they were recorded, and the compile time. Phone times are warm medians of writing the inputs, running and reading back, with [min, max] where recorded.

WhereGraph, word tableRequestsQuestions (boundary)Same argmaxMax probability differenceQuestions over 0.01Median ms per request
CPU, 8 threads, 3 jobs in parallelS256 fp32, float32 table1,2812,796 (1,602)2,796 (1,602)1.1e-50152.8
CPU, 8 threads, 3 jobs in parallelS256 wfp16, float16 table1,2812,796 (1,602)2,796 (1,602)0.000810305.8
CPU, 8 threads, 2 to 3 jobs in parallelS256 fp32, float16 table1,2812,796 (1,602)2,796 (1,602)0.000240135.9
CPU, 8 threads, 2 to 3 jobs in parallelS256 wfp16, float32 table1,2812,796 (1,602)2,796 (1,602)0.000800350.2
CPU, 8 threads, 3 jobs in parallelS512 fp32, float32 table1,8094,327 (1,830)4,327 (1,830)1.1e-50291.1
CPU, 8 threads, 2 to 3 jobs in parallelS512 wfp16, float16 table (shipped)1,8094,327 (1,830)4,327 (1,830)0.00160531.8
CPU, 8 threads, 2 jobs in parallelS512 fp32, float16 table1,8094,327 (1,830)4,327 (1,830)0.000240214.2
CPU, 8 threads, 1 to 2 jobs in parallelS512 wfp16, float32 table1,8094,327 (1,830)4,327 (1,830)0.00160467.8
CPU, idle machine, 20-request probeS512 fp32, float32 table2048481.7e-60203.6
Metal GPU, explicit FP32S512 fp32, float32 table1002222223.5e-6068.4
Metal GPU, explicit FP32S512 wfp16, float16 table (shipped)100222 (92)222 (92)0.0016068.2
Metal GPU, explicit FP32S256 wfp16, float16 table1002002000.00081027.6
Metal GPU, default precision (fp16 kernels)S512 wfp16, float16 table1002220: every output non-finite55.3
Metal GPU, default precision (fp16 kernels)S256 wfp16, float16 table1002000: every output non-finite
Metal GPU, default precision (fp16 kernels)S512 fp32, float32 table1002220: every output non-finite55.9
Galaxy S26 GPU, explicit FP32; thermal status 0 to 3, battery 31.0 to 44.9 °C; compile 3.69 sS512 wfp16, float16 table (shipped)5001,302 (754)1,302 (754)0.000830697.5 [575.5, 1,201.7]
Galaxy S26 GPU, explicit FP32; thermal status 2, battery 42.4 to 43.0 °C; compile 4.97 sS256 wfp16, float16 table (shipped)4191,074 (712)1,074 (712)0.000740407.6 [375.6, 484.7]
Galaxy S26 GPU, explicit FP32; thermal status 2 on the debug-gate runs before and after; the sample in android/sample/ with its Kotlin tokenizerShipped S256 (143 requests) and S512 (17 requests), float16 table160420420; ids and spans equal to the author's Collator on 160 of 160 requests0.000690301 (S256 299, S512 620)
Galaxy S26 GPU, default precision (fp16); thermal status 2, battery 43.0 to 42.3 °C; compile 5.12 sS512 wfp16 with a -1e4 mask (experiment), float16 table100274 (189)216 (142): finite but wrong0.6662481,011.9 [933.5, 1,064.3]
Galaxy S26 NPU (Hexagon, JIT, BURST mode)S512 wfp16, float16 table (shipped)nonecompile failed: the HTP prepare ran out of host memory 94 s after the compile started
Galaxy S26 NPU (Hexagon, JIT, BURST mode); thermal status 1 to 2, battery 40.8 to 41.9 °C; JIT compile 186.9 sS256 wfp16, float16 table (shipped)100276 (192)95 (54): finite but wrong0.9986265129.7 [124.2, 140.3]
Galaxy S26 NPU (Hexagon, JIT, BURST mode); thermal status 2, battery 40.9 to 42.4 °C; JIT compile 200.8 sS256 wfp16 with a -1e4 mask (experiment), float16 table100270 (187)102 (65); 6 more questions non-finite0.999259128.3
Galaxy S26 NPU (Hexagon, JIT, BURST mode); JIT compile 180.7 sS256 wfp16 with a -1e4 mask and SafeLayerNorm k=2 (experiment), float16 table100268112; 8 more questions non-finite0.998258111.8
Galaxy S26 NPU (Hexagon, JIT, BURST mode); JIT compile 171.8 sS256 wfp16 with a -1e4 mask and SafeLayerNorm k=3 (experiment), float16 table100276 (192)274 (190), all finite0.02547110.0 [106.9, 117.4]
Galaxy S26 NPU (Hexagon, JIT, BURST mode); JIT compile 182.5 sS256 wfp16 with a -1e4 mask and SafeLayerNorm k=4 (experiment), float16 table100276 (192)273 (189), all finite0.027911108.7 [102.6, 112.8]
Galaxy S26 CPU; Pixel phones: GPU, NPU, CPUanynot measured

S512 and S256 name the window size. The S256 CPU rows cover the 1,281 requests that fit 256 tokens. On the shipped S512 CPU row, every question kind matches: choice 1,802 of 1,802, score 716 of 716 and noul 1,809 of 1,809. The largest change in a score answer, the expected level, is 0.0009. The CPU probe compiled in 0.5 s, and the shipped S512 graph compiled in 3.25 s on Metal with explicit FP32.

The S26 S512 row with explicit FP32 uses 500 of the fixture requests: all 300 from ood-test.jsonl, the README example, the 8 invented tickets and 191 random requests from test.jsonl. The S26 S256 row covers the 419 of them that fit 256 tokens. On both rows the largest change in a score answer is 0.00079. The sample row uses 160 fixture requests: the README example, the 8 invented tickets, 100 from ood-test.jsonl and 51 from test.jsonl.

With explicit FP32 on the S26 GPU, each shipped graph ran all 1,785 operators in one LITERT_CL partition. The initial call took 620.9 ms at S512 and 428.1 ms at S256. The Kotlin float16 table lookup is not in these times: it took a median of 37.5 ms per S512 request and 21.7 ms per S256 request in the debug gate app. On the NPU, the shipped S256 graph and the k=3 graph each ran whole as one DispatchDelegate node. Where a row has non-finite outputs, its question count covers only the requests whose outputs were finite.

At the default Metal precision, every output is non-finite on all three graphs tried. On the S26 NPU, the shipped S256 graph gives finite but wrong answers. The next section explains both.

Accuracy on the author's public test file comes from the author's implementation at temperature 1.05, scored with the author's metrics code. On the first 1,500 states of data-fam/test.jsonl (3,474 questions), accuracy is 0.8584. By domain it is 0.9531 for banking77, 0.7560 for sst5 and 0.8752 for boolq, and score questions reach 0.5726. On the first 300 states of data-fam/ood-test.jsonl (818 questions), accuracy is 0.7017, with 0.4312 on score questions. The metric is argmax accuracy, and on desktop CPU the shipped graph picks the same option on every one of these questions, so it scores the same.

The author reports 0.854 in-domain and 0.690 out of domain. Those are the author's numbers, on a test set of 1,500 states and 3,508 questions (4,012 out-of-domain questions) that is not in the author's repository. The repository's data-fam/test.jsonl is a later generation, with 489 sst5 states against 506 and 513 boolq states against 496. Our numbers are on that public file, not on the author's test set.

Not measured:

  • Pixel phones: GPU, NPU and CPU.
  • The Galaxy S26 CPU.
  • Sustained runs.

No NPU graph ships in this repository. The NPU rows above are not a verified path.

fp16, the GPU default precision and the NPU

Use the GPU with explicit FP32 computation. In fp16 computation, two parts of the graph fail.

One is the attention mask constant. The graph uses torch.finfo(float32).min, -3.4e38, the stock transformers DeBERTa value. In fp16 this constant becomes -inf. At every real token (1 - mask) is 0, and 0 × -inf is NaN. At the default Metal precision, which runs fp16 kernels, every output is non-finite, for the fp32 graph and the wfp16 graphs alike. A torch fp16 run of the graph agrees: the earliest non-finite module is layers.0.layer.attention.output.dense, and 0 logits are finite. On the S26 NPU, the shipped S256 graph returns finite outputs, but only 95 of 276 answers match the author's implementation.

The other is LayerNorm. Apart from the mask, the fp32 values are small. On the 413-token / 89-option request, the residual stream peaks at 28.5, the attention scores at 32.4 and the FFN intermediate at 142. But LayerNorm sums squares over 1024 channels, and 1024 × 28.5² is 8.3e5, above the fp16 maximum of 65,504. A -1e4 mask alone keeps fp32 results bit-identical on 6 probe requests at S512. On Metal at default precision the outputs then become finite, but only 185 of 222 answers match. The S26 GPU at default precision matches 216 of 274, and the S26 NPU 102 of 270, with 6 more questions non-finite.

The fix has two parts, and it is exact in fp32: the -1e4 mask, and SafeLayerNorm with a shift of k=3, which computes all 49 LayerNorms on x·2⁻³ with eps·2⁻⁶. The graph stays bit-identical in fp32 on 5 probe requests. With both parts, Metal at default precision matches all 200 answers of 100 S256 requests, with a maximum probability difference of 0.0152. The S26 NPU matches 274 of 276 answers, with a maximum probability difference of 0.0254, at a median of 110.0 ms per 256-token request. The S26 GPU with explicit FP32 takes 407.6 ms for the same window. The two answers that differ are near-ties: the reference's top two probabilities differ by 0.0065 and 0.040. A shift of k=2 still overflows, with 8 questions non-finite. A shift of k=4 is no better, at 273 of 276.

This release ships only the two graphs above, unchanged: neither rewrite is applied to them. Run them on the GPU with explicit FP32. The k=3 NPU graph is not included, because two of its answers differ from the author's implementation. The S512 graph cannot be compiled for the NPU on the S26: the HTP prepare ran out of host memory 94 s after the compile started.

Explicit FP32 is the GPU setting verified here: GpuOptions(enforce_f32=True) on Metal and GpuOptions(precision = FP32) on the S26.

Limits

  • The agreement numbers measure how closely the conversion follows the author's implementation, not task accuracy. The conversion reproduces the author's answers, including the wrong ones.
  • English only, per the author. Training covers three public domains: banking77 support messages, SST-5 review sentences and BoolQ passages. The author says anything else is out of distribution and should be measured before use.
  • The window is 512 tokens, and the author's contract cuts the state to 256 tokens. Of the 1,809 fixture requests, 1,281 fit the S256 graph and all fit the S512 graph. The longest has 485 tokens.
  • A request has 128 option slots, shared by all its questions. The author's contract allows up to 255 options in one choice question, and such a request does not fit this graph. The fixtures need at most 89.
  • A request that does not fit the window or the option slots raises an error. Only the state is cut.
  • In the author's words, the model "reads the question only partly". The author reports a gap between in-domain and out-of-domain questions, and "unseen ordered scales are the weakest". On the public test file the gap is 0.8584 against 0.7017, and score questions out of domain reach 0.4312.
  • Confidence comes from a temperature fitted on the author's validation split. The author reports an in-domain ECE of 0.022 and advises re-calibrating on your own data for out-of-domain questions.
  • The model does not generate text. It chooses among the options you give it.
  • One desktop and one phone, the Galaxy S26, were tested. The Kotlin tokenizer is in android/sample/.
  • The phone GPU warms under load. Over the 500-request S512 run, the thermal status rose from 0 to 3, and later requests took longer than early ones. Sustained throughput is lower than the early requests show.

Provenance, conversion and license

  • Source: com-kotobalabs/open-jev-deberta-v3-large at revision 188ee67a5c93122b916e5acd5bdb0cb3623e380a. model.safetensors is 1,736,094,384 bytes, all F32, with 434,012,160 parameters, SHA-256 3f1d5bc3b6d3dc412ea2c499446fcbd212242d79baf16bc0eda155b4ebbac806. head.safetensors is 12,591,412 bytes, SHA-256 f101be67c5808a810abbcede9ebe702f7e4af4aa6f7298bb243842ce9e3a295a.
  • Model: a DebertaV2Model encoder with 24 layers, hidden size 1024, 16 heads, FFN size 4096, relative attention with 256 position buckets, no absolute position embeddings and a vocabulary of 128,100. For each option, the head applies a 3072-to-1024 linear layer, GELU and a 1024-to-1 linear layer to [mean of the question's text tokens; mean of the option's text tokens; their product]. Softmax runs within each question's options at temperature 1.05, which the author fitted on the validation split.
  • Training data, per the author: public gold labels only, from banking77 (CC-BY-4.0), SST-5 (the author names no license) and BoolQ (CC-BY-SA-3.0). The base model is microsoft/deberta-v3-large, which the author lists as MIT.
  • The author describes the model as an "independent reproduction of the shape" of a hosted decision API. It is "not affiliated" with that API's maker and uses none of its data or code.
  • Conversion: litert-torch 0.9.3 with torch 2.12.1 and transformers 4.57.6, at fixed shapes. The encoder layers are a DeBERTa-v2 layer rewritten for fixed shapes, in conversion/shaped_layer.py (SHA-256 prefix 1adada39), the same file used for the GLiNER2.5 conversion. It pre-expands the log-bucket relative positions as constants and uses rank-4 batch matmuls, a float mask and native GELU.
  • The author's span pooling becomes the two routing inputs, q_routing and o_routing. The question and option means are then matrix products inside the graph.
  • Each graph has 1,639 operators, or 1,785 in the wfp16 form. None is GATHER, GATHER_ND, CAST, SELECT_V2, BROADCAST_TO or MAXIMUM, and no BATCH_MATMUL has a constant left operand. No tensor is int64, and none is above rank 4.
  • Float16 weights: ai-edge-quantizer 0.8.0, FLOAT_CASTING, weight-only.
  • fp16 experiments: conversion/graph.py carries the two optional fp16 rewrites used in the experiments, the finite mask constant and SafeLayerNorm. conversion/export.py turns them on with --mask-neg and --ln-shift. The shipped graphs use neither.
  • Rewrite check: in PyTorch fp32, the rewritten graph is bit-identical to the stock DebertaV2Model path. The maximum difference is 0.0 on 5 probe requests at S256 and 6 at S512, including the 485-token / 89-option request. The rewritten graph's logits are within 9.5e-7 of the author's implementation.
  • Word table: the checkpoint's float32 word embeddings rounded to the nearest float16, with no inf or NaN. SHA-256 d1b86bce8ffa0aae67e7d99db5824e67b5241337d0b5132bf25f22bf9dd72b45. tokenizer.json has SHA-256 cd119378b0160677b7a1e561ba29ada83918c1d420b326516a585382e83d9d39.
  • Verification: the LiteRT CompiledModel Python API on desktop CPU and Metal, and the Kotlin CompiledModel API on the Galaxy S26. Every check is the same argmax plus the absolute probability difference against the author's implementation.

License: the author's model artifacts are licensed under Apache 2.0. These converted files, the conversion scripts, the Python host and the Kotlin block are released under the same license. DeBERTa-v3-large is MIT. The source model is published by Mithril (formerly Kotoba Cloud), operated by Kotoba Labs Inc. The license text is in LICENSE, and attribution is in NOTICE.

calibrated
deberta-v2
decision-model
litert
text-classification
tflite
typed-decisions