litert-community/Julia-1-LiteRT

Model

Julia-1 for LiteRT — Android GPU

1

4 commits

3 linked in READMEs

updated Sep 30, 2026

See the code

README

Julia-1 for LiteRT — Android GPU

The SupersonicLabs/Julia-1 decision model, converted to LiteRT with litert-torch, answers one question in a median of 80.7 ms (first session) to 112.1 ms (second session) of graph time on a Samsung Galaxy S26 GPU with explicit FP32 computation, at a 512-token window (three runs on 2026-09-30; each run's thermal state is below). Julia-1 reads a state (text or JSON), a typed question and 2 to 20 options, and returns one probability per option from one forward pass. A choice question returns the winning option, a score question the expected index on an ordered rubric, and a noul (yes/no) question the probability of true. On the phone, all 706 validation rows got the same answer as the author's runtime: probabilities within 0.0077 with the shipped float16 token table, and within 0.00005 with a float32 table. In one run on the phone's GPU, the Android sample answered the three questions of its support-ticket preset (151 prompt tokens in all) in 229.7 ms end to end. On desktop CPU, all 2,000 questions of the typed-decisions test got the same answer with the S1024 graph, and all 2,065 rows that fit 512 tokens with the S512 graph.

An invented support ticket on the left and the three answers the S512 graph returned for it on the right

The ticket is invented. The bars are the probabilities that julia1_s512_fp32.tflite with the float16 table returned for it on desktop CPU, the same run as the Python example below.

Other formats of Julia-1: SupersonicLabs/Julia-1-ONNX (the author's ONNX export and WebGPU adapter), zainmerchan/Julia-1-MLX and andrelucas/Julia-1-GGUF.

Files

FileBytesRole
julia1_s512_fp32.tflite185,095,348Graph for a 512-token window. Tested on the S26 GPU with explicit FP32 and on desktop CPU.
julia1_s1024_fp32.tflite188,503,220Graph for a 1,024-token window, for requests longer than 512 tokens. Tested on the S26 GPU with explicit FP32 (100 rows) and on desktop CPU (all 2,000 typed-decisions questions).
julia1_token_table_fp16.bin196,608,000Token table [256000, 384], little-endian float16, for the host lookup.
tokenizer.json34,363,188The source repository's tokenizer/tokenizer.json, unchanged.
julia_litert.pyPython host: encodes the request, looks up the table rows, runs the graph and decodes the answers.
conversion/Conversion and check scripts, and GateActivity.kt, the activity that ran on the phone.
android/Android sample: Kotlin host and Jetpack Compose UI.
REPRODUCE.mdHow to reproduce the conversion.
LICENSE, NOTICEApache License 2.0 text and attribution.
assets/hero.pngThe figure above.
SHA256SUMSSHA-256 checksums of the files.

The S512 graph, the token table and tokenizer.json together are 416,066,536 bytes (185,095,348 + 196,608,000 + 34,363,188).

Both graphs hold the mmBERT-small encoder, the question-type embedding, the two head layers and the option scorer, which gives one logit per position. They keep the checkpoint's float32 weights. The host does the rest: it tokenizes, builds the sequence with one <mask> marker per option, looks up each token's table row, reads the logits at the markers and applies softmax.

The float16 table is half the size of the checkpoint's float32 table (393,216,000 bytes). In every check it changes no answer and moves probabilities by up to 0.0077: desktop CPU (2,065 rows at S512, 2,000 at S1024) and the S26 GPU (706 rows at S512, 100 at S1024).

Minimal usage

Python — desktop CPU

Needs numpy, tokenizers and ai-edge-litert (tested with 2.1.6). Torch, transformers and the author's julia package are not needed. Run it from the directory that holds the downloaded files and julia_litert.py.

from julia_litert import JuliaLiteRT

engine = JuliaLiteRT(".", graph="julia1_s512_fp32.tflite", table="julia1_token_table_fp16.bin")
state = (
    "Order #4417 from Larkspur Home Goods was charged twice. The customer, "
    "Mira Okafor, wants the extra charge refunded before Friday."
)
questions = {
    "team": {
        "type": "choice",
        "instructions": "Which team should handle this request?",
        "criteria": {
            "billing": "Billing and payment disputes",
            "shipping": "Shipping and delivery",
            "access": "Account access and login",
        },
    },
    "urgency": {
        "type": "score",
        "instructions": "How urgent is this request?",
        "criteria": ["Not urgent", "Somewhat urgent", "Urgent", "Critical"],
    },
    "refund": {
        "type": "noul",
        "instructions": "Is the customer asking for money back?",
        "criteria": {"false": "No refund is requested", "true": "A refund is requested"},
    },
}
answers = engine.predict(state, questions)["answers"]
print(answers["team"]["choice"], answers["team"]["probabilities"])
print(answers["urgency"]["score"], answers["urgency"]["probabilities"])
print(answers["refund"]["noul"])

On a Mac CPU it prints the following (values rounded here to 4 decimals):

billing {'billing': 0.7925, 'shipping': 0.2073, 'access': 0.0002}
0.8169 {'0': 0.6492, '1': 0.0191, '2': 0.1975, '3': 0.1342}
0.6915

On the same request, the author's runtime (CPU FP32) gives the same three answers, and no probability differs by more than 0.001.

Kotlin — Android GPU with explicit FP32

import android.util.Half
import com.google.ai.edge.litert.Accelerator
import com.google.ai.edge.litert.CompiledModel
import java.io.File
import java.io.RandomAccessFile
import java.nio.ByteOrder
import java.nio.channels.FileChannel
import kotlin.math.exp

class JuliaGpu(dir: File) : AutoCloseable {
  private val options = CompiledModel.Options(Accelerator.GPU).apply {
    gpuOptions = CompiledModel.GpuOptions(precision = CompiledModel.GpuOptions.Precision.FP32)
  }
  private val model = CompiledModel.create(File(dir, "julia1_s512_fp32.tflite").absolutePath, options, null)
  private val inputs = listOf("inputs_embeds", "attention_mask", "qtype_onehot")
    .associateWith { model.createInputBuffer(it, "serving_default") }
  private val outputs = mapOf("token_logits" to model.createOutputBuffer("token_logits", "serving_default"))
  private val table = RandomAccessFile(File(dir, "julia1_token_table_fp16.bin"), "r").use {
    it.channel.map(FileChannel.MapMode.READ_ONLY, 0, it.length()).order(ByteOrder.LITTLE_ENDIAN).asShortBuffer()
  }
  private val embeds = FloatArray(512 * 384)

  /** ids and markers come from a tokenizer that matches julia_litert.py; qtype 0 choice, 1 score, 2 noul. */
  fun probabilities(ids: IntArray, markers: IntArray, qtype: Int): List<Double> {
    for (p in 0 until 512) {
      val row = (if (p < ids.size) ids[p] else 0) * 384
      for (c in 0 until 384) embeds[p * 384 + c] = Half.toFloat(table.get(row + c))
    }
    inputs.getValue("inputs_embeds").writeFloat(embeds)
    inputs.getValue("attention_mask").writeFloat(FloatArray(512) { if (it < ids.size) 1f else 0f })
    inputs.getValue("qtype_onehot").writeFloat(FloatArray(3) { if (it == qtype) 1f else 0f })
    model.run(inputs, outputs, "serving_default")
    val tokens = outputs.getValue("token_logits").readFloat()
    val z = markers.map { tokens[it].toDouble() }
    val top = z.max()
    val e = z.map { exp(it - top) }
    val sum = e.sum()
    return e.map { it / sum }
  }

  override fun close() {
    (inputs.values + outputs.values).forEach { it.close() }
    model.close()
  }
}

Create JuliaGpu once and call probabilities from one worker thread. It takes the token ids, the marker positions and the question type of one request, and returns one probability per option. The ids and markers must come from a tokenizer that matches julia_litert.py. The Android sample in android/ contains a Kotlin tokenizer and a strict sequence builder that follow julia_litert.py. Its Jetpack Compose UI has three presets (support ticket, agent trace, product review) with editable text. The LiteRT calls are the ones that produced the S26 numbers, with one difference: the measuring app passed an Environment where this block passes null. This float16 lookup is the one the second-session phone rows used.

Host contract

julia_litert.py implements this contract. It follows sequence() in the author's julia/data.py and the named questions in julia/typed.py.

  1. Build the ids: <bos> "{type} question: {question}" <eos> <mask> " option0" <mask> " option1" … <eos> state <eos>. Each option may have at most 48 tokens. A dict or list state is json.dumps(state, ensure_ascii=False). Remember the position of every <mask>. A request has 2 to 20 options. A noul request has two, ordered [false, true].
  2. Encode strictly: a request that does not fit raises an error instead of being truncated. Of the 2,000 questions in the LocalLLaMA/typed-decisions test, 35 need more than 512 tokens (up to 607). They fit the S1024 graph.
  3. Pad the ids on the right with id 0 to the window S. inputs_embeds [1,S,384] holds the table row of every position, padding included. attention_mask [1,S] is 1 for real tokens and 0 for padding. qtype_onehot [1,3] selects choice, score or noul.
  4. The graph returns token_logits [1,S]. Take the logits at the marker positions and apply softmax at temperature 1. The author ships no calibration: the checkpoint's temperature buffer is 1, and its inference-policy.json sets calibration to null.
  5. Read the answer. choice is the option with the highest probability. score is Σ i·pᵢ, the expected zero-based index, not rounded. noul is P(true).

In the named-question form, the criteria are the option texts: the mapping values for choice, the rubric strings for score, and [criteria.false, criteria.true] for noul, or the literal false and true when no criteria are given. On all 2,065 rows that fit 512 tokens, and on all 2,100 rows at 1,024 tokens, julia_litert.py produces the same ids, marker positions and question types as the author's encoding.

Measured quality and performance

The reference is the author's runtime: engine.logits from the julia package in the source repository, on CPU FP32 (torch 2.14.0, transformers 5.0.0, Apple M4 Max). Probabilities are compared at temperature 1.

The rows come from two sets. The LocalLLaMA/typed-decisions test at revision c76749ec58bd8c3d2ea706b31c333a9059c38f90 has 2,000 questions, and 1,965 of them fit 512 tokens. parity-cases.json in SupersonicLabs/Julia-1-ONNX (revision 82a2fadf) adds 100 requests, for 2,065 rows. Of these, 306 are boundary rows: the reference's top probability is below 0.9.

On the 100 parity requests, the reference reproduces the logits the author stored there: the same argmax on 100 of 100, and a maximum logit difference of 1.0e-4.

Device, acceleratorGraph, token tableRowsSame argmaxMax probability differenceRows over 0.01Median ms per question
S26 GPU, explicit FP32, first sessionS512 fp32, float32 table7067060.00005080.7
S26 NPU (Hexagon, JIT)S512 wfp16, float32 table7066840.42529627.7
S26 NPU (Hexagon, JIT)S512 fp32, float32 table7066850.41828629.3
Desktop CPUS512 fp32, float32 table2,0652,0650.00007091
Desktop CPUS512 fp32, float16 table (shipped)2,0652,0650.0077091
S26 GPU, explicit FP32, second session, run 1S512 fp32, float16 table (shipped)7067060.00770111.8
S26 GPU, explicit FP32, second session, run 2S512 fp32, float16 table (shipped)7067060.00770112.1
S26 GPU, explicit FP32, second sessionS1024 fp32, float16 table (shipped)1001000.00770298.5
S26 GPU, FP16S512 fp32not measured
S26 CPUS512 fp32not measured
S26 NPUS1024 fp32not measured

The S512 phone rows are all 306 boundary rows plus 400 others. The phone is one Samsung Galaxy S26 (SM-S942Q, Android 16) running LiteRT 2.2.0 through the Kotlin CompiledModel API, in a debug gate app. The GPU time is the median of 701 warm rows, from writing the inputs to the end of the output readback.

In the first session, warm rows took between 56.3 ms and 104.1 ms on the GPU. The battery went from 34.2 °C to 38.6 °C during the GPU run, and the thermal status was 0 before it. GPU compilation took 0.92 s, and all 1,704 operators ran in one LITERT_CL partition.

The second session ran the same S512 graph twice with the shipped float16 table. Run 1 started at thermal status 0 and 35.7 °C, run 2 at thermal status 1 and 38.2 °C, and run 2 ended at thermal status 2 and 41.3 °C. Warm rows took 56.0 ms to 134.9 ms and 56.7 ms to 134.2 ms. Compilation took 0.91 s and 0.88 s, and all 1,704 operators ran in one LITERT_CL partition. The minimum row time was 56.0 ms to 56.7 ms in all three S512 runs. The medians differ between the sessions, and the table file is not in the timed interval.

In the same session the S1024 graph ran 100 rows: all 35 typed-decisions questions longer than 512 tokens plus 65 others, 18 of them boundary rows. All 100 got the same argmax, with a maximum difference of 0.0077, at a median of 298.5 ms per question. Warm rows took 155.6 ms to 305.2 ms, compilation took 1.47 s, all 1,704 operators ran in one partition, and the battery went from 41.3 °C to 41.9 °C at thermal status 2.

The host token lookup is not in these times. In the gate app, a Kotlin loop over the float32 table took a median of 19.0 ms per question, and the loop over the float16 table (Half.toFloat) 23.6 ms in run 1.

wfp16 is a variant of the S512 graph that stores the FULLY_CONNECTED weights in float16 (93,383,808 bytes). It is not shipped. The NPU rows used JIT compilation with QualcommOptions BURST, which took 34 s for the wfp16 graph and 41 s for the fp32 graph.

The whole wfp16 graph ran on the NPU as one DispatchDelegate node (1,803 of 1,803 operators). The wfp16 NPU run started at thermal status 0 (battery 38.6 °C) and ended at 2 (42.2 °C). The fp32 NPU run started and ended at thermal status 2 (40.9 °C to 42.8 °C).

Desktop CPU means ai-edge-litert 2.1.6 on a Mac with 4 threads. Its 91 ms was measured on a loaded machine and is informational.

The Mac GPU (Metal, ai-edge-litert 2.1.6 on the same M4 Max) ran 400 rows, 306 of them boundary rows. With explicit FP32 (enforce_f32), all 400 got the same argmax, with a maximum difference of 0.00007, at 11.4 ms per question (informational). At the default precision (fp16), 375 of 400 got the same argmax, with a maximum difference of 0.57 and 285 rows over 0.01.

The author's reproduction script, scripts/reproduce_typed.py, run unchanged on CPU FP32 (1,024-token window, 4 threads), gives 426/600 choice, 542/800 score and 483/600 noul. The author's published CPU FP32 run (metrics/typed-cpu-20260926.json in the source repository) has the same numbers.

With the S1024 graph, the Python host gives the same answer as the author's runtime on all 2,000 typed-decisions questions, with either table: 426/600 choice, 542/800 score and 483/600 noul, the author's CPU FP32 numbers. The maximum probability difference is 0.00007 with the float32 table and 0.0077 with the shipped float16 table. With the S512 graph it gives the same answer on all 1,965 questions that fit 512 tokens, again with either table. For task accuracy, see the author's figures on the source model card.

fp16 and the NPU

Use the GPU with explicit FP32 computation. The fp16 paths change answers. The Hexagon NPU computes float graphs in fp16. The S26 NPU gave the same argmax on 684 of 706 rows at 27.7 ms per question, with probabilities off by up to 0.425. The Mac GPU at its default fp16 precision gave the same argmax on 375 of 400.

The answers are sensitive to fp16 rounding. Storing only the weights in fp16, with fp32 arithmetic, moves probabilities by up to 0.031 on desktop CPU.

Arithmetic in fp16 moves probabilities further, and the encoder is the sensitive part. In a Mac emulation on the 306 boundary rows, an fp16 encoder with an fp32 head moved them by up to 0.087, and an fp16 head alone by up to 0.0014.

The residual stream holds values of at most 36 through layer 11. The MLP of layer 11 then writes values up to 3,147, and the stream reaches 5,315 in layers 19–21. Without a rewrite, (x − mean)² inside LayerNorm reaches 2.8e7, which fp16 cannot represent.

Both graphs carry rewrites that are exact in fp32 and keep fp16 finite. Finite is not the same answer, as the NPU rows show.

No fp16 file and no NPU-specific file is shipped, and INT8 files were not built.

Limits

  • The numbers measure agreement with the author's runtime, not task accuracy. The conversion reproduces the author's answers, including the wrong ones.
  • One device was tested: one Galaxy S26 (SM-S942Q). Other phones have not been validated here.
  • The S1024 graph ran 100 rows on the S26 GPU, all at thermal status 2, and all 2,000 typed-decisions questions on desktop CPU. It was not tried on the NPU.
  • Phone timings are single runs at the stated thermal state. Sustained runs and the S26 CPU were not measured.
  • The validation used English text only: the typed-decisions test and the ONNX parity requests. The source model is multilingual. Other languages were not checked here.
  • Any app must produce the same ids and marker positions as julia_litert.py. The Android sample in android/ does so in Kotlin. On the S26, its tokenizer and builder matched julia_litert.py on all 2,065 rows that fit 512 tokens, and rejected the 35 that do not. Its GPU run got the same answer as the author's runtime on all 706 phone rows, with probabilities within 0.0077.
  • A request takes 2 to 20 options, the author's native limit.
  • The act head is not included. The checkpoint holds an act_head that the author's public API does not use: predict, logits and the named questions run with return_actions=False.

Provenance, conversion and license

  • Source: SupersonicLabs/Julia-1 at revision a85b127321d580d65176c89ced8273f305745d85. model.safetensors is 577,189,056 bytes, all float32, SHA-256 df853bf7fe424420011f3d0c47a05d7341aa9eefa7fb9f203ea4aada4ad95b72. The model has 144.3M parameters (the author's figure): the jhu-clsp/mmBERT-small encoder (22 layers, hidden size 384, vocabulary 256,000), a question-type embedding, two pre-norm transformer layers and a scorer whose logits are read at the <mask> markers.
  • Conversion: litert-torch 0.9.3 with torch 2.12.1, and transformers 5.0.0 to load the checkpoint. Shapes are fixed. The math is unchanged: host-side token lookup, rotary tables baked per window type, the sliding-window band as a constant, the question type as a one-hot matmul, the two head layers written out, and exact GELU. The S512 graph has 1,704 operators, no GATHER, no int64 tensor and no tensor above rank 4.
  • Rewrites that are exact in fp32: LayerNorm computed on x·2⁻ᵏ with epsilon·2⁻²ᵏ (k = 8 in layers 12–21 and the final norm, k = 2 in the head and scorer norms), attention masks of −1e4 instead of −1e9, and q scaled by 1/8 before QKᵀ. In PyTorch, the token logits with and without them are bit-identical on 5 probe rows.
  • Token table: the checkpoint's float32 embedding table, rounded to float16.
  • Verification: the LiteRT CompiledModel Python API on desktop CPU and Metal, and the Kotlin CompiledModel API on the Galaxy S26 GPU and NPU. Every check is the same argmax plus the absolute probability difference against the author's runtime. Correlation was never used as a check.

License: the source README says "The model artifacts are licensed under Apache 2.0." These converted files are released under the same license. mmBERT-small declares MIT. Julia-1 is by Supersonic Labs and builds on mmBERT-small by JHU CLSP. The conversion scripts are in conversion/, and REPRODUCE.md gives the steps. The license text is in LICENSE, and attribution is in NOTICE.

android
litert
modernbert
routing
text-classification
tflite
zero-shot-classification

litert-community/Julia-1-LiteRT

Model

Julia-1 for LiteRT — Android GPU

1

4 commits

3 linked in READMEs

updated Sep 30, 2026

See the code

README

Julia-1 for LiteRT — Android GPU

The SupersonicLabs/Julia-1 decision model, converted to LiteRT with litert-torch, answers one question in a median of 80.7 ms (first session) to 112.1 ms (second session) of graph time on a Samsung Galaxy S26 GPU with explicit FP32 computation, at a 512-token window (three runs on 2026-09-30; each run's thermal state is below). Julia-1 reads a state (text or JSON), a typed question and 2 to 20 options, and returns one probability per option from one forward pass. A choice question returns the winning option, a score question the expected index on an ordered rubric, and a noul (yes/no) question the probability of true. On the phone, all 706 validation rows got the same answer as the author's runtime: probabilities within 0.0077 with the shipped float16 token table, and within 0.00005 with a float32 table. In one run on the phone's GPU, the Android sample answered the three questions of its support-ticket preset (151 prompt tokens in all) in 229.7 ms end to end. On desktop CPU, all 2,000 questions of the typed-decisions test got the same answer with the S1024 graph, and all 2,065 rows that fit 512 tokens with the S512 graph.

An invented support ticket on the left and the three answers the S512 graph returned for it on the right

The ticket is invented. The bars are the probabilities that julia1_s512_fp32.tflite with the float16 table returned for it on desktop CPU, the same run as the Python example below.

Other formats of Julia-1: SupersonicLabs/Julia-1-ONNX (the author's ONNX export and WebGPU adapter), zainmerchan/Julia-1-MLX and andrelucas/Julia-1-GGUF.

Files

FileBytesRole
julia1_s512_fp32.tflite185,095,348Graph for a 512-token window. Tested on the S26 GPU with explicit FP32 and on desktop CPU.
julia1_s1024_fp32.tflite188,503,220Graph for a 1,024-token window, for requests longer than 512 tokens. Tested on the S26 GPU with explicit FP32 (100 rows) and on desktop CPU (all 2,000 typed-decisions questions).
julia1_token_table_fp16.bin196,608,000Token table [256000, 384], little-endian float16, for the host lookup.
tokenizer.json34,363,188The source repository's tokenizer/tokenizer.json, unchanged.
julia_litert.pyPython host: encodes the request, looks up the table rows, runs the graph and decodes the answers.
conversion/Conversion and check scripts, and GateActivity.kt, the activity that ran on the phone.
android/Android sample: Kotlin host and Jetpack Compose UI.
REPRODUCE.mdHow to reproduce the conversion.
LICENSE, NOTICEApache License 2.0 text and attribution.
assets/hero.pngThe figure above.
SHA256SUMSSHA-256 checksums of the files.

The S512 graph, the token table and tokenizer.json together are 416,066,536 bytes (185,095,348 + 196,608,000 + 34,363,188).

Both graphs hold the mmBERT-small encoder, the question-type embedding, the two head layers and the option scorer, which gives one logit per position. They keep the checkpoint's float32 weights. The host does the rest: it tokenizes, builds the sequence with one <mask> marker per option, looks up each token's table row, reads the logits at the markers and applies softmax.

The float16 table is half the size of the checkpoint's float32 table (393,216,000 bytes). In every check it changes no answer and moves probabilities by up to 0.0077: desktop CPU (2,065 rows at S512, 2,000 at S1024) and the S26 GPU (706 rows at S512, 100 at S1024).

Minimal usage

Python — desktop CPU

Needs numpy, tokenizers and ai-edge-litert (tested with 2.1.6). Torch, transformers and the author's julia package are not needed. Run it from the directory that holds the downloaded files and julia_litert.py.

from julia_litert import JuliaLiteRT

engine = JuliaLiteRT(".", graph="julia1_s512_fp32.tflite", table="julia1_token_table_fp16.bin")
state = (
    "Order #4417 from Larkspur Home Goods was charged twice. The customer, "
    "Mira Okafor, wants the extra charge refunded before Friday."
)
questions = {
    "team": {
        "type": "choice",
        "instructions": "Which team should handle this request?",
        "criteria": {
            "billing": "Billing and payment disputes",
            "shipping": "Shipping and delivery",
            "access": "Account access and login",
        },
    },
    "urgency": {
        "type": "score",
        "instructions": "How urgent is this request?",
        "criteria": ["Not urgent", "Somewhat urgent", "Urgent", "Critical"],
    },
    "refund": {
        "type": "noul",
        "instructions": "Is the customer asking for money back?",
        "criteria": {"false": "No refund is requested", "true": "A refund is requested"},
    },
}
answers = engine.predict(state, questions)["answers"]
print(answers["team"]["choice"], answers["team"]["probabilities"])
print(answers["urgency"]["score"], answers["urgency"]["probabilities"])
print(answers["refund"]["noul"])

On a Mac CPU it prints the following (values rounded here to 4 decimals):

billing {'billing': 0.7925, 'shipping': 0.2073, 'access': 0.0002}
0.8169 {'0': 0.6492, '1': 0.0191, '2': 0.1975, '3': 0.1342}
0.6915

On the same request, the author's runtime (CPU FP32) gives the same three answers, and no probability differs by more than 0.001.

Kotlin — Android GPU with explicit FP32

import android.util.Half
import com.google.ai.edge.litert.Accelerator
import com.google.ai.edge.litert.CompiledModel
import java.io.File
import java.io.RandomAccessFile
import java.nio.ByteOrder
import java.nio.channels.FileChannel
import kotlin.math.exp

class JuliaGpu(dir: File) : AutoCloseable {
  private val options = CompiledModel.Options(Accelerator.GPU).apply {
    gpuOptions = CompiledModel.GpuOptions(precision = CompiledModel.GpuOptions.Precision.FP32)
  }
  private val model = CompiledModel.create(File(dir, "julia1_s512_fp32.tflite").absolutePath, options, null)
  private val inputs = listOf("inputs_embeds", "attention_mask", "qtype_onehot")
    .associateWith { model.createInputBuffer(it, "serving_default") }
  private val outputs = mapOf("token_logits" to model.createOutputBuffer("token_logits", "serving_default"))
  private val table = RandomAccessFile(File(dir, "julia1_token_table_fp16.bin"), "r").use {
    it.channel.map(FileChannel.MapMode.READ_ONLY, 0, it.length()).order(ByteOrder.LITTLE_ENDIAN).asShortBuffer()
  }
  private val embeds = FloatArray(512 * 384)

  /** ids and markers come from a tokenizer that matches julia_litert.py; qtype 0 choice, 1 score, 2 noul. */
  fun probabilities(ids: IntArray, markers: IntArray, qtype: Int): List<Double> {
    for (p in 0 until 512) {
      val row = (if (p < ids.size) ids[p] else 0) * 384
      for (c in 0 until 384) embeds[p * 384 + c] = Half.toFloat(table.get(row + c))
    }
    inputs.getValue("inputs_embeds").writeFloat(embeds)
    inputs.getValue("attention_mask").writeFloat(FloatArray(512) { if (it < ids.size) 1f else 0f })
    inputs.getValue("qtype_onehot").writeFloat(FloatArray(3) { if (it == qtype) 1f else 0f })
    model.run(inputs, outputs, "serving_default")
    val tokens = outputs.getValue("token_logits").readFloat()
    val z = markers.map { tokens[it].toDouble() }
    val top = z.max()
    val e = z.map { exp(it - top) }
    val sum = e.sum()
    return e.map { it / sum }
  }

  override fun close() {
    (inputs.values + outputs.values).forEach { it.close() }
    model.close()
  }
}

Create JuliaGpu once and call probabilities from one worker thread. It takes the token ids, the marker positions and the question type of one request, and returns one probability per option. The ids and markers must come from a tokenizer that matches julia_litert.py. The Android sample in android/ contains a Kotlin tokenizer and a strict sequence builder that follow julia_litert.py. Its Jetpack Compose UI has three presets (support ticket, agent trace, product review) with editable text. The LiteRT calls are the ones that produced the S26 numbers, with one difference: the measuring app passed an Environment where this block passes null. This float16 lookup is the one the second-session phone rows used.

Host contract

julia_litert.py implements this contract. It follows sequence() in the author's julia/data.py and the named questions in julia/typed.py.

  1. Build the ids: <bos> "{type} question: {question}" <eos> <mask> " option0" <mask> " option1" … <eos> state <eos>. Each option may have at most 48 tokens. A dict or list state is json.dumps(state, ensure_ascii=False). Remember the position of every <mask>. A request has 2 to 20 options. A noul request has two, ordered [false, true].
  2. Encode strictly: a request that does not fit raises an error instead of being truncated. Of the 2,000 questions in the LocalLLaMA/typed-decisions test, 35 need more than 512 tokens (up to 607). They fit the S1024 graph.
  3. Pad the ids on the right with id 0 to the window S. inputs_embeds [1,S,384] holds the table row of every position, padding included. attention_mask [1,S] is 1 for real tokens and 0 for padding. qtype_onehot [1,3] selects choice, score or noul.
  4. The graph returns token_logits [1,S]. Take the logits at the marker positions and apply softmax at temperature 1. The author ships no calibration: the checkpoint's temperature buffer is 1, and its inference-policy.json sets calibration to null.
  5. Read the answer. choice is the option with the highest probability. score is Σ i·pᵢ, the expected zero-based index, not rounded. noul is P(true).

In the named-question form, the criteria are the option texts: the mapping values for choice, the rubric strings for score, and [criteria.false, criteria.true] for noul, or the literal false and true when no criteria are given. On all 2,065 rows that fit 512 tokens, and on all 2,100 rows at 1,024 tokens, julia_litert.py produces the same ids, marker positions and question types as the author's encoding.

Measured quality and performance

The reference is the author's runtime: engine.logits from the julia package in the source repository, on CPU FP32 (torch 2.14.0, transformers 5.0.0, Apple M4 Max). Probabilities are compared at temperature 1.

The rows come from two sets. The LocalLLaMA/typed-decisions test at revision c76749ec58bd8c3d2ea706b31c333a9059c38f90 has 2,000 questions, and 1,965 of them fit 512 tokens. parity-cases.json in SupersonicLabs/Julia-1-ONNX (revision 82a2fadf) adds 100 requests, for 2,065 rows. Of these, 306 are boundary rows: the reference's top probability is below 0.9.

On the 100 parity requests, the reference reproduces the logits the author stored there: the same argmax on 100 of 100, and a maximum logit difference of 1.0e-4.

Device, acceleratorGraph, token tableRowsSame argmaxMax probability differenceRows over 0.01Median ms per question
S26 GPU, explicit FP32, first sessionS512 fp32, float32 table7067060.00005080.7
S26 NPU (Hexagon, JIT)S512 wfp16, float32 table7066840.42529627.7
S26 NPU (Hexagon, JIT)S512 fp32, float32 table7066850.41828629.3
Desktop CPUS512 fp32, float32 table2,0652,0650.00007091
Desktop CPUS512 fp32, float16 table (shipped)2,0652,0650.0077091
S26 GPU, explicit FP32, second session, run 1S512 fp32, float16 table (shipped)7067060.00770111.8
S26 GPU, explicit FP32, second session, run 2S512 fp32, float16 table (shipped)7067060.00770112.1
S26 GPU, explicit FP32, second sessionS1024 fp32, float16 table (shipped)1001000.00770298.5
S26 GPU, FP16S512 fp32not measured
S26 CPUS512 fp32not measured
S26 NPUS1024 fp32not measured

The S512 phone rows are all 306 boundary rows plus 400 others. The phone is one Samsung Galaxy S26 (SM-S942Q, Android 16) running LiteRT 2.2.0 through the Kotlin CompiledModel API, in a debug gate app. The GPU time is the median of 701 warm rows, from writing the inputs to the end of the output readback.

In the first session, warm rows took between 56.3 ms and 104.1 ms on the GPU. The battery went from 34.2 °C to 38.6 °C during the GPU run, and the thermal status was 0 before it. GPU compilation took 0.92 s, and all 1,704 operators ran in one LITERT_CL partition.

The second session ran the same S512 graph twice with the shipped float16 table. Run 1 started at thermal status 0 and 35.7 °C, run 2 at thermal status 1 and 38.2 °C, and run 2 ended at thermal status 2 and 41.3 °C. Warm rows took 56.0 ms to 134.9 ms and 56.7 ms to 134.2 ms. Compilation took 0.91 s and 0.88 s, and all 1,704 operators ran in one LITERT_CL partition. The minimum row time was 56.0 ms to 56.7 ms in all three S512 runs. The medians differ between the sessions, and the table file is not in the timed interval.

In the same session the S1024 graph ran 100 rows: all 35 typed-decisions questions longer than 512 tokens plus 65 others, 18 of them boundary rows. All 100 got the same argmax, with a maximum difference of 0.0077, at a median of 298.5 ms per question. Warm rows took 155.6 ms to 305.2 ms, compilation took 1.47 s, all 1,704 operators ran in one partition, and the battery went from 41.3 °C to 41.9 °C at thermal status 2.

The host token lookup is not in these times. In the gate app, a Kotlin loop over the float32 table took a median of 19.0 ms per question, and the loop over the float16 table (Half.toFloat) 23.6 ms in run 1.

wfp16 is a variant of the S512 graph that stores the FULLY_CONNECTED weights in float16 (93,383,808 bytes). It is not shipped. The NPU rows used JIT compilation with QualcommOptions BURST, which took 34 s for the wfp16 graph and 41 s for the fp32 graph.

The whole wfp16 graph ran on the NPU as one DispatchDelegate node (1,803 of 1,803 operators). The wfp16 NPU run started at thermal status 0 (battery 38.6 °C) and ended at 2 (42.2 °C). The fp32 NPU run started and ended at thermal status 2 (40.9 °C to 42.8 °C).

Desktop CPU means ai-edge-litert 2.1.6 on a Mac with 4 threads. Its 91 ms was measured on a loaded machine and is informational.

The Mac GPU (Metal, ai-edge-litert 2.1.6 on the same M4 Max) ran 400 rows, 306 of them boundary rows. With explicit FP32 (enforce_f32), all 400 got the same argmax, with a maximum difference of 0.00007, at 11.4 ms per question (informational). At the default precision (fp16), 375 of 400 got the same argmax, with a maximum difference of 0.57 and 285 rows over 0.01.

The author's reproduction script, scripts/reproduce_typed.py, run unchanged on CPU FP32 (1,024-token window, 4 threads), gives 426/600 choice, 542/800 score and 483/600 noul. The author's published CPU FP32 run (metrics/typed-cpu-20260926.json in the source repository) has the same numbers.

With the S1024 graph, the Python host gives the same answer as the author's runtime on all 2,000 typed-decisions questions, with either table: 426/600 choice, 542/800 score and 483/600 noul, the author's CPU FP32 numbers. The maximum probability difference is 0.00007 with the float32 table and 0.0077 with the shipped float16 table. With the S512 graph it gives the same answer on all 1,965 questions that fit 512 tokens, again with either table. For task accuracy, see the author's figures on the source model card.

fp16 and the NPU

Use the GPU with explicit FP32 computation. The fp16 paths change answers. The Hexagon NPU computes float graphs in fp16. The S26 NPU gave the same argmax on 684 of 706 rows at 27.7 ms per question, with probabilities off by up to 0.425. The Mac GPU at its default fp16 precision gave the same argmax on 375 of 400.

The answers are sensitive to fp16 rounding. Storing only the weights in fp16, with fp32 arithmetic, moves probabilities by up to 0.031 on desktop CPU.

Arithmetic in fp16 moves probabilities further, and the encoder is the sensitive part. In a Mac emulation on the 306 boundary rows, an fp16 encoder with an fp32 head moved them by up to 0.087, and an fp16 head alone by up to 0.0014.

The residual stream holds values of at most 36 through layer 11. The MLP of layer 11 then writes values up to 3,147, and the stream reaches 5,315 in layers 19–21. Without a rewrite, (x − mean)² inside LayerNorm reaches 2.8e7, which fp16 cannot represent.

Both graphs carry rewrites that are exact in fp32 and keep fp16 finite. Finite is not the same answer, as the NPU rows show.

No fp16 file and no NPU-specific file is shipped, and INT8 files were not built.

Limits

  • The numbers measure agreement with the author's runtime, not task accuracy. The conversion reproduces the author's answers, including the wrong ones.
  • One device was tested: one Galaxy S26 (SM-S942Q). Other phones have not been validated here.
  • The S1024 graph ran 100 rows on the S26 GPU, all at thermal status 2, and all 2,000 typed-decisions questions on desktop CPU. It was not tried on the NPU.
  • Phone timings are single runs at the stated thermal state. Sustained runs and the S26 CPU were not measured.
  • The validation used English text only: the typed-decisions test and the ONNX parity requests. The source model is multilingual. Other languages were not checked here.
  • Any app must produce the same ids and marker positions as julia_litert.py. The Android sample in android/ does so in Kotlin. On the S26, its tokenizer and builder matched julia_litert.py on all 2,065 rows that fit 512 tokens, and rejected the 35 that do not. Its GPU run got the same answer as the author's runtime on all 706 phone rows, with probabilities within 0.0077.
  • A request takes 2 to 20 options, the author's native limit.
  • The act head is not included. The checkpoint holds an act_head that the author's public API does not use: predict, logits and the named questions run with return_actions=False.

Provenance, conversion and license

  • Source: SupersonicLabs/Julia-1 at revision a85b127321d580d65176c89ced8273f305745d85. model.safetensors is 577,189,056 bytes, all float32, SHA-256 df853bf7fe424420011f3d0c47a05d7341aa9eefa7fb9f203ea4aada4ad95b72. The model has 144.3M parameters (the author's figure): the jhu-clsp/mmBERT-small encoder (22 layers, hidden size 384, vocabulary 256,000), a question-type embedding, two pre-norm transformer layers and a scorer whose logits are read at the <mask> markers.
  • Conversion: litert-torch 0.9.3 with torch 2.12.1, and transformers 5.0.0 to load the checkpoint. Shapes are fixed. The math is unchanged: host-side token lookup, rotary tables baked per window type, the sliding-window band as a constant, the question type as a one-hot matmul, the two head layers written out, and exact GELU. The S512 graph has 1,704 operators, no GATHER, no int64 tensor and no tensor above rank 4.
  • Rewrites that are exact in fp32: LayerNorm computed on x·2⁻ᵏ with epsilon·2⁻²ᵏ (k = 8 in layers 12–21 and the final norm, k = 2 in the head and scorer norms), attention masks of −1e4 instead of −1e9, and q scaled by 1/8 before QKᵀ. In PyTorch, the token logits with and without them are bit-identical on 5 probe rows.
  • Token table: the checkpoint's float32 embedding table, rounded to float16.
  • Verification: the LiteRT CompiledModel Python API on desktop CPU and Metal, and the Kotlin CompiledModel API on the Galaxy S26 GPU and NPU. Every check is the same argmax plus the absolute probability difference against the author's runtime. Correlation was never used as a check.

License: the source README says "The model artifacts are licensed under Apache 2.0." These converted files are released under the same license. mmBERT-small declares MIT. Julia-1 is by Supersonic Labs and builds on mmBERT-small by JHU CLSP. The conversion scripts are in conversion/, and REPRODUCE.md gives the steps. The license text is in LICENSE, and attribution is in NOTICE.

android
litert
modernbert
routing
text-classification
tflite
zero-shot-classification