0
stars
5
commits
1
linked in READMEs
Jul 30, 2026
updated
Built with Inkling (Thinking Machines Lab).
MLX (Apple Silicon) conversion of thinkingmachines/Inkling-Small, quantized to 6-bit (affine group quant, group size 64).
Code / loader: github.com/PipeNetwork/inkling-mlx
Inkling Small is a 276B-total / 12B-active sparse-MoE, natively multimodal model (text + image/video + audio → text). This is the full multimodal conversion: all three towers (text backbone, HMLP vision, dMel audio) are ported; the multi-token-prediction head is dropped (inference-irrelevant).
| Variant | Size | Text ppl | Notes |
|---|---|---|---|
| 8bit | ~280 GB | 5.569 | near-lossless |
| 6bit | ~214 GB | 5.569 | high quality |
| 4bit | ~148 GB | 5.452 | balanced default |
| 3bit | ~115 GB | 6.706 | ⚠️ experimental — visibly degraded |
Perplexity is teacher-forcing over one fixed held-out set (prose / code / reasoning / multilingual) — identical inputs across builds, so the columns compare directly. 4-bit shows no measurable loss vs 8-bit.
There is also a REAP-pruned build: REAP25-4bit keeps 4-bit precision with 192 of 256 routed experts, fitting a 128 GB Mac at ~112 GB for no measurable perplexity cost, with vision and speech intact.
No bf16 build is published. The MLX bf16 conversion is bit-identical to the upstream checkpoint (name-mapping and layout only — the dtype cast is a no-op), so it would carry nothing thinkingmachines/Inkling-Small does not already have, and at ~527 GB it does not fit a 512 GB Mac. If you want it as a requant source, scripts/convert_all.sh regenerates it from the upstream weights in about three minutes.
MLX supports FP4 modes and Thinking Machines ships an Inkling-NVFP4 checkpoint — so for the record, we benchmarked round-trip reconstruction error (‖W − Ŵ‖ / ‖W‖ vs bf16) on real Inkling expert weights:
| Scheme | bits/weight | reconstruction error |
|---|---|---|
| affine int4 (group 64) | 4.50 | ~9.1% |
| nvfp4 (group 16) | 4.50 | ~10.2% |
| mxfp4 (group 32) | 4.25 | ~12.3% |
Affine int4 is the most faithful: it is asymmetric (per-group scale and zero-point, 16 uniform levels), which centers on Inkling's near-Gaussian expert weights better than symmetric FP4's fixed non-uniform levels. FP4's real payoff is heavy-tailed activations and native Blackwell FP4 tensor cores — neither helps weight fidelity on Apple Silicon, where MLX would dequantize FP4 anyway. So these builds use affine int4.
inkling_mlx loaderThe inkling_mm_model architecture is not in stock mlx-lm / mlx-vlm, so this
repo bundles a minimal, numerically-validated MLX implementation under inkling_mlx/.
pip install mlx mlx-lm transformers
from inkling_mlx.load import load
from inkling_mlx.generate import greedy_generate
from transformers import AutoTokenizer
model, config = load("/path/to/this/repo")
tok = AutoTokenizer.from_pretrained("/path/to/this/repo", trust_remote_code=True)
ids = tok("The capital of France is")["input_ids"]
print(tok.decode(greedy_generate(model, config, ids, max_new_tokens=64)))
Needs an Apple-Silicon Mac with enough unified memory to hold the weights (≈ the size above).
InklingProcessor — image patchify/normalize, audio log-mel→dMel,
validated ~1e-7 vs the reference) are included. Pass images/audio via the processor.Conversion is streaming (tensor-by-tensor; the ~527 GB bf16 model never fully loads into RAM) and was validated with fp32 numerical parity against transformers PR #47347. License: Apache-2.0 (inherits the base model).
5 commits
0
stars
5
commits
1
linked in READMEs
Jul 30, 2026
updated
Built with Inkling (Thinking Machines Lab).
MLX (Apple Silicon) conversion of thinkingmachines/Inkling-Small, quantized to 6-bit (affine group quant, group size 64).
Code / loader: github.com/PipeNetwork/inkling-mlx
Inkling Small is a 276B-total / 12B-active sparse-MoE, natively multimodal model (text + image/video + audio → text). This is the full multimodal conversion: all three towers (text backbone, HMLP vision, dMel audio) are ported; the multi-token-prediction head is dropped (inference-irrelevant).
| Variant | Size | Text ppl | Notes |
|---|---|---|---|
| 8bit | ~280 GB | 5.569 | near-lossless |
| 6bit | ~214 GB | 5.569 | high quality |
| 4bit | ~148 GB | 5.452 | balanced default |
| 3bit | ~115 GB | 6.706 | ⚠️ experimental — visibly degraded |
Perplexity is teacher-forcing over one fixed held-out set (prose / code / reasoning / multilingual) — identical inputs across builds, so the columns compare directly. 4-bit shows no measurable loss vs 8-bit.
There is also a REAP-pruned build: REAP25-4bit keeps 4-bit precision with 192 of 256 routed experts, fitting a 128 GB Mac at ~112 GB for no measurable perplexity cost, with vision and speech intact.
No bf16 build is published. The MLX bf16 conversion is bit-identical to the upstream checkpoint (name-mapping and layout only — the dtype cast is a no-op), so it would carry nothing thinkingmachines/Inkling-Small does not already have, and at ~527 GB it does not fit a 512 GB Mac. If you want it as a requant source, scripts/convert_all.sh regenerates it from the upstream weights in about three minutes.
MLX supports FP4 modes and Thinking Machines ships an Inkling-NVFP4 checkpoint — so for the record, we benchmarked round-trip reconstruction error (‖W − Ŵ‖ / ‖W‖ vs bf16) on real Inkling expert weights:
| Scheme | bits/weight | reconstruction error |
|---|---|---|
| affine int4 (group 64) | 4.50 | ~9.1% |
| nvfp4 (group 16) | 4.50 | ~10.2% |
| mxfp4 (group 32) | 4.25 | ~12.3% |
Affine int4 is the most faithful: it is asymmetric (per-group scale and zero-point, 16 uniform levels), which centers on Inkling's near-Gaussian expert weights better than symmetric FP4's fixed non-uniform levels. FP4's real payoff is heavy-tailed activations and native Blackwell FP4 tensor cores — neither helps weight fidelity on Apple Silicon, where MLX would dequantize FP4 anyway. So these builds use affine int4.
inkling_mlx loaderThe inkling_mm_model architecture is not in stock mlx-lm / mlx-vlm, so this
repo bundles a minimal, numerically-validated MLX implementation under inkling_mlx/.
pip install mlx mlx-lm transformers
from inkling_mlx.load import load
from inkling_mlx.generate import greedy_generate
from transformers import AutoTokenizer
model, config = load("/path/to/this/repo")
tok = AutoTokenizer.from_pretrained("/path/to/this/repo", trust_remote_code=True)
ids = tok("The capital of France is")["input_ids"]
print(tok.decode(greedy_generate(model, config, ids, max_new_tokens=64)))
Needs an Apple-Silicon Mac with enough unified memory to hold the weights (≈ the size above).
InklingProcessor — image patchify/normalize, audio log-mel→dMel,
validated ~1e-7 vs the reference) are included. Pass images/audio via the processor.Conversion is streaming (tensor-by-tensor; the ~527 GB bf16 model never fully loads into RAM) and was validated with fp32 numerical parity against transformers PR #47347. License: Apache-2.0 (inherits the base model).
5 commits