3
stars
12
commits
1
linked in READMEs
Sep 8, 2026
updated
397B intelligence on a 128 GB Mac. This model fits in 112 GB — the first 397B quantization that runs on M4 Pro/Max 128 GB machines. Uses reasoning mode for 86.5% MMLU accuracy.
LM Studio, Ollama, oMLX do NOT support JANG format. Use MLX Studio or
pip install "jang[mlx]>=2.1.5".
JANG — Jang Adaptive N-bit Grading | The GGUF Equivalent for MLX
JANG is fully open-source. Quantization engine, research, and full commit history: github.com/jjang-ai/jangq. Created by Jinho Jang.
<think>...</think> step-by-step problem solvingPer-subject comparison across all modes. Both JANG and MLX 4-bit tested with and without reasoning.
| Subject | JANG No-Think | JANG Reasoning | MLX 4-bit No-Think | MLX 4-bit Reasoning |
|---|---|---|---|---|
| Abstract Algebra | 8/20 | 10/20 | 10/20 | 17/20 |
| Anatomy | 17/20 | 19/20 | 18/20 | 19/20 |
| Astronomy | 20/20 | 20/20 | 19/20 | 19/20 |
| College CS | 17/20 | 18/20 | 15/20 | 18/20 |
| College Physics | 17/20 | 18/20 | 15/20 | 19/20 |
| HS Biology | 19/20 | 20/20 | 19/20 | 19/20 |
| HS Chemistry | 17/20 | 18/20 | 17/20 | 19/20 |
| HS Mathematics | 8/20 | 10/20 | 12/20 | 19/20 |
| Logical Fallacies | 20/20 | 20/20 | 19/20 | 20/20 |
| World Religions | 19/20 | 20/20 | 19/20 | 19/20 |
| Total | 162/200 (81.0%) | 173/200 (86.5%) | 163/200 (81.5%) | 188/200 (94.0%) |
| JANG_1L | JANG_2L | MLX 4-bit | MLX 2/3-bit | |
|---|---|---|---|---|
| MMLU (no-think) | 81.0% | 79.5% | 81.5% | NaN -- cannot run |
| MMLU (reasoning) | 86.5% | 92.0% | 94.0% | NaN -- cannot run |
| Size | 112 GB | 187 GB | 209 GB | N/A |
| GPU RAM | 110 GB | 184 GB | ~210 GB | N/A |
| Speed | 36.1 tok/s | 36.0 tok/s | ~36 tok/s | N/A |
| Fits 128 GB? | YES | No | No | N/A |
JANG_1L is 97 GB smaller than MLX 4-bit and fits on 128 GB Macs where MLX 4-bit (209 GB) cannot run. MLX 2-bit and 3-bit produce NaN -- cannot run (float16 overflow on 512-expert models). JANG solves this with bfloat16.
| Metric | Value |
|---|---|
| Source | Qwen3.5-397B-A17B |
| Architecture | Hybrid MoE + SSM (GatedDeltaNet + Full Attention) |
| Experts | 512 per layer, top-10 active (17B active params) |
| Profile | JANG_1L (CRITICAL=8, IMPORTANT=8, COMPRESS=2) |
| Average bits | 2.13 bpw |
| Disk size | 112 GB |
| GPU RAM | 110 GB (peak 120 GB) |
| Speed | 36.1 tok/s generation, 96 tok/s prefill |
| Compute | bfloat16 (auto-detected) |
| VLM | 333 vision tensors, 31.6 tok/s |
pip install "jang[mlx]>=2.1.5"pip install "jang[mlx]>=2.1.5"
from jang_tools.loader import load_jang_model
from mlx_lm import generate
model, tokenizer = load_jang_model("JANGQ-AI/Qwen3.5-397B-A17B-JANG_1L")
# With reasoning
messages = [{"role": "user", "content": "Prove that sqrt(2) is irrational."}]
prompt = tokenizer.apply_chat_template(messages, tokenize=False,
add_generation_prompt=True, enable_thinking=True)
result = generate(model, tokenizer, prompt=prompt, max_tokens=2048)
# Without reasoning (faster)
prompt = tokenizer.apply_chat_template(messages, tokenize=False,
add_generation_prompt=True, enable_thinking=False)
result = generate(model, tokenizer, prompt=prompt, max_tokens=100)
from jang_tools.loader import load_jang_vlm_model
from mlx_vlm import generate as vlm_generate
model, processor = load_jang_vlm_model("JANGQ-AI/Qwen3.5-397B-A17B-JANG_1L")
messages = [{"role": "user", "content": [
{"type": "image"},
{"type": "text", "text": "Describe this image."},
]}]
prompt = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
result = vlm_generate(model, processor, prompt=prompt, image=["photo.jpg"], max_tokens=200)
JANG — Created by Jinho Jang (eric@jangq.ai) · @dealignai
GitHub · PyPI · HuggingFace
JANG_1L은 Qwen3.5-397B를 128 GB Mac에서 실행할 수 있는 최초의 양자화입니다. 112 GB, 36 tok/s, 86.5% MMLU.
pip install "jang[mlx]>=2.1.5"
12 commits
3
stars
12
commits
1
linked in READMEs
Sep 8, 2026
updated
397B intelligence on a 128 GB Mac. This model fits in 112 GB — the first 397B quantization that runs on M4 Pro/Max 128 GB machines. Uses reasoning mode for 86.5% MMLU accuracy.
LM Studio, Ollama, oMLX do NOT support JANG format. Use MLX Studio or
pip install "jang[mlx]>=2.1.5".
JANG — Jang Adaptive N-bit Grading | The GGUF Equivalent for MLX
JANG is fully open-source. Quantization engine, research, and full commit history: github.com/jjang-ai/jangq. Created by Jinho Jang.
<think>...</think> step-by-step problem solvingPer-subject comparison across all modes. Both JANG and MLX 4-bit tested with and without reasoning.
| Subject | JANG No-Think | JANG Reasoning | MLX 4-bit No-Think | MLX 4-bit Reasoning |
|---|---|---|---|---|
| Abstract Algebra | 8/20 | 10/20 | 10/20 | 17/20 |
| Anatomy | 17/20 | 19/20 | 18/20 | 19/20 |
| Astronomy | 20/20 | 20/20 | 19/20 | 19/20 |
| College CS | 17/20 | 18/20 | 15/20 | 18/20 |
| College Physics | 17/20 | 18/20 | 15/20 | 19/20 |
| HS Biology | 19/20 | 20/20 | 19/20 | 19/20 |
| HS Chemistry | 17/20 | 18/20 | 17/20 | 19/20 |
| HS Mathematics | 8/20 | 10/20 | 12/20 | 19/20 |
| Logical Fallacies | 20/20 | 20/20 | 19/20 | 20/20 |
| World Religions | 19/20 | 20/20 | 19/20 | 19/20 |
| Total | 162/200 (81.0%) | 173/200 (86.5%) | 163/200 (81.5%) | 188/200 (94.0%) |
| JANG_1L | JANG_2L | MLX 4-bit | MLX 2/3-bit | |
|---|---|---|---|---|
| MMLU (no-think) | 81.0% | 79.5% | 81.5% | NaN -- cannot run |
| MMLU (reasoning) | 86.5% | 92.0% | 94.0% | NaN -- cannot run |
| Size | 112 GB | 187 GB | 209 GB | N/A |
| GPU RAM | 110 GB | 184 GB | ~210 GB | N/A |
| Speed | 36.1 tok/s | 36.0 tok/s | ~36 tok/s | N/A |
| Fits 128 GB? | YES | No | No | N/A |
JANG_1L is 97 GB smaller than MLX 4-bit and fits on 128 GB Macs where MLX 4-bit (209 GB) cannot run. MLX 2-bit and 3-bit produce NaN -- cannot run (float16 overflow on 512-expert models). JANG solves this with bfloat16.
| Metric | Value |
|---|---|
| Source | Qwen3.5-397B-A17B |
| Architecture | Hybrid MoE + SSM (GatedDeltaNet + Full Attention) |
| Experts | 512 per layer, top-10 active (17B active params) |
| Profile | JANG_1L (CRITICAL=8, IMPORTANT=8, COMPRESS=2) |
| Average bits | 2.13 bpw |
| Disk size | 112 GB |
| GPU RAM | 110 GB (peak 120 GB) |
| Speed | 36.1 tok/s generation, 96 tok/s prefill |
| Compute | bfloat16 (auto-detected) |
| VLM | 333 vision tensors, 31.6 tok/s |
pip install "jang[mlx]>=2.1.5"pip install "jang[mlx]>=2.1.5"
from jang_tools.loader import load_jang_model
from mlx_lm import generate
model, tokenizer = load_jang_model("JANGQ-AI/Qwen3.5-397B-A17B-JANG_1L")
# With reasoning
messages = [{"role": "user", "content": "Prove that sqrt(2) is irrational."}]
prompt = tokenizer.apply_chat_template(messages, tokenize=False,
add_generation_prompt=True, enable_thinking=True)
result = generate(model, tokenizer, prompt=prompt, max_tokens=2048)
# Without reasoning (faster)
prompt = tokenizer.apply_chat_template(messages, tokenize=False,
add_generation_prompt=True, enable_thinking=False)
result = generate(model, tokenizer, prompt=prompt, max_tokens=100)
from jang_tools.loader import load_jang_vlm_model
from mlx_vlm import generate as vlm_generate
model, processor = load_jang_vlm_model("JANGQ-AI/Qwen3.5-397B-A17B-JANG_1L")
messages = [{"role": "user", "content": [
{"type": "image"},
{"type": "text", "text": "Describe this image."},
]}]
prompt = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
result = vlm_generate(model, processor, prompt=prompt, image=["photo.jpg"], max_tokens=200)
JANG — Created by Jinho Jang (eric@jangq.ai) · @dealignai
GitHub · PyPI · HuggingFace
JANG_1L은 Qwen3.5-397B를 128 GB Mac에서 실행할 수 있는 최초의 양자화입니다. 112 GB, 36 tok/s, 86.5% MMLU.
pip install "jang[mlx]>=2.1.5"
12 commits