Developed by Empero
[!Note] This repository contains model weights and configuration files in the Hugging Face Transformers format.
These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, and other standard runtimes with Qwen3.6 architecture support.
Qwen3.8-35B-A3B is a distillation of the Qwen3.8 frontier models into the Qwen3.6-35B-A3B Mixture-of-Experts architecture. The student was trained on curated teacher traces from our internal Qwen3.8 distillation datasets β dense chain-of-thought spanning mathematics, code, general reasoning, instruction following, and tool use, quality-filtered before training.
The objective: bring the reasoning behavior of frontier-scale teachers into a sparse 35B that activates only 3B parameters per token and deploys on a single GPU.
<think> block learned directly from Qwen3.8 teacher traces rather than synthetic self-generated reasoning.Measured with lm-evaluation-harness, HF backend, bfloat16, identical settings and seed for base and student. Zero-shot, loglikelihood scoring.
| Task | Metric | Qwen3.6-35B-A3B (base) | Qwen3.8-35B-A3B | Ξ |
|---|---|---|---|---|
| MMLU (57 subjects) | acc | 0.838 | 0.834 | β0.004 |
| ARC-Challenge | acc | 0.548 | 0.582 | +0.034 |
| ARC-Challenge | acc_norm | 0.548 | 0.591 | +0.044 |
| ARC-Easy | acc | 0.819 | 0.830 | +0.011 |
| ARC-Easy | acc_norm | 0.717 | 0.766 | +0.048 |
The MMLU difference is within noise (standard error 0.003 on each measurement). The ARC gains are outside it.
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "empero-ai/Qwen3.8-35B-A3B-Distilled"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16, device_map="auto")
messages = [{"role": "user", "content": "A snail is at the bottom of a 10-meter well. Each day it climbs 3 meters, each night it slips back 2. How many days until it escapes?"}]
inputs = tok.apply_chat_template(messages, add_generation_prompt=True,
return_tensors="pt", return_dict=True).to(model.device)
out = model.generate(**inputs, max_new_tokens=16384,
temperature=0.6, top_p=0.95, top_k=20, do_sample=True)
print(tok.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
A recent transformers release with Qwen3.6 support is required, along with the Gated DeltaNet kernels (flash-linear-attention and a CUDA-matched causal_conv1d build) β without them the linear-attention layers fall back to slow, memory-hungry PyTorch ops.
AutoModelForCausalLM loads the text path (34.7B parameters). The vision tower is retained in the checkpoint and is reachable via AutoModelForImageTextToText.
temperature=0.6, top_p=0.95, top_k=20. Greedy decoding on long generations is a known repetition-loop failure mode for reasoning models in this class.max_new_tokens (16,384 recommended); every answer opens with a <think> block. Parse and strip the <think>...</think> span for end users.Sign up for the Empero newsletter at empero.org for releases, evals, and research notes.
If this model helped you, consider supporting the project:
bc1qx6zepu6sfkvshgdmc4ewu6pk6rpadvpgffpp7vltc1qv2mefzps2vtjcpwfx8xxdrpplrcvltswm68r7xWeights are released under Apache-2.0, inherited from the Qwen3.6-35B-A3B base. Shared for research and experimentation, as-is.
Developed by Empero
[!Note] This repository contains model weights and configuration files in the Hugging Face Transformers format.
These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, and other standard runtimes with Qwen3.6 architecture support.
Qwen3.8-35B-A3B is a distillation of the Qwen3.8 frontier models into the Qwen3.6-35B-A3B Mixture-of-Experts architecture. The student was trained on curated teacher traces from our internal Qwen3.8 distillation datasets β dense chain-of-thought spanning mathematics, code, general reasoning, instruction following, and tool use, quality-filtered before training.
The objective: bring the reasoning behavior of frontier-scale teachers into a sparse 35B that activates only 3B parameters per token and deploys on a single GPU.
<think> block learned directly from Qwen3.8 teacher traces rather than synthetic self-generated reasoning.Measured with lm-evaluation-harness, HF backend, bfloat16, identical settings and seed for base and student. Zero-shot, loglikelihood scoring.
| Task | Metric | Qwen3.6-35B-A3B (base) | Qwen3.8-35B-A3B | Ξ |
|---|---|---|---|---|
| MMLU (57 subjects) | acc | 0.838 | 0.834 | β0.004 |
| ARC-Challenge | acc | 0.548 | 0.582 | +0.034 |
| ARC-Challenge | acc_norm | 0.548 | 0.591 | +0.044 |
| ARC-Easy | acc | 0.819 | 0.830 | +0.011 |
| ARC-Easy | acc_norm | 0.717 | 0.766 | +0.048 |
The MMLU difference is within noise (standard error 0.003 on each measurement). The ARC gains are outside it.
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "empero-ai/Qwen3.8-35B-A3B-Distilled"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16, device_map="auto")
messages = [{"role": "user", "content": "A snail is at the bottom of a 10-meter well. Each day it climbs 3 meters, each night it slips back 2. How many days until it escapes?"}]
inputs = tok.apply_chat_template(messages, add_generation_prompt=True,
return_tensors="pt", return_dict=True).to(model.device)
out = model.generate(**inputs, max_new_tokens=16384,
temperature=0.6, top_p=0.95, top_k=20, do_sample=True)
print(tok.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
A recent transformers release with Qwen3.6 support is required, along with the Gated DeltaNet kernels (flash-linear-attention and a CUDA-matched causal_conv1d build) β without them the linear-attention layers fall back to slow, memory-hungry PyTorch ops.
AutoModelForCausalLM loads the text path (34.7B parameters). The vision tower is retained in the checkpoint and is reachable via AutoModelForImageTextToText.
temperature=0.6, top_p=0.95, top_k=20. Greedy decoding on long generations is a known repetition-loop failure mode for reasoning models in this class.max_new_tokens (16,384 recommended); every answer opens with a <think> block. Parse and strip the <think>...</think> span for end users.Sign up for the Empero newsletter at empero.org for releases, evals, and research notes.
If this model helped you, consider supporting the project:
bc1qx6zepu6sfkvshgdmc4ewu6pk6rpadvpgffpp7vltc1qv2mefzps2vtjcpwfx8xxdrpplrcvltswm68r7xWeights are released under Apache-2.0, inherited from the Qwen3.6-35B-A3B base. Shared for research and experimentation, as-is.