A from-scratch Mixture-of-Experts LLM: 3.87B total / 1.45B active parameters, pretrained on 86.5B tokens and instruction-tuned with SFT.
This checkpoint is the SFT release (pretrain β 2-stage SFT). A DPO stage was tried and dropped because it lowered code, math and instruction-following scores (see below).
| Architecture | Decoder-only MoE Transformer, every layer MoE. GQA (16Q/4KV, head_dim 128) + QK-RMSNorm + RoPE (ΞΈ=10K) + SwiGLU experts + RMSNorm, weight tying |
| Parameters | 3,869.1M total Β· 1,453.2M active per token (1,142.0M non-embedding active), measured |
| Layers Β· d_model | 32 Β· 2048 |
| Experts | 16 per layer, top-4 routing with renormalized probabilities, expert width 1024 |
| Router losses | load-balancing aux loss 0.01 (global-batch statistics) + router z-loss 1e-3 |
| Context length | 4096 |
| Vocab | 151,936 (Qwen3 tokenizer), tied embeddings |
| Weights | stored as bfloat16 (~7.7GB), cast from the float32 training export |
| Pretrain | 54,250 steps Β· 86.5B tokens (web Β· code in 11 languages Β· math Β· curated), 1M-token batches; the last 20.4B tokens are an LR decay phase (web 30 / code 35 / math 15 / curated 20); parts trained with DiLoCo on 2Γ GH200 |
| Post-training | SFT stage 1: 2.21B tokens (general 29 / code 44 / math 27). SFT stage 2: 0.28B tokens Γ 2 epochs (execution-verified code, competition math, precise instruction following) |
| HF format | Weights map 1:1 onto Qwen3MoeForCausalLM (the training model uses fused expert tensors; the export splits them per expert). Export check: logits relative difference 1.0e-6 and 100% next-token agreement against the training checkpoint |
The model uses the ChatML template (<|im_start|>role\n...<|im_end|>) shipped in tokenizer_config.json; generation ends at <|im_end|>.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("YOON1v/Apex-2", dtype=torch.bfloat16, device_map="cuda")
tokenizer = AutoTokenizer.from_pretrained("YOON1v/Apex-2")
messages = [{"role": "user", "content": "Write a Python function that checks if a number is prime."}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt", return_dict=True).to("cuda")
out = model.generate(**inputs, max_new_tokens=512, do_sample=True, temperature=0.7, top_p=0.9)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
vLLM loads it directly as well: LLM(model="YOON1v/Apex-2", dtype="bfloat16", max_model_len=4096).
Measured by us with vLLM greedy decoding and the chat template (prompt + answer β€ 4096 tokens). Model-written code was executed in a sandbox.
| Area | Benchmark | Apex-2 (SFT) | Base |
|---|---|---|---|
| Code | HumanEval | 43.9 | 36.6 |
| Code | HumanEval+ | 41.5 | 32.9 |
| Code | MBPP | 56.3 | 54.8 |
| Code | MBPP+ | 48.9 | 46.3 |
| Code | MultiPL-E HumanEval C++ | 36.0 | β |
| Code | MultiPL-E MBPP C++ | 41.6 | β |
| Code | LiveCodeBench v5βv6 | 3.2 | β |
| Code | β LiveCodeBench easy (84) | 11.9 | β |
| Code | β LiveCodeBench medium (104) | 1.0 | β |
| Code | β LiveCodeBench hard (154) | 0.0 | β |
| Code | CRUXEval-O | 8.5 | β |
| Code | CRUXEval-I | 3.2 | β |
| Math | GSM8K | 32.4 (0-shot CoT) | 15.1 (8-shot) |
| Math | MATH-500 (0-shot CoT) | 21.0 | β |
| Instruction following | IFEval prompt strict | 44.7 | β |
| Instruction following | IFEval instruction strict | 56.6 | β |
| Knowledge | MMLU (5-shot) | 28.6 | 28.2 |
| Commonsense | HellaSwag (acc_norm) | 62.2 | 60.8 |
| Commonsense | ARC-e (acc_norm) | 64.1 | 66.7 |
| Commonsense | ARC-c (acc_norm) | 38.4 | 39.7 |
| Commonsense | PIQA (acc_norm) | 74.8 | 73.9 |
| Commonsense | WinoGrande (acc) | 61.8 | 60.1 |
| Commonsense | LAMBADA (acc) | 52.8 | 54.5 |
Compared with similar-size models (official numbers; protocols differ, so this is a rough comparison):
| Model | Pretraining tokens | HumanEval | HumanEval+ | MBPP | MBPP+ | GSM8K | IFEval | MMLU |
|---|---|---|---|---|---|---|---|---|
| Apex-2 (3.87B Β· 1.45B active) | 0.087T | 43.9 | 41.5 | 56.3 | 48.9 | 32.4 | 44.7 | 28.6 |
| Apex-1 DPO (1.1B dense) | 0.02T | 8.5 | β | 5.2 | β | 1.9 | β | 24.9 |
| Qwen2.5-1.5B-Instruct | 18T | 61.6 | β | 63.2 | β | 73.2 | 42.5 | 50.7 |
| Qwen2.5-Coder-1.5B-Instruct | 5.5T | 70.7 | 66.5 | 69.2 | 59.4 | β | β | β |
| Qwen3-1.7B (non-thinking) | 36T | β | β | β | β | β | 68.2 | 64.4 |
| Llama-3.2-1B-Instruct | 9T | β | β | β | β | 44.4 | 59.5* | 49.3 |
| Gemma-3-1B-it | 2T | 41.5 | β | 35.2 | β | 62.8 | 80.2* | 38.8 |
| OLMoE-1B-7B (1.3B active) | 5.1T | 62.3 | 54.4 | β | β | 72.4 | 66.4* | 55.1 |
| DeepSeek-Coder-1.3B | 2T | 65.9 | 60.4 | 65.3 | 54.8 | β | β | β |
DeepSeek-Coder numbers are the Qwen2.5-Coder report's re-evaluation. * Different IFEval metric: Apex-2 and Qwen report prompt-level strict; the others report an average of metrics or do not specify.
With 1/23 to 1/400 of the pretraining data of these models, Apex-2's base model matches Qwen2.5-1.5B (18T tokens) on HumanEval+ (32.9 vs 32.9). The big gaps are knowledge (MMLU) and math, which mainly track pretraining scale.
[!NOTE] SFT data was 13-gram decontaminated against HumanEval(+), MBPP(+), GSM8K test, MATH, MMLU, ARC, HellaSwag, PIQA, WinoGrande, IFEval and LAMBADA. For LiveCodeBench only the newer v5βv6 problems (after 2024-08) were used, to reduce contamination risk.
Why no DPO: DPO on allenai/Dolci-Instruct-DPO (220K pairs) made answers 2.3Γ longer (320 β 733 tokens on average). Chosen-answer likelihood also fell during training. In evaluation HumanEval+ went 41.5 β 32.3, MBPP+ 48.9 β 37.3, GSM8K 32.4 β 14.3 and IFEval 44.7 β 35.7, so this SFT checkpoint is the release.
Some SFT sources contain synthetic data generated by third-party models (e.g. GPT-4-class, Llama-3.1-405B, Qwen2.5-Coder); check each dataset's terms for your use case.
Apache 2.0.
A from-scratch Mixture-of-Experts LLM: 3.87B total / 1.45B active parameters, pretrained on 86.5B tokens and instruction-tuned with SFT.
This checkpoint is the SFT release (pretrain β 2-stage SFT). A DPO stage was tried and dropped because it lowered code, math and instruction-following scores (see below).
| Architecture | Decoder-only MoE Transformer, every layer MoE. GQA (16Q/4KV, head_dim 128) + QK-RMSNorm + RoPE (ΞΈ=10K) + SwiGLU experts + RMSNorm, weight tying |
| Parameters | 3,869.1M total Β· 1,453.2M active per token (1,142.0M non-embedding active), measured |
| Layers Β· d_model | 32 Β· 2048 |
| Experts | 16 per layer, top-4 routing with renormalized probabilities, expert width 1024 |
| Router losses | load-balancing aux loss 0.01 (global-batch statistics) + router z-loss 1e-3 |
| Context length | 4096 |
| Vocab | 151,936 (Qwen3 tokenizer), tied embeddings |
| Weights | stored as bfloat16 (~7.7GB), cast from the float32 training export |
| Pretrain | 54,250 steps Β· 86.5B tokens (web Β· code in 11 languages Β· math Β· curated), 1M-token batches; the last 20.4B tokens are an LR decay phase (web 30 / code 35 / math 15 / curated 20); parts trained with DiLoCo on 2Γ GH200 |
| Post-training | SFT stage 1: 2.21B tokens (general 29 / code 44 / math 27). SFT stage 2: 0.28B tokens Γ 2 epochs (execution-verified code, competition math, precise instruction following) |
| HF format | Weights map 1:1 onto Qwen3MoeForCausalLM (the training model uses fused expert tensors; the export splits them per expert). Export check: logits relative difference 1.0e-6 and 100% next-token agreement against the training checkpoint |
The model uses the ChatML template (<|im_start|>role\n...<|im_end|>) shipped in tokenizer_config.json; generation ends at <|im_end|>.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("YOON1v/Apex-2", dtype=torch.bfloat16, device_map="cuda")
tokenizer = AutoTokenizer.from_pretrained("YOON1v/Apex-2")
messages = [{"role": "user", "content": "Write a Python function that checks if a number is prime."}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt", return_dict=True).to("cuda")
out = model.generate(**inputs, max_new_tokens=512, do_sample=True, temperature=0.7, top_p=0.9)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
vLLM loads it directly as well: LLM(model="YOON1v/Apex-2", dtype="bfloat16", max_model_len=4096).
Measured by us with vLLM greedy decoding and the chat template (prompt + answer β€ 4096 tokens). Model-written code was executed in a sandbox.
| Area | Benchmark | Apex-2 (SFT) | Base |
|---|---|---|---|
| Code | HumanEval | 43.9 | 36.6 |
| Code | HumanEval+ | 41.5 | 32.9 |
| Code | MBPP | 56.3 | 54.8 |
| Code | MBPP+ | 48.9 | 46.3 |
| Code | MultiPL-E HumanEval C++ | 36.0 | β |
| Code | MultiPL-E MBPP C++ | 41.6 | β |
| Code | LiveCodeBench v5βv6 | 3.2 | β |
| Code | β LiveCodeBench easy (84) | 11.9 | β |
| Code | β LiveCodeBench medium (104) | 1.0 | β |
| Code | β LiveCodeBench hard (154) | 0.0 | β |
| Code | CRUXEval-O | 8.5 | β |
| Code | CRUXEval-I | 3.2 | β |
| Math | GSM8K | 32.4 (0-shot CoT) | 15.1 (8-shot) |
| Math | MATH-500 (0-shot CoT) | 21.0 | β |
| Instruction following | IFEval prompt strict | 44.7 | β |
| Instruction following | IFEval instruction strict | 56.6 | β |
| Knowledge | MMLU (5-shot) | 28.6 | 28.2 |
| Commonsense | HellaSwag (acc_norm) | 62.2 | 60.8 |
| Commonsense | ARC-e (acc_norm) | 64.1 | 66.7 |
| Commonsense | ARC-c (acc_norm) | 38.4 | 39.7 |
| Commonsense | PIQA (acc_norm) | 74.8 | 73.9 |
| Commonsense | WinoGrande (acc) | 61.8 | 60.1 |
| Commonsense | LAMBADA (acc) | 52.8 | 54.5 |
Compared with similar-size models (official numbers; protocols differ, so this is a rough comparison):
| Model | Pretraining tokens | HumanEval | HumanEval+ | MBPP | MBPP+ | GSM8K | IFEval | MMLU |
|---|---|---|---|---|---|---|---|---|
| Apex-2 (3.87B Β· 1.45B active) | 0.087T | 43.9 | 41.5 | 56.3 | 48.9 | 32.4 | 44.7 | 28.6 |
| Apex-1 DPO (1.1B dense) | 0.02T | 8.5 | β | 5.2 | β | 1.9 | β | 24.9 |
| Qwen2.5-1.5B-Instruct | 18T | 61.6 | β | 63.2 | β | 73.2 | 42.5 | 50.7 |
| Qwen2.5-Coder-1.5B-Instruct | 5.5T | 70.7 | 66.5 | 69.2 | 59.4 | β | β | β |
| Qwen3-1.7B (non-thinking) | 36T | β | β | β | β | β | 68.2 | 64.4 |
| Llama-3.2-1B-Instruct | 9T | β | β | β | β | 44.4 | 59.5* | 49.3 |
| Gemma-3-1B-it | 2T | 41.5 | β | 35.2 | β | 62.8 | 80.2* | 38.8 |
| OLMoE-1B-7B (1.3B active) | 5.1T | 62.3 | 54.4 | β | β | 72.4 | 66.4* | 55.1 |
| DeepSeek-Coder-1.3B | 2T | 65.9 | 60.4 | 65.3 | 54.8 | β | β | β |
DeepSeek-Coder numbers are the Qwen2.5-Coder report's re-evaluation. * Different IFEval metric: Apex-2 and Qwen report prompt-level strict; the others report an average of metrics or do not specify.
With 1/23 to 1/400 of the pretraining data of these models, Apex-2's base model matches Qwen2.5-1.5B (18T tokens) on HumanEval+ (32.9 vs 32.9). The big gaps are knowledge (MMLU) and math, which mainly track pretraining scale.
[!NOTE] SFT data was 13-gram decontaminated against HumanEval(+), MBPP(+), GSM8K test, MATH, MMLU, ARC, HellaSwag, PIQA, WinoGrande, IFEval and LAMBADA. For LiveCodeBench only the newer v5βv6 problems (after 2024-08) were used, to reduce contamination risk.
Why no DPO: DPO on allenai/Dolci-Instruct-DPO (220K pairs) made answers 2.3Γ longer (320 β 733 tokens on average). Chosen-answer likelihood also fell during training. In evaluation HumanEval+ went 41.5 β 32.3, MBPP+ 48.9 β 37.3, GSM8K 32.4 β 14.3 and IFEval 44.7 β 35.7, so this SFT checkpoint is the release.
Some SFT sources contain synthetic data generated by third-party models (e.g. GPT-4-class, Llama-3.1-405B, Qwen2.5-Coder); check each dataset's terms for your use case.
Apache 2.0.