YOON1v/Apex-2

Model

APEX-2

6

3 commits

1 linked in READMEs

updated Sep 29, 2026

See the code

README

APEX-2

A from-scratch Mixture-of-Experts LLM: 3.87B total / 1.45B active parameters, pretrained on 86.5B tokens and instruction-tuned with SFT.

This checkpoint is the SFT release (pretrain β†’ 2-stage SFT). A DPO stage was tried and dropped because it lowered code, math and instruction-following scores (see below).

Model details

ArchitectureDecoder-only MoE Transformer, every layer MoE. GQA (16Q/4KV, head_dim 128) + QK-RMSNorm + RoPE (ΞΈ=10K) + SwiGLU experts + RMSNorm, weight tying
Parameters3,869.1M total Β· 1,453.2M active per token (1,142.0M non-embedding active), measured
Layers Β· d_model32 Β· 2048
Experts16 per layer, top-4 routing with renormalized probabilities, expert width 1024
Router lossesload-balancing aux loss 0.01 (global-batch statistics) + router z-loss 1e-3
Context length4096
Vocab151,936 (Qwen3 tokenizer), tied embeddings
Weightsstored as bfloat16 (~7.7GB), cast from the float32 training export
Pretrain54,250 steps Β· 86.5B tokens (web Β· code in 11 languages Β· math Β· curated), 1M-token batches; the last 20.4B tokens are an LR decay phase (web 30 / code 35 / math 15 / curated 20); parts trained with DiLoCo on 2Γ— GH200
Post-trainingSFT stage 1: 2.21B tokens (general 29 / code 44 / math 27). SFT stage 2: 0.28B tokens Γ— 2 epochs (execution-verified code, competition math, precise instruction following)
HF formatWeights map 1:1 onto Qwen3MoeForCausalLM (the training model uses fused expert tensors; the export splits them per expert). Export check: logits relative difference 1.0e-6 and 100% next-token agreement against the training checkpoint

Usage

The model uses the ChatML template (<|im_start|>role\n...<|im_end|>) shipped in tokenizer_config.json; generation ends at <|im_end|>.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("YOON1v/Apex-2", dtype=torch.bfloat16, device_map="cuda")
tokenizer = AutoTokenizer.from_pretrained("YOON1v/Apex-2")

messages = [{"role": "user", "content": "Write a Python function that checks if a number is prime."}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt", return_dict=True).to("cuda")
out = model.generate(**inputs, max_new_tokens=512, do_sample=True, temperature=0.7, top_p=0.9)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

vLLM loads it directly as well: LLM(model="YOON1v/Apex-2", dtype="bfloat16", max_model_len=4096).

Benchmarks

Measured by us with vLLM greedy decoding and the chat template (prompt + answer ≀ 4096 tokens). Model-written code was executed in a sandbox.

AreaBenchmarkApex-2 (SFT)Base
CodeHumanEval43.936.6
CodeHumanEval+41.532.9
CodeMBPP56.354.8
CodeMBPP+48.946.3
CodeMultiPL-E HumanEval C++36.0β€”
CodeMultiPL-E MBPP C++41.6β€”
CodeLiveCodeBench v5–v63.2β€”
Codeβ”” LiveCodeBench easy (84)11.9β€”
Codeβ”” LiveCodeBench medium (104)1.0β€”
Codeβ”” LiveCodeBench hard (154)0.0β€”
CodeCRUXEval-O8.5β€”
CodeCRUXEval-I3.2β€”
MathGSM8K32.4 (0-shot CoT)15.1 (8-shot)
MathMATH-500 (0-shot CoT)21.0β€”
Instruction followingIFEval prompt strict44.7β€”
Instruction followingIFEval instruction strict56.6β€”
KnowledgeMMLU (5-shot)28.628.2
CommonsenseHellaSwag (acc_norm)62.260.8
CommonsenseARC-e (acc_norm)64.166.7
CommonsenseARC-c (acc_norm)38.439.7
CommonsensePIQA (acc_norm)74.873.9
CommonsenseWinoGrande (acc)61.860.1
CommonsenseLAMBADA (acc)52.854.5

Compared with similar-size models (official numbers; protocols differ, so this is a rough comparison):

ModelPretraining tokensHumanEvalHumanEval+MBPPMBPP+GSM8KIFEvalMMLU
Apex-2 (3.87B Β· 1.45B active)0.087T43.941.556.348.932.444.728.6
Apex-1 DPO (1.1B dense)0.02T8.5β€”5.2β€”1.9β€”24.9
Qwen2.5-1.5B-Instruct18T61.6β€”63.2β€”73.242.550.7
Qwen2.5-Coder-1.5B-Instruct5.5T70.766.569.259.4β€”β€”β€”
Qwen3-1.7B (non-thinking)36Tβ€”β€”β€”β€”β€”68.264.4
Llama-3.2-1B-Instruct9Tβ€”β€”β€”β€”44.459.5*49.3
Gemma-3-1B-it2T41.5β€”35.2β€”62.880.2*38.8
OLMoE-1B-7B (1.3B active)5.1T62.354.4β€”β€”72.466.4*55.1
DeepSeek-Coder-1.3B2T65.960.465.354.8β€”β€”β€”

DeepSeek-Coder numbers are the Qwen2.5-Coder report's re-evaluation. * Different IFEval metric: Apex-2 and Qwen report prompt-level strict; the others report an average of metrics or do not specify.

With 1/23 to 1/400 of the pretraining data of these models, Apex-2's base model matches Qwen2.5-1.5B (18T tokens) on HumanEval+ (32.9 vs 32.9). The big gaps are knowledge (MMLU) and math, which mainly track pretraining scale.

[!NOTE] SFT data was 13-gram decontaminated against HumanEval(+), MBPP(+), GSM8K test, MATH, MMLU, ARC, HellaSwag, PIQA, WinoGrande, IFEval and LAMBADA. For LiveCodeBench only the newer v5–v6 problems (after 2024-08) were used, to reduce contamination risk.

Why no DPO: DPO on allenai/Dolci-Instruct-DPO (220K pairs) made answers 2.3Γ— longer (320 β†’ 733 tokens on average). Chosen-answer likelihood also fell during training. In evaluation HumanEval+ went 41.5 β†’ 32.3, MBPP+ 48.9 β†’ 37.3, GSM8K 32.4 β†’ 14.3 and IFEval 44.7 β†’ 35.7, so this SFT checkpoint is the release.

Known limitations

  • English-centric. Other languages, including Korean, barely work (multilingual data was excluded from SFT).
  • Limited world knowledge (MMLU ~29%), so it often states wrong facts with confidence.
  • Weak at competitive programming (LiveCodeBench medium ~1%) and at predicting code execution (CRUXEval).
  • 4096-token context. No tool use, no thinking mode.

Training data

  • Pretrain: 40 deduplicated sources, including:
    • web: FineWeb-Edu, DCLM, FinePDFs
    • code: StarCoder-family and The Stack-derived code in 11 languages, OpenCoder corpora, synthetic code
    • math: FineMath, InfiWebMath, Nemotron-CC-Math
    • curated: Wikipedia, StackExchange, peS2o, Cosmopedia, Gutenberg
  • SFT:
    • general: allenai/Dolci-Instruct-SFT (tool-use, multilingual and identity subsets removed)
    • code: OpenCoder opc-sft-stage1/2, nvidia/OpenCodeInstruct (unit-test pass rate β‰₯ 0.9), bigcode self-oss-instruct, m-a-p/Code-Feedback
    • math: nvidia/OpenMathInstruct-2, AI-MO/NuminaMath-CoT

Some SFT sources contain synthetic data generated by third-party models (e.g. GPT-4-class, Llama-3.1-405B, Qwen2.5-Coder); check each dataset's terms for your use case.

License

Apache 2.0.

causal-lm
code
conversational
endpoints_compatible
from-scratch
mixture-of-experts
model-index
qwen3_moe
safetensors
text-generation
transformers

YOON1v/Apex-2

Model

APEX-2

6

3 commits

1 linked in READMEs

updated Sep 29, 2026

See the code

README

APEX-2

A from-scratch Mixture-of-Experts LLM: 3.87B total / 1.45B active parameters, pretrained on 86.5B tokens and instruction-tuned with SFT.

This checkpoint is the SFT release (pretrain β†’ 2-stage SFT). A DPO stage was tried and dropped because it lowered code, math and instruction-following scores (see below).

Model details

ArchitectureDecoder-only MoE Transformer, every layer MoE. GQA (16Q/4KV, head_dim 128) + QK-RMSNorm + RoPE (ΞΈ=10K) + SwiGLU experts + RMSNorm, weight tying
Parameters3,869.1M total Β· 1,453.2M active per token (1,142.0M non-embedding active), measured
Layers Β· d_model32 Β· 2048
Experts16 per layer, top-4 routing with renormalized probabilities, expert width 1024
Router lossesload-balancing aux loss 0.01 (global-batch statistics) + router z-loss 1e-3
Context length4096
Vocab151,936 (Qwen3 tokenizer), tied embeddings
Weightsstored as bfloat16 (~7.7GB), cast from the float32 training export
Pretrain54,250 steps Β· 86.5B tokens (web Β· code in 11 languages Β· math Β· curated), 1M-token batches; the last 20.4B tokens are an LR decay phase (web 30 / code 35 / math 15 / curated 20); parts trained with DiLoCo on 2Γ— GH200
Post-trainingSFT stage 1: 2.21B tokens (general 29 / code 44 / math 27). SFT stage 2: 0.28B tokens Γ— 2 epochs (execution-verified code, competition math, precise instruction following)
HF formatWeights map 1:1 onto Qwen3MoeForCausalLM (the training model uses fused expert tensors; the export splits them per expert). Export check: logits relative difference 1.0e-6 and 100% next-token agreement against the training checkpoint

Usage

The model uses the ChatML template (<|im_start|>role\n...<|im_end|>) shipped in tokenizer_config.json; generation ends at <|im_end|>.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("YOON1v/Apex-2", dtype=torch.bfloat16, device_map="cuda")
tokenizer = AutoTokenizer.from_pretrained("YOON1v/Apex-2")

messages = [{"role": "user", "content": "Write a Python function that checks if a number is prime."}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt", return_dict=True).to("cuda")
out = model.generate(**inputs, max_new_tokens=512, do_sample=True, temperature=0.7, top_p=0.9)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

vLLM loads it directly as well: LLM(model="YOON1v/Apex-2", dtype="bfloat16", max_model_len=4096).

Benchmarks

Measured by us with vLLM greedy decoding and the chat template (prompt + answer ≀ 4096 tokens). Model-written code was executed in a sandbox.

AreaBenchmarkApex-2 (SFT)Base
CodeHumanEval43.936.6
CodeHumanEval+41.532.9
CodeMBPP56.354.8
CodeMBPP+48.946.3
CodeMultiPL-E HumanEval C++36.0β€”
CodeMultiPL-E MBPP C++41.6β€”
CodeLiveCodeBench v5–v63.2β€”
Codeβ”” LiveCodeBench easy (84)11.9β€”
Codeβ”” LiveCodeBench medium (104)1.0β€”
Codeβ”” LiveCodeBench hard (154)0.0β€”
CodeCRUXEval-O8.5β€”
CodeCRUXEval-I3.2β€”
MathGSM8K32.4 (0-shot CoT)15.1 (8-shot)
MathMATH-500 (0-shot CoT)21.0β€”
Instruction followingIFEval prompt strict44.7β€”
Instruction followingIFEval instruction strict56.6β€”
KnowledgeMMLU (5-shot)28.628.2
CommonsenseHellaSwag (acc_norm)62.260.8
CommonsenseARC-e (acc_norm)64.166.7
CommonsenseARC-c (acc_norm)38.439.7
CommonsensePIQA (acc_norm)74.873.9
CommonsenseWinoGrande (acc)61.860.1
CommonsenseLAMBADA (acc)52.854.5

Compared with similar-size models (official numbers; protocols differ, so this is a rough comparison):

ModelPretraining tokensHumanEvalHumanEval+MBPPMBPP+GSM8KIFEvalMMLU
Apex-2 (3.87B Β· 1.45B active)0.087T43.941.556.348.932.444.728.6
Apex-1 DPO (1.1B dense)0.02T8.5β€”5.2β€”1.9β€”24.9
Qwen2.5-1.5B-Instruct18T61.6β€”63.2β€”73.242.550.7
Qwen2.5-Coder-1.5B-Instruct5.5T70.766.569.259.4β€”β€”β€”
Qwen3-1.7B (non-thinking)36Tβ€”β€”β€”β€”β€”68.264.4
Llama-3.2-1B-Instruct9Tβ€”β€”β€”β€”44.459.5*49.3
Gemma-3-1B-it2T41.5β€”35.2β€”62.880.2*38.8
OLMoE-1B-7B (1.3B active)5.1T62.354.4β€”β€”72.466.4*55.1
DeepSeek-Coder-1.3B2T65.960.465.354.8β€”β€”β€”

DeepSeek-Coder numbers are the Qwen2.5-Coder report's re-evaluation. * Different IFEval metric: Apex-2 and Qwen report prompt-level strict; the others report an average of metrics or do not specify.

With 1/23 to 1/400 of the pretraining data of these models, Apex-2's base model matches Qwen2.5-1.5B (18T tokens) on HumanEval+ (32.9 vs 32.9). The big gaps are knowledge (MMLU) and math, which mainly track pretraining scale.

[!NOTE] SFT data was 13-gram decontaminated against HumanEval(+), MBPP(+), GSM8K test, MATH, MMLU, ARC, HellaSwag, PIQA, WinoGrande, IFEval and LAMBADA. For LiveCodeBench only the newer v5–v6 problems (after 2024-08) were used, to reduce contamination risk.

Why no DPO: DPO on allenai/Dolci-Instruct-DPO (220K pairs) made answers 2.3Γ— longer (320 β†’ 733 tokens on average). Chosen-answer likelihood also fell during training. In evaluation HumanEval+ went 41.5 β†’ 32.3, MBPP+ 48.9 β†’ 37.3, GSM8K 32.4 β†’ 14.3 and IFEval 44.7 β†’ 35.7, so this SFT checkpoint is the release.

Known limitations

  • English-centric. Other languages, including Korean, barely work (multilingual data was excluded from SFT).
  • Limited world knowledge (MMLU ~29%), so it often states wrong facts with confidence.
  • Weak at competitive programming (LiveCodeBench medium ~1%) and at predicting code execution (CRUXEval).
  • 4096-token context. No tool use, no thinking mode.

Training data

  • Pretrain: 40 deduplicated sources, including:
    • web: FineWeb-Edu, DCLM, FinePDFs
    • code: StarCoder-family and The Stack-derived code in 11 languages, OpenCoder corpora, synthetic code
    • math: FineMath, InfiWebMath, Nemotron-CC-Math
    • curated: Wikipedia, StackExchange, peS2o, Cosmopedia, Gutenberg
  • SFT:
    • general: allenai/Dolci-Instruct-SFT (tool-use, multilingual and identity subsets removed)
    • code: OpenCoder opc-sft-stage1/2, nvidia/OpenCodeInstruct (unit-test pass rate β‰₯ 0.9), bigcode self-oss-instruct, m-a-p/Code-Feedback
    • math: nvidia/OpenMathInstruct-2, AI-MO/NuminaMath-CoT

Some SFT sources contain synthetic data generated by third-party models (e.g. GPT-4-class, Llama-3.1-405B, Qwen2.5-Coder); check each dataset's terms for your use case.

License

Apache 2.0.

causal-lm
code
conversational
endpoints_compatible
from-scratch
mixture-of-experts
model-index
qwen3_moe
safetensors
text-generation
transformers