Note: This is a toy project designed to explore and understand the model inference structure of the
kailibrary. It focuses on architectural mapping and experimentation rather than production use.
This project provides a unified inference engine and CLI for South Korea's top Sovereign AI LLM models:
The codebase allows for checking architecture compatibility, creating random weights for testing, and running inference (text generation/chat) through a unified interface.
| Model | Developer | Params (Total/Active) | Architecture | HuggingFace |
|---|---|---|---|---|
| K-EXAONE | LG AI Research | 236B / 23B | MoE, LLLG Hybrid Attention | LGAI-EXAONE/K-EXAONE-236B-A23B |
| Solar-Open | Upstage | 102.6B / 12B | MoE, GQA | upstage/Solar-Open-100B |
| A.X-K1 | SK Telecom | 519B / 33B | MoE (1 Dense + 60 MoE) | skt/A.X-K1 |
| VAETKI | NC AI | 112.2B / 10.1B | MoE, Edge Optimized | NC-AI-consortium-VAETKI/VAETKI |
| HyperCLOVA Omni | Naver | 8B | Dense, Multimodal | naver-hyperclovax/HyperCLOVAX-SEED-Omni-8B |
| HyperCLOVA Think | Naver | 32B | Dense, VLM + Thinking | naver-hyperclovax/HyperCLOVAX-SEED-Think-32B |
| Spec | K-EXAONE | Solar-Open | A.X-K1 | VAETKI |
|---|---|---|---|---|
| Layers | 48 | 48 | 61 | 48 |
| Hidden Size | 4,096 | 4,096 | 6,144 | 4,096 |
| Attention Heads | 64 | 64 | 64 | 64 |
| KV Heads (GQA) | 8 | 8 | 8 | 8 |
| Routed Experts | 128 | 128 | 192 | 128 |
| Shared Experts | 1 | 1 | 1 | 0 |
| Top-K Selection | 8 | 8 | 8 | 8 |
| MoE Intermediate | 1,280 | 1,280 | 2,048 | 1,280 |
| Dense Intermediate | 10,240 | 10,240 | 7,168 | 10,240 |
| Vocab Size | 153,600 | 196,608 | 163,840 | 128,000 |
| Context Length | 256K | 128K | 128K | 128K |
| RoPE Theta | 1M | 1M | 1M | 1M |
| QK Norm | Yes | No | No | No |
| Attention | LLLG Hybrid | Standard GQA | Standard GQA | Standard GQA |
| Spec | HyperCLOVA Omni | HyperCLOVA Think |
|---|---|---|
| Layers | 36 | 72 |
| Hidden Size | 4,096 | 5,120 |
| Attention Heads | 32 | 40 |
| KV Heads (GQA) | 8 | 8 |
| Head Dim | 128 | 128 |
| Intermediate Size | 12,288 | 24,192 |
| Vocab Size | 200,704 | 128,256 |
| Context Length | 8K (32K with RoPE) | 128K |
| RoPE Theta | 5M | 50M |
| Modality | Text, Vision, Audio | Text, Vision |
| Special Features | Omni-modal I/O | Thinking Mode (<think>) |
# Install dependencies
pip install -r requirements.txt
# Install the package in editable mode
pip install -e .
Run the CLI to interact with a model. You can list available model types:
python cli.py --list-models
To run with a specific model path (e.g., a mini model created for testing):
python cli.py --model-path ./solar_open_mini --model-type solar_open
Since these models are massive, you can create "mini" versions with random weights to test the architecture and memory flow:
# Create Solar-Open mini model
python create_random_weights.py --model-type solar_open
# Create K-EXAONE mini model
python create_random_weights.py --model-type k_exaone
# Create A.X-K1 mini model
python create_random_weights.py --model-type ax_k1
# Create VAETKI mini model
python create_random_weights.py --model-type vaetki
# Create HyperCLOVA Omni mini model
python create_random_weights.py --model-type hyperclova_omni
# Create HyperCLOVA Think mini model
python create_random_weights.py --model-type hyperclova_think
These will create directories like ./solar_open_mini containing config.json, model.safetensors, and tokenizer files downloaded from HuggingFace.
To skip tokenizer download:
python create_random_weights.py --model-type hyperclova_omni --no-tokenizer
You can use the KAIInferenceEngine in your Python code:
from kai.inference import KAIInferenceEngine
# Load engine
engine = KAIInferenceEngine(
model_path="./k_exaone_mini",
model_type="k_exaone",
device="cuda" # or "cpu", "mps"
)
# Chat interface
response = engine.chat([
{"role": "user", "content": "안녕하세요! K-AI에 대해 알려주세요."}
])
print(response)
kai/: Main package
core/: Shared components (Attention, MLP, MoE, Norm, RoPE)models/: Model-specific implementations
k_exaone/: K-EXAONE 236B (LG)solar_open/: Solar Open 100B (Upstage)ax_k1/: A.X-K1 519B (SKT)vaetki/: VAETKI 112B (NC)hyperclova_omni/: HyperCLOVAX SEED Omni 8B (Naver)hyperclova_think/: HyperCLOVAX SEED Think 32B (Naver)inference/: Unified inference enginecli.py: Command-line interfacecreate_random_weights.py: Utility to generate test weights1 commits
Python
94.9%
Jinja
3.6%
Just
1.5%
Note: This is a toy project designed to explore and understand the model inference structure of the
kailibrary. It focuses on architectural mapping and experimentation rather than production use.
This project provides a unified inference engine and CLI for South Korea's top Sovereign AI LLM models:
The codebase allows for checking architecture compatibility, creating random weights for testing, and running inference (text generation/chat) through a unified interface.
| Model | Developer | Params (Total/Active) | Architecture | HuggingFace |
|---|---|---|---|---|
| K-EXAONE | LG AI Research | 236B / 23B | MoE, LLLG Hybrid Attention | LGAI-EXAONE/K-EXAONE-236B-A23B |
| Solar-Open | Upstage | 102.6B / 12B | MoE, GQA | upstage/Solar-Open-100B |
| A.X-K1 | SK Telecom | 519B / 33B | MoE (1 Dense + 60 MoE) | skt/A.X-K1 |
| VAETKI | NC AI | 112.2B / 10.1B | MoE, Edge Optimized | NC-AI-consortium-VAETKI/VAETKI |
| HyperCLOVA Omni | Naver | 8B | Dense, Multimodal | naver-hyperclovax/HyperCLOVAX-SEED-Omni-8B |
| HyperCLOVA Think | Naver | 32B | Dense, VLM + Thinking | naver-hyperclovax/HyperCLOVAX-SEED-Think-32B |
| Spec | K-EXAONE | Solar-Open | A.X-K1 | VAETKI |
|---|---|---|---|---|
| Layers | 48 | 48 | 61 | 48 |
| Hidden Size | 4,096 | 4,096 | 6,144 | 4,096 |
| Attention Heads | 64 | 64 | 64 | 64 |
| KV Heads (GQA) | 8 | 8 | 8 | 8 |
| Routed Experts | 128 | 128 | 192 | 128 |
| Shared Experts | 1 | 1 | 1 | 0 |
| Top-K Selection | 8 | 8 | 8 | 8 |
| MoE Intermediate | 1,280 | 1,280 | 2,048 | 1,280 |
| Dense Intermediate | 10,240 | 10,240 | 7,168 | 10,240 |
| Vocab Size | 153,600 | 196,608 | 163,840 | 128,000 |
| Context Length | 256K | 128K | 128K | 128K |
| RoPE Theta | 1M | 1M | 1M | 1M |
| QK Norm | Yes | No | No | No |
| Attention | LLLG Hybrid | Standard GQA | Standard GQA | Standard GQA |
| Spec | HyperCLOVA Omni | HyperCLOVA Think |
|---|---|---|
| Layers | 36 | 72 |
| Hidden Size | 4,096 | 5,120 |
| Attention Heads | 32 | 40 |
| KV Heads (GQA) | 8 | 8 |
| Head Dim | 128 | 128 |
| Intermediate Size | 12,288 | 24,192 |
| Vocab Size | 200,704 | 128,256 |
| Context Length | 8K (32K with RoPE) | 128K |
| RoPE Theta | 5M | 50M |
| Modality | Text, Vision, Audio | Text, Vision |
| Special Features | Omni-modal I/O | Thinking Mode (<think>) |
# Install dependencies
pip install -r requirements.txt
# Install the package in editable mode
pip install -e .
Run the CLI to interact with a model. You can list available model types:
python cli.py --list-models
To run with a specific model path (e.g., a mini model created for testing):
python cli.py --model-path ./solar_open_mini --model-type solar_open
Since these models are massive, you can create "mini" versions with random weights to test the architecture and memory flow:
# Create Solar-Open mini model
python create_random_weights.py --model-type solar_open
# Create K-EXAONE mini model
python create_random_weights.py --model-type k_exaone
# Create A.X-K1 mini model
python create_random_weights.py --model-type ax_k1
# Create VAETKI mini model
python create_random_weights.py --model-type vaetki
# Create HyperCLOVA Omni mini model
python create_random_weights.py --model-type hyperclova_omni
# Create HyperCLOVA Think mini model
python create_random_weights.py --model-type hyperclova_think
These will create directories like ./solar_open_mini containing config.json, model.safetensors, and tokenizer files downloaded from HuggingFace.
To skip tokenizer download:
python create_random_weights.py --model-type hyperclova_omni --no-tokenizer
You can use the KAIInferenceEngine in your Python code:
from kai.inference import KAIInferenceEngine
# Load engine
engine = KAIInferenceEngine(
model_path="./k_exaone_mini",
model_type="k_exaone",
device="cuda" # or "cpu", "mps"
)
# Chat interface
response = engine.chat([
{"role": "user", "content": "안녕하세요! K-AI에 대해 알려주세요."}
])
print(response)
kai/: Main package
core/: Shared components (Attention, MLP, MoE, Norm, RoPE)models/: Model-specific implementations
k_exaone/: K-EXAONE 236B (LG)solar_open/: Solar Open 100B (Upstage)ax_k1/: A.X-K1 519B (SKT)vaetki/: VAETKI 112B (NC)hyperclova_omni/: HyperCLOVAX SEED Omni 8B (Naver)hyperclova_think/: HyperCLOVAX SEED Think 32B (Naver)inference/: Unified inference enginecli.py: Command-line interfacecreate_random_weights.py: Utility to generate test weights1 commits
Python
94.9%
Jinja
3.6%
Just
1.5%