NC-AI-consortium-VAETKI/VAETKI

Model

62

stars

115

commits

1

repos using this model

2

linked in READMEs

Aug 30, 2026

updated

causal-lm
conversational
custom_code
safetensors
text-generation
vaetki
VAETKI

README

[๐Ÿค— Models] | [๐Ÿ’ป Github] | [๐Ÿ“ Technical Report]

VAETKI ๋ชจ๋ธ ์†Œ๊ฐœ

VAETKI๋Š” NC-AI๋ฅผ ์ค‘์‹ฌ์œผ๋กœ ์ด 13๊ฐœ ๊ธฐ๊ด€์ด ์ฐธ์—ฌํ•˜๋Š” NC-AI ์ปจ์†Œ์‹œ์—„์—์„œ ๊ณต๋™ ๊ฐœ๋ฐœํ•œ ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ์ž…๋‹ˆ๋‹ค. ๋Œ€๊ทœ๋ชจ ํ˜‘๋ ฅ ์ฒด๊ณ„๋ฅผ ๊ธฐ๋ฐ˜์œผ๋กœ ๊ตฌ์ถ•๋œ VAETKI๋Š” ํšจ์œจ์„ฑ๊ณผ ํ™•์žฅ์„ฑ์„ ํ•ต์‹ฌ ๋ชฉํ‘œ๋กœ ์„ค๊ณ„๋˜์—ˆ์œผ๋ฉฐ, ์ด๋ฅผ ์œ„ํ•ด Mixture-of-Experts (MoE) ์•„ํ‚คํ…์ฒ˜๋ฅผ ์ฑ„ํƒํ•˜์˜€์Šต๋‹ˆ๋‹ค.

VAETKI๋Š” ์—ฐ๊ตฌ ๋ฐ ์‹ค์„œ๋น„์Šค ํ™˜๊ฒฝ ๋ชจ๋‘๋ฅผ ๊ณ ๋ คํ•ด ์„ค๊ณ„๋œ ๋ชจ๋ธ๋กœ์„œ, ํ–ฅํ›„ ๊ณ ๋‚œ๋„ ์ถ”๋ก  ์ค‘์‹ฌ ํƒœ์Šคํฌ, ์ „๋ฌธ ์ง€์‹ ๊ธฐ๋ฐ˜ ์‘์šฉ, ์—์ด์ „ํŠธํ˜• ํ™œ์šฉ ์‹œ๋‚˜๋ฆฌ์˜ค ๋“ฑ ๋‹ค์–‘ํ•œ ๋ถ„์•ผ์—์„œ ํ™œ์šฉ ๊ฐ€๋Šฅ์„ฑ์„ ํ™•์žฅํ•ด ๋‚˜๊ฐˆ ์ˆ˜ ์žˆ๋„๋ก ๊ฐœ๋ฐœ๋˜๊ณ  ์žˆ์œผ๋ฉฐ, ์•„๋ž˜์™€ ๊ฐ™์€ ์ฃผ์š” ํŠน์ง•์„ ๊ฐ€์ง€๊ณ  ์žˆ์Šต๋‹ˆ๋‹ค:

  • Tool agent์˜ ๊ฒฝ์šฐ non-thinking mode๋กœ ์ž‘๋™ํ•˜๋ฉฐ, ๊ทธ ์™ธ์˜ ๋ชจ๋“  ์ž‘์—…์€ thinking mode๋กœ ์ž‘๋™ํ•ฉ๋‹ˆ๋‹ค.
  • ์ง€์‹œ ์‚ฌํ•ญ์„ ์ •ํ™•ํžˆ ๋”ฐ๋ฅด๋„๋ก ์„ค๊ณ„๋œ ์ธ๊ฐ„ ์„ ํ˜ธ ์ •๋ ฌ์„ ํ†ตํ•ด ๋ณด๋‹ค ์ž์—ฐ์Šค๋Ÿฝ๊ณ  ์ผ๊ด€๋œ ๋Œ€ํ™”๋ฅผ ์ œ๊ณตํ•ฉ๋‹ˆ๋‹ค.
  • ์˜์–ด, ํ•œ๊ตญ์–ด, ์ค‘๊ตญ์–ด ๋ฐ ์ผ๋ณธ์–ด๋กœ ๊ตฌ์„ฑ๋œ ์ง€์‹œ ์ดํ–‰๊ณผ ๋ฒˆ์—ญ์„ ์ง€์›ํ•ฉ๋‹ˆ๋‹ค.

1. VAETKI Highlights

VAETKI is a large language model developed by the NC-AI consortium, a collaborative initiative led by NC-AI with participation from a total of 13 organizations. Designed with scalability and efficiency as primary goals, VAETKI adopts a Mixture-of-Experts (MoE) architecture to effectively balance performance and computational cost.

VAETKI is developed with both research and real-world applications in mind. It is intended to serve as a flexible foundation for a wide range of use cases, including advanced reasoning tasks, domain-specific knowledge applications, and agent-oriented systems, with the following key features:

  • Tool agent tasks operate in non-thinking mode, while all other tasks run in thinking mode.
  • Human preference alignment is designed to ensure accurate instruction following and more natural, consistent conversations.
  • Support of English, Korean, Chinese, and Japanese for instruction following and translation.

2. Model Overview

VAETKI-100B-A10B has the following features:

  • Type: Causal (Auto-regressive) Language Models
  • Architecture: Transformers, MoE
  • Developed by: NC-AI consortium
  • Training Stage: Pre-training & Post-training
  • Number of Parameters: 112.2B in total and 10.1B activated
  • Number of Paramaters (Non-Embedding): 111.3B
  • Number of Layers: 48
  • Number of Attention Heads: 24
  • Number of Experts: 128
  • Number of Activated Experts: 8
  • Context Length: 32k tokens
  • Vocabulary Size: 126k
  • Languages: Korean, English, Chinese, and Japanese
  • License: MIT
  • Related URLs: https://github.com/wbl-ncai/VAETKI/tree/releases/v1.0.0

For more details, please refer to our Technical Report.

3. How to Use

See the Quickstart for more details.

4. Training Details

Training Data

  • NIA-Supported Multilingual & Reasoning Datasets: To enhance multilingual processing and complex reasoning capabilities, we constructed a large-scale dataset with the support of the National Information Society Agency (NIA). During the pre-training phase, we secured 7.6 billion tokens by integrating Chinese and Japanese corpora with data specifically tailored for long-context comprehension and Chain-of-Thought (CoT) reasoning. In the subsequent post-training stage, we developed an additional 10-billion-token datasetโ€”focusing on specialized Korean studies and mathematical reasoningโ€”to maximize the model's linguistic nuance and logical performance, ultimately refining the overall maturity of the foundation model.

Training Procedure

  • Hardware
    • Platform: Naver Cloud MLX Platform
    • GPUs: NVIDIA H100 80GB HBM3 ร— 1,016
    • Interconnect: InfiniBand 400 Gb/s, 6 lanes (4 lanes were used for RDMA-based inter-node communication)
  • Software: The model architecture configuration, training loop, checkpointing, and distributed optimization logic were implemented based on Megatron-Core v0.14, with selective modifications to accommodate experimental requirements. The implementation includes internal modifications to the original frameworks for research and optimization purposes, and this model does not claim full compatibility with original upstream implementations.
  • Hyperparameters | Hyperparameters | Value | |-----------------|-------| | Learning rate | 2e-4 โ†’ 1e-4 โ†’ 8e-5 | | Batch size | 8M Tokens โ†’ 32M Tokens โ†’ 46M Tokens | | Context Length | 4096 โ†’ 4096 โ†’ 32768 |

5. Evaluation Results

We evaluate VAETKI-100B-A10B on various benchmarks and compare it with other models, as shown below.

LanguageTasksBenchmark (Metric)gpt-oss-120b (medium)VAETKI-100B-A10B
ArchitectureMoEMoE
# Total Params117B112B
# Activated Params5.1B10B
KoreanGeneralKMMLU-Pro61.958.4
GeneralCLIcK73.075.5
GeneralKoBALT46.047.5
ReasoningHRM8K83.370.6
EnglishGeneralMMLU-Pro79.171.0
ReasoningGPQA-Diamond73.153.2
ReasoningHLE (text only)8.65.9
ReasoningIFBench63.152.3
ReasoningIFEval83.686.0

6. Limitations

  • Limitations: This model may produce inaccurate or incomplete outputs, including hallucinated content, particularly for ambiguous prompts or tasks requiring high factual accuracy. It may have limitations in complex multi-step reasoning, precise mathematical computation, and strict correctness in code generation. The model does not have the ability to independently verify information.
  • (Potential) Biases: The training data may contain social or cultural biases, which can be reflected in the modelโ€™s outputs. Despite mitigation efforts, biases related to gender, ethnicity, nationality, or religion may still occur.
  • Out-of-Scope Use: This model is not designed for use in safety-critical or regulated domains, such as medical, legal, financial, or military applications. It should not be relied upon for decisions where errors could lead to harm.

7. License

This model repository is licensed under the MIT License. The use of VAETKI models is subject to the Model License. For information on third-party open-source software and data licenses used in this model, please refer to the NOTICE.md file.

8. Citation

@misc{ncai2025vaetkitechnicalreport,
      title={VAETKI Technical Report}, 
      author={NC-AI Consortium},
      year={2025},
      howpublished={\url{https://github.com/wbl-ncai/VAETKI/raw/releases/v1.0.0/VAETKI_Technical_Report.pdf}},
      note={Version 1.0.0}
}

9. Contact

If you are interested to leave a message or have any questions, please contact us at wbl.ncai.hf@gmail.com.

Contributors

nc-ai-consortium

111 commits

jungseob

2 commits

NC
ncai

1 commits

WB
wbl

1 commits

NC-AI-consortium-VAETKI/VAETKI

Model

62

stars

115

commits

1

repos using this model

2

linked in READMEs

Aug 30, 2026

updated

causal-lm
conversational
custom_code
safetensors
text-generation
vaetki
VAETKI

README

[๐Ÿค— Models] | [๐Ÿ’ป Github] | [๐Ÿ“ Technical Report]

VAETKI ๋ชจ๋ธ ์†Œ๊ฐœ

VAETKI๋Š” NC-AI๋ฅผ ์ค‘์‹ฌ์œผ๋กœ ์ด 13๊ฐœ ๊ธฐ๊ด€์ด ์ฐธ์—ฌํ•˜๋Š” NC-AI ์ปจ์†Œ์‹œ์—„์—์„œ ๊ณต๋™ ๊ฐœ๋ฐœํ•œ ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ์ž…๋‹ˆ๋‹ค. ๋Œ€๊ทœ๋ชจ ํ˜‘๋ ฅ ์ฒด๊ณ„๋ฅผ ๊ธฐ๋ฐ˜์œผ๋กœ ๊ตฌ์ถ•๋œ VAETKI๋Š” ํšจ์œจ์„ฑ๊ณผ ํ™•์žฅ์„ฑ์„ ํ•ต์‹ฌ ๋ชฉํ‘œ๋กœ ์„ค๊ณ„๋˜์—ˆ์œผ๋ฉฐ, ์ด๋ฅผ ์œ„ํ•ด Mixture-of-Experts (MoE) ์•„ํ‚คํ…์ฒ˜๋ฅผ ์ฑ„ํƒํ•˜์˜€์Šต๋‹ˆ๋‹ค.

VAETKI๋Š” ์—ฐ๊ตฌ ๋ฐ ์‹ค์„œ๋น„์Šค ํ™˜๊ฒฝ ๋ชจ๋‘๋ฅผ ๊ณ ๋ คํ•ด ์„ค๊ณ„๋œ ๋ชจ๋ธ๋กœ์„œ, ํ–ฅํ›„ ๊ณ ๋‚œ๋„ ์ถ”๋ก  ์ค‘์‹ฌ ํƒœ์Šคํฌ, ์ „๋ฌธ ์ง€์‹ ๊ธฐ๋ฐ˜ ์‘์šฉ, ์—์ด์ „ํŠธํ˜• ํ™œ์šฉ ์‹œ๋‚˜๋ฆฌ์˜ค ๋“ฑ ๋‹ค์–‘ํ•œ ๋ถ„์•ผ์—์„œ ํ™œ์šฉ ๊ฐ€๋Šฅ์„ฑ์„ ํ™•์žฅํ•ด ๋‚˜๊ฐˆ ์ˆ˜ ์žˆ๋„๋ก ๊ฐœ๋ฐœ๋˜๊ณ  ์žˆ์œผ๋ฉฐ, ์•„๋ž˜์™€ ๊ฐ™์€ ์ฃผ์š” ํŠน์ง•์„ ๊ฐ€์ง€๊ณ  ์žˆ์Šต๋‹ˆ๋‹ค:

  • Tool agent์˜ ๊ฒฝ์šฐ non-thinking mode๋กœ ์ž‘๋™ํ•˜๋ฉฐ, ๊ทธ ์™ธ์˜ ๋ชจ๋“  ์ž‘์—…์€ thinking mode๋กœ ์ž‘๋™ํ•ฉ๋‹ˆ๋‹ค.
  • ์ง€์‹œ ์‚ฌํ•ญ์„ ์ •ํ™•ํžˆ ๋”ฐ๋ฅด๋„๋ก ์„ค๊ณ„๋œ ์ธ๊ฐ„ ์„ ํ˜ธ ์ •๋ ฌ์„ ํ†ตํ•ด ๋ณด๋‹ค ์ž์—ฐ์Šค๋Ÿฝ๊ณ  ์ผ๊ด€๋œ ๋Œ€ํ™”๋ฅผ ์ œ๊ณตํ•ฉ๋‹ˆ๋‹ค.
  • ์˜์–ด, ํ•œ๊ตญ์–ด, ์ค‘๊ตญ์–ด ๋ฐ ์ผ๋ณธ์–ด๋กœ ๊ตฌ์„ฑ๋œ ์ง€์‹œ ์ดํ–‰๊ณผ ๋ฒˆ์—ญ์„ ์ง€์›ํ•ฉ๋‹ˆ๋‹ค.

1. VAETKI Highlights

VAETKI is a large language model developed by the NC-AI consortium, a collaborative initiative led by NC-AI with participation from a total of 13 organizations. Designed with scalability and efficiency as primary goals, VAETKI adopts a Mixture-of-Experts (MoE) architecture to effectively balance performance and computational cost.

VAETKI is developed with both research and real-world applications in mind. It is intended to serve as a flexible foundation for a wide range of use cases, including advanced reasoning tasks, domain-specific knowledge applications, and agent-oriented systems, with the following key features:

  • Tool agent tasks operate in non-thinking mode, while all other tasks run in thinking mode.
  • Human preference alignment is designed to ensure accurate instruction following and more natural, consistent conversations.
  • Support of English, Korean, Chinese, and Japanese for instruction following and translation.

2. Model Overview

VAETKI-100B-A10B has the following features:

  • Type: Causal (Auto-regressive) Language Models
  • Architecture: Transformers, MoE
  • Developed by: NC-AI consortium
  • Training Stage: Pre-training & Post-training
  • Number of Parameters: 112.2B in total and 10.1B activated
  • Number of Paramaters (Non-Embedding): 111.3B
  • Number of Layers: 48
  • Number of Attention Heads: 24
  • Number of Experts: 128
  • Number of Activated Experts: 8
  • Context Length: 32k tokens
  • Vocabulary Size: 126k
  • Languages: Korean, English, Chinese, and Japanese
  • License: MIT
  • Related URLs: https://github.com/wbl-ncai/VAETKI/tree/releases/v1.0.0

For more details, please refer to our Technical Report.

3. How to Use

See the Quickstart for more details.

4. Training Details

Training Data

  • NIA-Supported Multilingual & Reasoning Datasets: To enhance multilingual processing and complex reasoning capabilities, we constructed a large-scale dataset with the support of the National Information Society Agency (NIA). During the pre-training phase, we secured 7.6 billion tokens by integrating Chinese and Japanese corpora with data specifically tailored for long-context comprehension and Chain-of-Thought (CoT) reasoning. In the subsequent post-training stage, we developed an additional 10-billion-token datasetโ€”focusing on specialized Korean studies and mathematical reasoningโ€”to maximize the model's linguistic nuance and logical performance, ultimately refining the overall maturity of the foundation model.

Training Procedure

  • Hardware
    • Platform: Naver Cloud MLX Platform
    • GPUs: NVIDIA H100 80GB HBM3 ร— 1,016
    • Interconnect: InfiniBand 400 Gb/s, 6 lanes (4 lanes were used for RDMA-based inter-node communication)
  • Software: The model architecture configuration, training loop, checkpointing, and distributed optimization logic were implemented based on Megatron-Core v0.14, with selective modifications to accommodate experimental requirements. The implementation includes internal modifications to the original frameworks for research and optimization purposes, and this model does not claim full compatibility with original upstream implementations.
  • Hyperparameters | Hyperparameters | Value | |-----------------|-------| | Learning rate | 2e-4 โ†’ 1e-4 โ†’ 8e-5 | | Batch size | 8M Tokens โ†’ 32M Tokens โ†’ 46M Tokens | | Context Length | 4096 โ†’ 4096 โ†’ 32768 |

5. Evaluation Results

We evaluate VAETKI-100B-A10B on various benchmarks and compare it with other models, as shown below.

LanguageTasksBenchmark (Metric)gpt-oss-120b (medium)VAETKI-100B-A10B
ArchitectureMoEMoE
# Total Params117B112B
# Activated Params5.1B10B
KoreanGeneralKMMLU-Pro61.958.4
GeneralCLIcK73.075.5
GeneralKoBALT46.047.5
ReasoningHRM8K83.370.6
EnglishGeneralMMLU-Pro79.171.0
ReasoningGPQA-Diamond73.153.2
ReasoningHLE (text only)8.65.9
ReasoningIFBench63.152.3
ReasoningIFEval83.686.0

6. Limitations

  • Limitations: This model may produce inaccurate or incomplete outputs, including hallucinated content, particularly for ambiguous prompts or tasks requiring high factual accuracy. It may have limitations in complex multi-step reasoning, precise mathematical computation, and strict correctness in code generation. The model does not have the ability to independently verify information.
  • (Potential) Biases: The training data may contain social or cultural biases, which can be reflected in the modelโ€™s outputs. Despite mitigation efforts, biases related to gender, ethnicity, nationality, or religion may still occur.
  • Out-of-Scope Use: This model is not designed for use in safety-critical or regulated domains, such as medical, legal, financial, or military applications. It should not be relied upon for decisions where errors could lead to harm.

7. License

This model repository is licensed under the MIT License. The use of VAETKI models is subject to the Model License. For information on third-party open-source software and data licenses used in this model, please refer to the NOTICE.md file.

8. Citation

@misc{ncai2025vaetkitechnicalreport,
      title={VAETKI Technical Report}, 
      author={NC-AI Consortium},
      year={2025},
      howpublished={\url{https://github.com/wbl-ncai/VAETKI/raw/releases/v1.0.0/VAETKI_Technical_Report.pdf}},
      note={Version 1.0.0}
}

9. Contact

If you are interested to leave a message or have any questions, please contact us at wbl.ncai.hf@gmail.com.

Contributors

nc-ai-consortium

111 commits

jungseob

2 commits

NC
ncai

1 commits

WB
wbl

1 commits