| Check out the NVFP4+GPTQ weights by FuriosaAI! ➡️ link |
We introduce K-EXAONE, a large-scale multilingual language model developed by LG AI Research. Built using a Mixture-of-Experts architecture, K-EXAONE features 236 billion total parameters, with 23 billion active during inference. Performance evaluations across various benchmarks demonstrate that K-EXAONE excels in reasoning, agentic capabilities, general knowledge, multilingual understanding, and long-context processing.
For more details, please refer to the technical report and blog.

The following table shows the evaluation results of the K-EXAONE model in reasoning mode, compared to our previous model, EXAONE-4.0, and other competing models. The evaluation details can be found in the technical report.
| K-EXAONE (Reasoning) | EXAONE 4.0 (Reasoning) | GPT-OSS (Reasoning: High) | Qwen3-Thinking-2507 | DeepSeek-V3.2 (Reasoning) | ||
|---|---|---|---|---|---|---|
| Architecture | MoE | Dense | MoE | MoE | MoE | |
| Total Params | 236B | 32B | 117B | 235B | 671B | |
| Active Params | 23B | 32B | 5.1B | 22B | 37B | |
| World Knowledge | ||||||
| MMLU-Pro | 83.8 | 81.8 | 80.7 | 84.4 | 85.0 | |
| GPQA-Diamond | 79.1 | 75.4 | 80.1 | 81.1 | 82.4 | |
| Humanity's Last Exam | 13.6 | 10.6 | 14.9 | 18.2 | 25.1 | |
| Math | ||||||
| IMO-AnswerBench | 76.3 | 66.1 | 75.6 | 74.8 | 78.3 | |
| AIME 2025 | 92.8 | 85.3 | 92.5 | 92.3 | 93.1 | |
| HMMT Nov 2025 | 86.8 | 78.1 | 84.9 | 88.8 | 90.2 | |
| Coding / Agentic Coding | ||||||
| LiveCodeBench Pro 25Q2 (Medium) | 25.9 | 4.8 | 35.4 | 16.0 | 27.9 | |
| LiveCodeBench v6 | 80.7 | 66.7 | 81.9 | 74.1 | 79.4 | |
| Terminal-Bench 2.0 | 29.0 | - | 18.7 | 13.3 | 46.4 | |
| SWE-Bench Verified | 49.4 | - | 62.4 | 25.0 | 73.1 | |
| Agentic Tool Use | ||||||
| τ2-Bench (Retail) | 78.6 | 67.5 | 69.1 | 71.9 | 77.9 | |
| τ2-Bench (Airline) | 60.4 | 52.0 | 60.5 | 58.0 | 66.0 | |
| τ2-Bench (Telecom) | 73.5 | 23.7 | 60.3 | 45.6 | 85.8 | |
| BrowseComp | 31.4 | - | - | - | 51.4 | |
| Instruction Following | ||||||
| IFBench | 67.3 | 36.0 | 69.5 | 52.6 | 62.5 | |
| IFEval | 89.7 | 84.7 | 89.5 | 87.8 | 92.6 | |
| Long Context Understanding | ||||||
| AA-LCR | 53.5 | 14.0 | 50.7 | 67.0 | 65.0 | |
| OpenAI-MRCR | 52.3 | 20.1 | 29.9 | 58.6 | 57.7 | |
| Korean | ||||||
| KMMLU-Pro | 67.3 | 67.7 | 62.4 | 71.6 | 72.1 | |
| KoBALT | 61.8 | 25.4 | 54.3 | 56.1 | 62.7 | |
| CLIcK | 83.9 | 78.8 | 74.6 | 81.3 | 86.3 | |
| HRM8K | 90.9 | 89.4 | 91.6 | 92.0 | 90.6 | |
| Ko-LongBench | 86.8 | 68.0 | 82.2 | 83.2 | 87.9 | |
| Multilinguality | ||||||
| MMMLU | 85.7 | 83.2 | 83.8 | 87.3 | 88.0 | |
| WMT24++ | 90.5 | 80.8 | 93.6 | 94.7 | 90.0 | |
| Safety | ||||||
| Wild-Jailbreak | 89.9 | 62.8 | 98.2 | 85.5 | 79.1 | |
| KGC-Safety | 96.1 | 58.0 | 92.5 | 66.2 | 73.0 | |
K-EXAONE is supported by multiple libraries. Please install the required libraries as needed for your use case.
You should install transformers >= 5.1.0 for the K-EXAONE model.
To serve the K-EXAONE model on a vLLM server, you should install both Transformers and vLLM (vllm >= 0.14.0).
You should install both Transformers and SGLang to serve the K-EXAONE model on SGLang server. You can install the latest version of SGLang from source using the following commands.
git clone https://github.com/sgl-project/sglang.git
pip install -e sglang/python
To use the K-EXAONE model with llama.cpp library, you should install llama.cpp >= b7737.
You can use the K-EXAONE model with the Transformers library version 5.1.0 or later.
For tasks that require accurate results, you can run the K-EXAONE model in reasoning mode as below.
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "LGAI-EXAONE/K-EXAONE-236B-A23B"
model = AutoModelForCausalLM.from_pretrained(
model_name,
dtype="bfloat16",
device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained(model_name)
messages = [
{"role": "system", "content": "You are K-EXAONE, a large language model developed by LG AI Research in South Korea, built to serve as a helpful and reliable assistant."},
{"role": "user", "content": "Which one is bigger, 3.9 vs 3.12?"}
]
input_ids = tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_tensors="pt",
enable_thinking=True, # skippable (default: True)
)
generated_ids = model.generate(
**input_ids.to(model.device),
max_new_tokens=16384,
temperature=1.0,
top_p=0.95,
do_sample=True,
)
output_ids = generated_ids[0][input_ids['input_ids'].shape[-1]:]
print(tokenizer.decode(output_ids, skip_special_tokens=True))
For tasks where latency matters more than accuracy, you can run the K-EXAONE model in non-reasoning mode as below.
messages = [
{"role": "system", "content": "You are K-EXAONE, a large language model developed by LG AI Research in South Korea, built to serve as a helpful and reliable assistant."},
{"role": "user", "content": "Explain how wonderful you are"}
]
input_ids = tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_tensors="pt",
enable_thinking=False,
)
generated_ids = model.generate(
**input_ids.to(model.device),
max_new_tokens=1024,
temperature=1.0,
top_p=0.95,
do_sample=True,
)
output_ids = generated_ids[0][input_ids['input_ids'].shape[-1]:]
print(tokenizer.decode(output_ids, skip_special_tokens=True))
For your AI-powered agent, you can leverage K-EXAONE’s tool calling capability. The K-EXAONE model is compatible with both OpenAI and HuggingFace tool calling specifications. The example below demonstrates tool calling using HuggingFace’s docstring-to-tool-schema utility.
Please check the example file for an example of a search agent conversation using K-EXAONE.
from transformers.utils import get_json_schema
def roll_dice(max_num: int):
"""
Roll a dice with the number 1 to N. User can select the number N.
Args:
max_num: The maximum number on the dice.
"""
return random.randint(1, max_num)
tool_schema = get_json_schema(roll_dice)
tools = [tool_schema]
messages = [
{"role": "system", "content": "You are K-EXAONE, a large language model developed by LG AI Research in South Korea, built to serve as a helpful and reliable assistant."},
{"role": "user", "content": "Roll a D20 twice and sum the results."}
]
input_ids = tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_tensors="pt",
tools=tools,
)
generated_ids = model.generate(
**input_ids.to(model.device),
max_new_tokens=16384,
temperature=1.0,
top_p=0.95,
do_sample=True,
)
output_ids = generated_ids[0][input_ids['input_ids'].shape[-1]:]
print(tokenizer.decode(output_ids, skip_special_tokens=True))
You should install the llama.cpp library with the version of b7737 or after.
After you install the library, you need to prepare a model file in GGUF format as below:
# Download GGUF model weights (e.g. Q4_K_M)
hf download LGAI-EXAONE/K-EXAONE-236B-A23B-GGUF --include "*Q4_K_M*" --local-dir .
# Or convert huggingface model into GGUF format on your own
hf download LGAI-EXAONE/K-EXAONE-236B-A23B --local-dir $YOUR_MODEL_DIR
python convert_hf_to_gguf.py $YOUR_MODEL_DIR --outtype bf16 --outfile K-EXAONE-236B-A23B-BF16.gguf
# If you want to use the lower precision than BF16, you need to quantize the model
./llama-quantize K-EXAONE-236B-A23B-BF16.gguf K-EXAONE-236B-A23B-Q4_K_M.gguf Q4_K_M
You can test the model with simple chat CLI by running the command below:
./llama-cli -m K-EXAONE-236B-A23B-Q4_K_M.gguf \
-ngl 99 \
-fa on -sm row \
--temp 1.0 --top-p 0.95 --min-p 0 \
-c 131072 -n 32768 \
--no-context-shift \
--jinja
You can also launch a server by running the command below:
./llama-server -m K-EXAONE-236B-A23B-Q4_K_M.gguf \
-ngl 99 \
-fa on -sm row \
--temp 1.0 --top-p 0.95 --min-p 0 \
-c 131072 -n 32768 \
--no-context-shift \
--jinja \
--host 0.0.0.0 --port 8080
When the server is ready, you can test the model using the chat-style UI at http://localhost:8080, and access the OpenAI-compatible API at http://localhost:8080/v1.
Ollama and LM-Studio are powered by llama.cpp, so they should be updated once llama.cpp officially supports K-EXAONE. We will update this section once each library supports K-EXAONE.
TensorRT-LLM provides official support for the K-EXAONE model. Please refer to the EXAONE Documentation in the TensorRT-LLM repository for more information.
We support the K-EXAONE model on vLLM. You need to install vllm >= 0.14.0.
Practically, you can serve the model with a 256K context length using tensor parallel on 4 H200 GPUs.
After you install the vLLM library with an EXAONE-MoE implementation, you can run the vLLM server by following command:
vllm serve LGAI-EXAONE/K-EXAONE-236B-A23B \
--reasoning-parser deepseek_v3 \
--tensor-parallel-size 4 \
--enable-auto-tool-choice \
--tool-call-parser hermes
An OpenAI-compatible API server will be available at http://localhost:8000/v1.
You can test the vLLM server by sending a chat completion request as below:
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "LGAI-EXAONE/K-EXAONE-236B-A23B",
"messages": [
{"role": "user", "content": "How many r'\''s in \"strawberry\"?"}
],
"max_tokens": 16384,
"temperature": 1.0,
"top_p": 0.95,
"chat_template_kwargs": {"enable_thinking": true}
}'
If you are interested in using MTP weights for speculative decoding, add according options as below.
vllm serve LGAI-EXAONE/K-EXAONE-236B-A23B \
--reasoning-parser deepseek_v3 \
--tensor-parallel-size 4 \
--enable-auto-tool-choice \
--tool-call-parser hermes \
--no-enable-prefix-caching \
--speculative_config '{
"method": "mtp",
"num_speculative_tokens": 2
}'
We support the K-EXAONE model on SGLang. You need to install the latest version of the SGLang library from source. Please check the requirements section. Practically, you can serve the model with a 256K context length using tensor parallel on 4 H200 GPUs.
python -m sglang.launch_server \
--model LGAI-EXAONE/K-EXAONE-236B-A23B \
--tp-size 4 \
--reasoning-parser qwen3
A SGLang server will be available at http://localhost:30000.
[!NOTE] Currently, using the OpenAI-compatible server is incompatible with the
transformers>=5.0.0rc0, so you need to use SGLang native API for now. For native API, please refer to the official documentation.Once the issue is resolved, we will update this section accordingly.
You can test the SGLang server by sending a request as below:
from transformers import AutoTokenizer
import requests
model_name = "LGAI-EXAONE/K-EXAONE-236B-A23B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
messages = [
{"role": "user", "content": "How many r'\''s in \"strawberry\"?"}
]
input_text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
return_tensors="pt",
)
response = requests.post(
f"http://localhost:30000/generate",
json={
"text": input_text,
"sampling_params": {
"temperature": 1.0,
"top_p": 0.95,
"max_new_tokens": 16384,
},
},
)
print(response.json()['text'])
If you are interested in in using MTP weights for speculative decoding, add according options as below.
python -m sglang.launch_server \
--model LGAI-EXAONE/K-EXAONE-236B-A23B \
--tp-size 4 \
--reasoning-parser qwen3 \
--speculative-algorithm EAGLE \
--speculative-num-steps 3 \
--speculative-eagle-topk 1 \
--speculative-num-draft-tokens 4
[!IMPORTANT] To achieve the expected performance, we recommend using the following configurations:
- We strongly recommend to use
temperature=1.0,top_p=0.95,presence_penalty=0.0for best performance.- Different from EXAONE-4.0, K-EXAONE uses
enable_thinking=Trueas default. Thus, you need to setenable_thinking=Falsewhen you want to use non-reasoning mode.
The K-EXAONE language model has certain limitations and may occasionally generate inappropriate responses. The language model generates responses based on the output probability of tokens, and it is determined during learning from training data. While we have made every effort to exclude personal, harmful, and biased information from the training data, some problematic content may still be included, potentially leading to undesirable responses. Please note that the text generated by K-EXAONE language model does not reflect the views of LG AI Research.
LG AI Research strives to reduce potential risks that may arise from K-EXAONE language models. Users are not allowed to engage in any malicious activities (e.g., keying in illegal information) that may induce the creation of inappropriate outputs violating LG AI's ethical principles when using K-EXAONE language models.
The model is licensed under K-EXAONE AI Model License Agreement
@article{k-exaone,
title={K-EXAONE Technical Report},
author={{LG AI Research}},
journal={arXiv preprint arXiv:2601.01739},
year={2025}
}
LG AI Research Technical Support: contact_us@lgresearch.ai
14 commits
| Check out the NVFP4+GPTQ weights by FuriosaAI! ➡️ link |
We introduce K-EXAONE, a large-scale multilingual language model developed by LG AI Research. Built using a Mixture-of-Experts architecture, K-EXAONE features 236 billion total parameters, with 23 billion active during inference. Performance evaluations across various benchmarks demonstrate that K-EXAONE excels in reasoning, agentic capabilities, general knowledge, multilingual understanding, and long-context processing.
For more details, please refer to the technical report and blog.

The following table shows the evaluation results of the K-EXAONE model in reasoning mode, compared to our previous model, EXAONE-4.0, and other competing models. The evaluation details can be found in the technical report.
| K-EXAONE (Reasoning) | EXAONE 4.0 (Reasoning) | GPT-OSS (Reasoning: High) | Qwen3-Thinking-2507 | DeepSeek-V3.2 (Reasoning) | ||
|---|---|---|---|---|---|---|
| Architecture | MoE | Dense | MoE | MoE | MoE | |
| Total Params | 236B | 32B | 117B | 235B | 671B | |
| Active Params | 23B | 32B | 5.1B | 22B | 37B | |
| World Knowledge | ||||||
| MMLU-Pro | 83.8 | 81.8 | 80.7 | 84.4 | 85.0 | |
| GPQA-Diamond | 79.1 | 75.4 | 80.1 | 81.1 | 82.4 | |
| Humanity's Last Exam | 13.6 | 10.6 | 14.9 | 18.2 | 25.1 | |
| Math | ||||||
| IMO-AnswerBench | 76.3 | 66.1 | 75.6 | 74.8 | 78.3 | |
| AIME 2025 | 92.8 | 85.3 | 92.5 | 92.3 | 93.1 | |
| HMMT Nov 2025 | 86.8 | 78.1 | 84.9 | 88.8 | 90.2 | |
| Coding / Agentic Coding | ||||||
| LiveCodeBench Pro 25Q2 (Medium) | 25.9 | 4.8 | 35.4 | 16.0 | 27.9 | |
| LiveCodeBench v6 | 80.7 | 66.7 | 81.9 | 74.1 | 79.4 | |
| Terminal-Bench 2.0 | 29.0 | - | 18.7 | 13.3 | 46.4 | |
| SWE-Bench Verified | 49.4 | - | 62.4 | 25.0 | 73.1 | |
| Agentic Tool Use | ||||||
| τ2-Bench (Retail) | 78.6 | 67.5 | 69.1 | 71.9 | 77.9 | |
| τ2-Bench (Airline) | 60.4 | 52.0 | 60.5 | 58.0 | 66.0 | |
| τ2-Bench (Telecom) | 73.5 | 23.7 | 60.3 | 45.6 | 85.8 | |
| BrowseComp | 31.4 | - | - | - | 51.4 | |
| Instruction Following | ||||||
| IFBench | 67.3 | 36.0 | 69.5 | 52.6 | 62.5 | |
| IFEval | 89.7 | 84.7 | 89.5 | 87.8 | 92.6 | |
| Long Context Understanding | ||||||
| AA-LCR | 53.5 | 14.0 | 50.7 | 67.0 | 65.0 | |
| OpenAI-MRCR | 52.3 | 20.1 | 29.9 | 58.6 | 57.7 | |
| Korean | ||||||
| KMMLU-Pro | 67.3 | 67.7 | 62.4 | 71.6 | 72.1 | |
| KoBALT | 61.8 | 25.4 | 54.3 | 56.1 | 62.7 | |
| CLIcK | 83.9 | 78.8 | 74.6 | 81.3 | 86.3 | |
| HRM8K | 90.9 | 89.4 | 91.6 | 92.0 | 90.6 | |
| Ko-LongBench | 86.8 | 68.0 | 82.2 | 83.2 | 87.9 | |
| Multilinguality | ||||||
| MMMLU | 85.7 | 83.2 | 83.8 | 87.3 | 88.0 | |
| WMT24++ | 90.5 | 80.8 | 93.6 | 94.7 | 90.0 | |
| Safety | ||||||
| Wild-Jailbreak | 89.9 | 62.8 | 98.2 | 85.5 | 79.1 | |
| KGC-Safety | 96.1 | 58.0 | 92.5 | 66.2 | 73.0 | |
K-EXAONE is supported by multiple libraries. Please install the required libraries as needed for your use case.
You should install transformers >= 5.1.0 for the K-EXAONE model.
To serve the K-EXAONE model on a vLLM server, you should install both Transformers and vLLM (vllm >= 0.14.0).
You should install both Transformers and SGLang to serve the K-EXAONE model on SGLang server. You can install the latest version of SGLang from source using the following commands.
git clone https://github.com/sgl-project/sglang.git
pip install -e sglang/python
To use the K-EXAONE model with llama.cpp library, you should install llama.cpp >= b7737.
You can use the K-EXAONE model with the Transformers library version 5.1.0 or later.
For tasks that require accurate results, you can run the K-EXAONE model in reasoning mode as below.
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "LGAI-EXAONE/K-EXAONE-236B-A23B"
model = AutoModelForCausalLM.from_pretrained(
model_name,
dtype="bfloat16",
device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained(model_name)
messages = [
{"role": "system", "content": "You are K-EXAONE, a large language model developed by LG AI Research in South Korea, built to serve as a helpful and reliable assistant."},
{"role": "user", "content": "Which one is bigger, 3.9 vs 3.12?"}
]
input_ids = tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_tensors="pt",
enable_thinking=True, # skippable (default: True)
)
generated_ids = model.generate(
**input_ids.to(model.device),
max_new_tokens=16384,
temperature=1.0,
top_p=0.95,
do_sample=True,
)
output_ids = generated_ids[0][input_ids['input_ids'].shape[-1]:]
print(tokenizer.decode(output_ids, skip_special_tokens=True))
For tasks where latency matters more than accuracy, you can run the K-EXAONE model in non-reasoning mode as below.
messages = [
{"role": "system", "content": "You are K-EXAONE, a large language model developed by LG AI Research in South Korea, built to serve as a helpful and reliable assistant."},
{"role": "user", "content": "Explain how wonderful you are"}
]
input_ids = tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_tensors="pt",
enable_thinking=False,
)
generated_ids = model.generate(
**input_ids.to(model.device),
max_new_tokens=1024,
temperature=1.0,
top_p=0.95,
do_sample=True,
)
output_ids = generated_ids[0][input_ids['input_ids'].shape[-1]:]
print(tokenizer.decode(output_ids, skip_special_tokens=True))
For your AI-powered agent, you can leverage K-EXAONE’s tool calling capability. The K-EXAONE model is compatible with both OpenAI and HuggingFace tool calling specifications. The example below demonstrates tool calling using HuggingFace’s docstring-to-tool-schema utility.
Please check the example file for an example of a search agent conversation using K-EXAONE.
from transformers.utils import get_json_schema
def roll_dice(max_num: int):
"""
Roll a dice with the number 1 to N. User can select the number N.
Args:
max_num: The maximum number on the dice.
"""
return random.randint(1, max_num)
tool_schema = get_json_schema(roll_dice)
tools = [tool_schema]
messages = [
{"role": "system", "content": "You are K-EXAONE, a large language model developed by LG AI Research in South Korea, built to serve as a helpful and reliable assistant."},
{"role": "user", "content": "Roll a D20 twice and sum the results."}
]
input_ids = tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_tensors="pt",
tools=tools,
)
generated_ids = model.generate(
**input_ids.to(model.device),
max_new_tokens=16384,
temperature=1.0,
top_p=0.95,
do_sample=True,
)
output_ids = generated_ids[0][input_ids['input_ids'].shape[-1]:]
print(tokenizer.decode(output_ids, skip_special_tokens=True))
You should install the llama.cpp library with the version of b7737 or after.
After you install the library, you need to prepare a model file in GGUF format as below:
# Download GGUF model weights (e.g. Q4_K_M)
hf download LGAI-EXAONE/K-EXAONE-236B-A23B-GGUF --include "*Q4_K_M*" --local-dir .
# Or convert huggingface model into GGUF format on your own
hf download LGAI-EXAONE/K-EXAONE-236B-A23B --local-dir $YOUR_MODEL_DIR
python convert_hf_to_gguf.py $YOUR_MODEL_DIR --outtype bf16 --outfile K-EXAONE-236B-A23B-BF16.gguf
# If you want to use the lower precision than BF16, you need to quantize the model
./llama-quantize K-EXAONE-236B-A23B-BF16.gguf K-EXAONE-236B-A23B-Q4_K_M.gguf Q4_K_M
You can test the model with simple chat CLI by running the command below:
./llama-cli -m K-EXAONE-236B-A23B-Q4_K_M.gguf \
-ngl 99 \
-fa on -sm row \
--temp 1.0 --top-p 0.95 --min-p 0 \
-c 131072 -n 32768 \
--no-context-shift \
--jinja
You can also launch a server by running the command below:
./llama-server -m K-EXAONE-236B-A23B-Q4_K_M.gguf \
-ngl 99 \
-fa on -sm row \
--temp 1.0 --top-p 0.95 --min-p 0 \
-c 131072 -n 32768 \
--no-context-shift \
--jinja \
--host 0.0.0.0 --port 8080
When the server is ready, you can test the model using the chat-style UI at http://localhost:8080, and access the OpenAI-compatible API at http://localhost:8080/v1.
Ollama and LM-Studio are powered by llama.cpp, so they should be updated once llama.cpp officially supports K-EXAONE. We will update this section once each library supports K-EXAONE.
TensorRT-LLM provides official support for the K-EXAONE model. Please refer to the EXAONE Documentation in the TensorRT-LLM repository for more information.
We support the K-EXAONE model on vLLM. You need to install vllm >= 0.14.0.
Practically, you can serve the model with a 256K context length using tensor parallel on 4 H200 GPUs.
After you install the vLLM library with an EXAONE-MoE implementation, you can run the vLLM server by following command:
vllm serve LGAI-EXAONE/K-EXAONE-236B-A23B \
--reasoning-parser deepseek_v3 \
--tensor-parallel-size 4 \
--enable-auto-tool-choice \
--tool-call-parser hermes
An OpenAI-compatible API server will be available at http://localhost:8000/v1.
You can test the vLLM server by sending a chat completion request as below:
curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "LGAI-EXAONE/K-EXAONE-236B-A23B",
"messages": [
{"role": "user", "content": "How many r'\''s in \"strawberry\"?"}
],
"max_tokens": 16384,
"temperature": 1.0,
"top_p": 0.95,
"chat_template_kwargs": {"enable_thinking": true}
}'
If you are interested in using MTP weights for speculative decoding, add according options as below.
vllm serve LGAI-EXAONE/K-EXAONE-236B-A23B \
--reasoning-parser deepseek_v3 \
--tensor-parallel-size 4 \
--enable-auto-tool-choice \
--tool-call-parser hermes \
--no-enable-prefix-caching \
--speculative_config '{
"method": "mtp",
"num_speculative_tokens": 2
}'
We support the K-EXAONE model on SGLang. You need to install the latest version of the SGLang library from source. Please check the requirements section. Practically, you can serve the model with a 256K context length using tensor parallel on 4 H200 GPUs.
python -m sglang.launch_server \
--model LGAI-EXAONE/K-EXAONE-236B-A23B \
--tp-size 4 \
--reasoning-parser qwen3
A SGLang server will be available at http://localhost:30000.
[!NOTE] Currently, using the OpenAI-compatible server is incompatible with the
transformers>=5.0.0rc0, so you need to use SGLang native API for now. For native API, please refer to the official documentation.Once the issue is resolved, we will update this section accordingly.
You can test the SGLang server by sending a request as below:
from transformers import AutoTokenizer
import requests
model_name = "LGAI-EXAONE/K-EXAONE-236B-A23B"
tokenizer = AutoTokenizer.from_pretrained(model_name)
messages = [
{"role": "user", "content": "How many r'\''s in \"strawberry\"?"}
]
input_text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
return_tensors="pt",
)
response = requests.post(
f"http://localhost:30000/generate",
json={
"text": input_text,
"sampling_params": {
"temperature": 1.0,
"top_p": 0.95,
"max_new_tokens": 16384,
},
},
)
print(response.json()['text'])
If you are interested in in using MTP weights for speculative decoding, add according options as below.
python -m sglang.launch_server \
--model LGAI-EXAONE/K-EXAONE-236B-A23B \
--tp-size 4 \
--reasoning-parser qwen3 \
--speculative-algorithm EAGLE \
--speculative-num-steps 3 \
--speculative-eagle-topk 1 \
--speculative-num-draft-tokens 4
[!IMPORTANT] To achieve the expected performance, we recommend using the following configurations:
- We strongly recommend to use
temperature=1.0,top_p=0.95,presence_penalty=0.0for best performance.- Different from EXAONE-4.0, K-EXAONE uses
enable_thinking=Trueas default. Thus, you need to setenable_thinking=Falsewhen you want to use non-reasoning mode.
The K-EXAONE language model has certain limitations and may occasionally generate inappropriate responses. The language model generates responses based on the output probability of tokens, and it is determined during learning from training data. While we have made every effort to exclude personal, harmful, and biased information from the training data, some problematic content may still be included, potentially leading to undesirable responses. Please note that the text generated by K-EXAONE language model does not reflect the views of LG AI Research.
LG AI Research strives to reduce potential risks that may arise from K-EXAONE language models. Users are not allowed to engage in any malicious activities (e.g., keying in illegal information) that may induce the creation of inappropriate outputs violating LG AI's ethical principles when using K-EXAONE language models.
The model is licensed under K-EXAONE AI Model License Agreement
@article{k-exaone,
title={K-EXAONE Technical Report},
author={{LG AI Research}},
journal={arXiv preprint arXiv:2601.01739},
year={2025}
}
LG AI Research Technical Support: contact_us@lgresearch.ai
14 commits