MoonshotAI/Kimi-Vendor-Verifier

Kimi-Vendor-Verifier

Python

163

19 commits

updated Sep 17, 2026

See the code

README

Kimi Vendor Verifier

English | 中文

K3 Evaluation Results

Listed in the order of submission time.

model: kimi-k3

thinking effort: max

ProviderOCRBenchMMMU Pro VisionBEAM (1M)DeepSWE
Moonshot0.890.820.310.675
Fireworks0.890.820.30370.664
Baseten0.8890.8040.32190.693
Together0.8970.8200.31600.678
DigitalOcean0.890.816TBDTBD
Inferact (vLLM ref.)0.8910.8180.31880.695
Nebius0.8780.8140.29130.673
Modal0.8870.8170.3220.658
RadixArk (SGlang ref.)0.8950.8200.31070.667

Overview

TypeTaskDescription
Inspect-ai benchmarkOCRBenchOCR text recognition benchmark
Inspect-ai benchmarkMMMU Pro VisionMultimodal understanding benchmark (visual QA)
Inspect-ai benchmarkAIME 2025Mathematical reasoning benchmark; no longer used for K3
Standalone scriptBEAM (1M)Long-term memory benchmark over 1M-token conversations, see beam/
Pytest verifiertests/params/API parameter constraint pre-flight validation
Pytest verifiertests/tool_call_json_schema/Tool-call argument validation for walle-valid MFJS schemas
Pytest verifiertests/k3_features/K3 feature contract validation, including dynamic tools, response_format, tool_choice, and thinking effort
Pytest verifiertests/prompt_tokens/Verifies vendor-reported usage.prompt_tokens against expected constants
Agent benchmarkDeepSWEMulti-step tool-use and coding-agent evaluation, run on the Pier platform

Environment Setup

Install Dependencies

uv sync && uv pip install -e .

Configure Environment Variables

export KIMI_API_KEY="your-api-key"
export KIMI_BASE_URL="your-base-url"

Or copy .env.example to .env and fill in the configuration.

Pre-flight Validation

Before running benchmarks, complete the API parameter, tool-call schema, K3 feature, and prompt-token pre-flight checks.

Parameter Constraint Validation: tests/params/

Validate that the API correctly constrains immutable parameters such as temperature, top_p, presence_penalty, frequency_penalty, and n.

# Kimi official API
uv run pytest tests/params --smoke-model kimi/your-model-id --think-mode kimi -v

# Open-source deployment (vLLM/SGLang/KTransformers)
uv run pytest tests/params --smoke-model your-model-id --think-mode opensource -v

Run formal benchmarks only after all tests pass.

Tool Call JSON Schema Validation: tests/tool_call_json_schema/

Validate that the vendor can use walle-valid MFJS schemas as tool-call parameters and return tool_calls[].function.arguments that conform to the schema.

uv run pytest -n 4 tests/tool_call_json_schema \
    --base-url "${KIMI_BASE_URL}" \
    --api-key "${KIMI_API_KEY}" \
    --smoke-model "${MODEL_NAME}" \
    --think-mode "$THINK_MODE" \
    --thinking \
    --reruns 3 \
    --reruns-delay 2 \
    --tool-json-report=tool-call-schema-report.json \
    -ra -v

This check runs walle valid.jsonl cases, sends each schema as tools[].function.parameters, forces a tool call, and uses jsonschema to validate the returned function.arguments locally. Each selected case is run once with stream=false and once with stream=true; streaming chunks are reassembled before validation. In GitLab CI, this check runs as the verify_tool_call_json_schema job; the same retry options are applied. In addition to the JUnit XML report, a JSON report (tool-call-schema-report.json) is produced containing the selected cases, per-case outcomes, and the same summary block (total / by_status / by_selection_reason / by_mode) that the original standalone script printed at the end of its log.

Arguments

ArgumentMeaningDefault
--smoke-modelVendor model nameMODEL_NAME env var
--base-urlAPI base URLKIMI_BASE_URL env var, otherwise Moonshot API
--api-keyAPI keyKIMI_API_KEY env var
--case-dirwalle validator_cases directorytestdata/walle_validator_cases/validator_cases
--selectionCase selection mode: all / explicit / objectall
--max-casesMaximum number of selected cases to runNo limit
--thinkingEnable thinking mode for the tool-call requestOff
--think-modeThinking parameter format: kimi, opensource, or noneTHINK_MODE env var, otherwise kimi
--max-tokensMaximum output tokens for each tool-call response2048
--tool-json-reportPath to the additional JSON report artifacttool-call-schema-report.json

K3 Feature Validation: tests/k3_features/

Validate K3 API features such as dynamic tools, response_format, tool_choice, and thinking effort.

uv run pytest tests/k3_features \
    --base-url "${KIMI_BASE_URL}" \
    --api-key "${KIMI_API_KEY}" \
    --smoke-model "${MODEL_NAME}" \
    -ra -v
ArgumentMeaningDefault
--smoke-modelVendor model nameMODEL_NAME env var
--base-urlAPI base URLKIMI_BASE_URL env var, otherwise Moonshot API
--api-keyAPI keyKIMI_API_KEY env var

The CI job verify_k3_features runs the pytest suite and reports results from the test log.

Prompt Token Validation: tests/prompt_tokens/

Validate that vendor-reported usage.prompt_tokens is accurate. This check sends the text cases in testdata/prompt_token_cases/cases.jsonl and vision cases in testdata/prompt_token_cases/vision_cases.jsonl with stream=true, then compares the stream-reported usage.prompt_tokens against the expected constant stored in each case. See testdata/prompt_token_cases/README.md for the case format.

uv run pytest tests/prompt_tokens \
    --base-url "${KIMI_BASE_URL}" \
    --api-key "${KIMI_API_KEY}" \
    --smoke-model "${MODEL_NAME}" \
    -ra -v

Why These Tests Are Not in Inspect-ai

tests/k3_features/, tests/prompt_tokens/, and tests/tool_call_json_schema/ validate raw chat-completions API behavior, including non-standard messages[].tools dynamic-tool declarations, malformed requests that should return HTTP 400, the location of usage fields in streaming chunks, and raw tools[].function.parameters JSON schema handling. inspect-ai is better suited for scored benchmark evaluations and may normalize or reject these raw payloads before they reach the vendor endpoint, so these tests remain standalone pytest suites.

K2.6 and Earlier

BenchmarkModeTemperatureTopPMax TokensEpochs
OCRBenchNon-Thinking0.60.95163841
OCRBenchThinking1.00.95163841
MMMUNon-Thinking0.60.95655361
MMMUThinking1.00.95655361
AIME 2025Non-Thinking0.60.959830432
AIME 2025Thinking1.00.959830432

K2.7

BenchmarkMode (thinking, preserve thinking)TemperatureTopPMax TokensEpochs
OCRBenchThinking, Keep all1.00.95163841
MMMUThinking, Keep all1.00.95655361
AIME 2025Thinking, Keep all1.00.959830432

K3

K3 no longer uses AIME 2025 and adds the reasoning_effort parameter.

BenchmarkMode (thinking, preserve thinking, effort)TemperatureTopPMax TokensEpochs
OCRBenchThinking, Keep all, low1.00.95163841
OCRBenchThinking, Keep all, high1.00.95163841
OCRBenchThinking, Keep all, max1.00.95163841
MMMUThinking, Keep all, low1.00.95983041
MMMUThinking, Keep all, high1.00.95983041
MMMUThinking, Keep all, max1.00.95983041

Running Inspect-ai Benchmarks

OCRBench

Non-Thinking

uv run python eval.py ocrbench --model kimi/your-model-id \
    --think-mode kimi --max-tokens 16384 --stream

Thinking

uv run python eval.py ocrbench --model kimi/your-model-id \
    --thinking --think-mode kimi --max-tokens 16384 --stream

Thinking + reasoning effort (K3)

uv run python eval.py ocrbench \
    --model "${THINK_MODE}/${MODEL_NAME}" \
    --max-tokens 16384 \
    --thinking \
    --think-mode "$THINK_MODE" \
    --stream \
    --max-connections 50 \
    --temperature 1.0 \
    --top-p 0.95 \
    --thinking-effort high

MMMU Pro Vision

Non-Thinking

uv run python eval.py mmmu --model kimi/your-model-id \
    --think-mode kimi --max-tokens 65536 --stream

Thinking

uv run python eval.py mmmu --model kimi/your-model-id \
    --thinking --think-mode kimi --max-tokens 65536 --stream

Thinking + reasoning effort (K3)

uv run python eval.py mmmu \
    --model "$MODEL" \
    --max-tokens 98304 \
    --thinking \
    --think-mode "$THINK_MODE" \
    --stream \
    --max-connections 50 \
    --temperature 1.0 \
    --top-p 0.95 \
    --thinking-effort high

AIME 2025

Non-Thinking

uv run python eval.py aime2025 --model kimi/your-model-id \
    --think-mode kimi --max-tokens 98304 --stream

Thinking

uv run python eval.py aime2025 --model kimi/your-model-id \
    --thinking --think-mode kimi --max-tokens 98304 --stream

Run OCRBench first as a quick deployment sanity check, then run the full MMMU and other benchmark evaluations.

Agent Benchmark (DeepSWE)

To cover that capability, we also add an agent benchmark based on the open-source DeepSWE benchmark. DeepSWE contains 113 coding-agent tasks and is used here through the open-source Pier platform. The agent runtime uses Kimi-code 0.23.6 or later.

Before running the benchmark on Pier, register kimi-code on the Pier agent first. See the Kimi-code 0.23.6 setup reference: https://github.com/MoonshotAI/kimi-code/tree/%40moonshot-ai/kimi-code%400.23.6

Inspect-ai Parameters

Available Parameters

ParameterDescriptionDefault
benchmarkEvaluation task: ocrbench, mmmu, aime2025ocrbench
--modelModel identifier, for example kimi/your-model-idRequired
--max-tokensMaximum output tokens, see recommended benchmark parametersRequired
--thinkingEnable thinking modeOff
--think-modeThinking parameter format: none, kimi, or opensourcenone
--temperatureSampling temperatureOptional; determined by server or model defaults if omitted
--top-pTop-p samplingOptional; determined by server or model defaults if omitted
--streamEnable streaming, recommended for long reasoning requestsOff
--max-connectionsMaximum concurrent connectionsPer benchmark
--epochsNumber of sampling epochsPer benchmark
--client-timeoutHTTP timeout in seconds86400
--thinking-effortK3 reasoning effort parameter: low, high, maxNone

Thinking Mode Parameters

Model TypeParameter Combinationextra_body Sent
Kimi official + thinking off--think-mode kimi{"thinking": {"type": "disabled"}}
Kimi official + thinking on--thinking --think-mode kimi{"thinking": {"type": "enabled"}}
Kimi official + reasoning effort low--thinking --think-mode kimi --thinking-effort low{"thinking": {"type": "enabled", "keep": "all", "effort": "low"}}
Kimi official + reasoning effort high--thinking --think-mode kimi --thinking-effort high{"thinking": {"type": "enabled", "keep": "all", "effort": "high"}}
Kimi official + reasoning effort max--thinking --think-mode kimi --thinking-effort max{"thinking": {"type": "enabled", "keep": "all", "effort": "max"}}
Open-source framework + thinking off--think-mode opensource{"chat_template_kwargs": {"thinking": false}}
Open-source framework + thinking on--thinking --think-mode opensource{"chat_template_kwargs": {"thinking": true}}
Open-source framework + reasoning effort low--thinking --think-mode opensource --thinking-effort low{"chat_template_kwargs": {"thinking": true, "preserve_thinking": true, "thinking_effort": "low"}}
Open-source framework + reasoning effort high--thinking --think-mode opensource --thinking-effort high{"chat_template_kwargs": {"thinking": true, "preserve_thinking": true, "thinking_effort": "high"}}
Open-source framework + reasoning effort max--thinking --think-mode opensource --thinking-effort max{"chat_template_kwargs": {"thinking": true, "preserve_thinking": true, "thinking_effort": "max"}}

Viewing Results

# View logs with inspect view
uv run inspect view

# Logs are saved in the logs/ directory

Resuming an Interrupted Benchmark

uv run inspect eval-retry logs/<log-file>.eval

Notes

AIME 2025

AIME evaluation produces many output tokens; please note the following:

  1. Timeout settings: the client default is --client-timeout 86400 (24 hours), and the server, gateway, or proxy timeout should also be long enough.
  2. Streaming: strongly recommended with --stream; non-streaming requests are more prone to timeout in thinking mode.
  3. Concurrency control: if you see many 429s or RemoteProtocolErrors, lower --max-connections.
  4. Quick validation: first run all samples with --epochs 1, then run the full evaluation after configuration is confirmed.
# Step 1: Quick validation (30 samples x 1 epoch)
uv run python eval.py aime2025 --model kimi/your-model-id \
    --thinking --think-mode kimi --max-tokens 98304 --stream --epochs 1

# Step 2: Full evaluation (30 samples x 32 epochs)
uv run python eval.py aime2025 --model kimi/your-model-id \
    --thinking --think-mode kimi --max-tokens 98304 --stream

Automatic Retry Mechanism

The following network-related errors are automatically retried with exponential backoff (1-60 seconds), no manual configuration needed:

Error TypeDescription
RateLimitError / 429Server-side rate limiting
APIConnectionErrorConnection failure
ReadError / RemoteProtocolErrorNetwork read error

Non-network errors, such as model output format issues, are not retried and are logged directly for later analysis.

Project Structure

├── eval.py              # Main evaluation CLI entrypoint
├── tests/params/        # Pre-flight parameter validation
├── tests/tool_call_json_schema/ # walle tool-call schema validation
├── tests/prompt_tokens/ # usage.prompt_tokens accuracy validation
├── tests/k3_features/   # K3 feature contract validation
├── kimi_model.py        # Kimi Model API implementation
├── aime2025.py          # AIME 2025 evaluation task
├── mmmu_pro_vision.py   # MMMU Pro Vision evaluation task
├── ocr_bench.py         # OCRBench evaluation task
├── testdata/            # JSON schema and prompt-token test data
├── beam/                # BEAM 1M context data and scripts, see the directory README
├── logs/                # Evaluation logs
└── pyproject.toml       # Project configuration

Contact

If you have any questions or suggestions, please contact contact-kvv@kimi.com.

License

This project is licensed under the MIT License. See LICENSE for details.

Contributors

Zwysilence

11 commits

dijin-dev

5 commits

xiaochen-dev

3 commits

MoonshotAI/Kimi-Vendor-Verifier

Kimi-Vendor-Verifier

Python

163

19 commits

updated Sep 17, 2026

See the code

README

Kimi Vendor Verifier

English | 中文

K3 Evaluation Results

Listed in the order of submission time.

model: kimi-k3

thinking effort: max

ProviderOCRBenchMMMU Pro VisionBEAM (1M)DeepSWE
Moonshot0.890.820.310.675
Fireworks0.890.820.30370.664
Baseten0.8890.8040.32190.693
Together0.8970.8200.31600.678
DigitalOcean0.890.816TBDTBD
Inferact (vLLM ref.)0.8910.8180.31880.695
Nebius0.8780.8140.29130.673
Modal0.8870.8170.3220.658
RadixArk (SGlang ref.)0.8950.8200.31070.667

Overview

TypeTaskDescription
Inspect-ai benchmarkOCRBenchOCR text recognition benchmark
Inspect-ai benchmarkMMMU Pro VisionMultimodal understanding benchmark (visual QA)
Inspect-ai benchmarkAIME 2025Mathematical reasoning benchmark; no longer used for K3
Standalone scriptBEAM (1M)Long-term memory benchmark over 1M-token conversations, see beam/
Pytest verifiertests/params/API parameter constraint pre-flight validation
Pytest verifiertests/tool_call_json_schema/Tool-call argument validation for walle-valid MFJS schemas
Pytest verifiertests/k3_features/K3 feature contract validation, including dynamic tools, response_format, tool_choice, and thinking effort
Pytest verifiertests/prompt_tokens/Verifies vendor-reported usage.prompt_tokens against expected constants
Agent benchmarkDeepSWEMulti-step tool-use and coding-agent evaluation, run on the Pier platform

Environment Setup

Install Dependencies

uv sync && uv pip install -e .

Configure Environment Variables

export KIMI_API_KEY="your-api-key"
export KIMI_BASE_URL="your-base-url"

Or copy .env.example to .env and fill in the configuration.

Pre-flight Validation

Before running benchmarks, complete the API parameter, tool-call schema, K3 feature, and prompt-token pre-flight checks.

Parameter Constraint Validation: tests/params/

Validate that the API correctly constrains immutable parameters such as temperature, top_p, presence_penalty, frequency_penalty, and n.

# Kimi official API
uv run pytest tests/params --smoke-model kimi/your-model-id --think-mode kimi -v

# Open-source deployment (vLLM/SGLang/KTransformers)
uv run pytest tests/params --smoke-model your-model-id --think-mode opensource -v

Run formal benchmarks only after all tests pass.

Tool Call JSON Schema Validation: tests/tool_call_json_schema/

Validate that the vendor can use walle-valid MFJS schemas as tool-call parameters and return tool_calls[].function.arguments that conform to the schema.

uv run pytest -n 4 tests/tool_call_json_schema \
    --base-url "${KIMI_BASE_URL}" \
    --api-key "${KIMI_API_KEY}" \
    --smoke-model "${MODEL_NAME}" \
    --think-mode "$THINK_MODE" \
    --thinking \
    --reruns 3 \
    --reruns-delay 2 \
    --tool-json-report=tool-call-schema-report.json \
    -ra -v

This check runs walle valid.jsonl cases, sends each schema as tools[].function.parameters, forces a tool call, and uses jsonschema to validate the returned function.arguments locally. Each selected case is run once with stream=false and once with stream=true; streaming chunks are reassembled before validation. In GitLab CI, this check runs as the verify_tool_call_json_schema job; the same retry options are applied. In addition to the JUnit XML report, a JSON report (tool-call-schema-report.json) is produced containing the selected cases, per-case outcomes, and the same summary block (total / by_status / by_selection_reason / by_mode) that the original standalone script printed at the end of its log.

Arguments

ArgumentMeaningDefault
--smoke-modelVendor model nameMODEL_NAME env var
--base-urlAPI base URLKIMI_BASE_URL env var, otherwise Moonshot API
--api-keyAPI keyKIMI_API_KEY env var
--case-dirwalle validator_cases directorytestdata/walle_validator_cases/validator_cases
--selectionCase selection mode: all / explicit / objectall
--max-casesMaximum number of selected cases to runNo limit
--thinkingEnable thinking mode for the tool-call requestOff
--think-modeThinking parameter format: kimi, opensource, or noneTHINK_MODE env var, otherwise kimi
--max-tokensMaximum output tokens for each tool-call response2048
--tool-json-reportPath to the additional JSON report artifacttool-call-schema-report.json

K3 Feature Validation: tests/k3_features/

Validate K3 API features such as dynamic tools, response_format, tool_choice, and thinking effort.

uv run pytest tests/k3_features \
    --base-url "${KIMI_BASE_URL}" \
    --api-key "${KIMI_API_KEY}" \
    --smoke-model "${MODEL_NAME}" \
    -ra -v
ArgumentMeaningDefault
--smoke-modelVendor model nameMODEL_NAME env var
--base-urlAPI base URLKIMI_BASE_URL env var, otherwise Moonshot API
--api-keyAPI keyKIMI_API_KEY env var

The CI job verify_k3_features runs the pytest suite and reports results from the test log.

Prompt Token Validation: tests/prompt_tokens/

Validate that vendor-reported usage.prompt_tokens is accurate. This check sends the text cases in testdata/prompt_token_cases/cases.jsonl and vision cases in testdata/prompt_token_cases/vision_cases.jsonl with stream=true, then compares the stream-reported usage.prompt_tokens against the expected constant stored in each case. See testdata/prompt_token_cases/README.md for the case format.

uv run pytest tests/prompt_tokens \
    --base-url "${KIMI_BASE_URL}" \
    --api-key "${KIMI_API_KEY}" \
    --smoke-model "${MODEL_NAME}" \
    -ra -v

Why These Tests Are Not in Inspect-ai

tests/k3_features/, tests/prompt_tokens/, and tests/tool_call_json_schema/ validate raw chat-completions API behavior, including non-standard messages[].tools dynamic-tool declarations, malformed requests that should return HTTP 400, the location of usage fields in streaming chunks, and raw tools[].function.parameters JSON schema handling. inspect-ai is better suited for scored benchmark evaluations and may normalize or reject these raw payloads before they reach the vendor endpoint, so these tests remain standalone pytest suites.

K2.6 and Earlier

BenchmarkModeTemperatureTopPMax TokensEpochs
OCRBenchNon-Thinking0.60.95163841
OCRBenchThinking1.00.95163841
MMMUNon-Thinking0.60.95655361
MMMUThinking1.00.95655361
AIME 2025Non-Thinking0.60.959830432
AIME 2025Thinking1.00.959830432

K2.7

BenchmarkMode (thinking, preserve thinking)TemperatureTopPMax TokensEpochs
OCRBenchThinking, Keep all1.00.95163841
MMMUThinking, Keep all1.00.95655361
AIME 2025Thinking, Keep all1.00.959830432

K3

K3 no longer uses AIME 2025 and adds the reasoning_effort parameter.

BenchmarkMode (thinking, preserve thinking, effort)TemperatureTopPMax TokensEpochs
OCRBenchThinking, Keep all, low1.00.95163841
OCRBenchThinking, Keep all, high1.00.95163841
OCRBenchThinking, Keep all, max1.00.95163841
MMMUThinking, Keep all, low1.00.95983041
MMMUThinking, Keep all, high1.00.95983041
MMMUThinking, Keep all, max1.00.95983041

Running Inspect-ai Benchmarks

OCRBench

Non-Thinking

uv run python eval.py ocrbench --model kimi/your-model-id \
    --think-mode kimi --max-tokens 16384 --stream

Thinking

uv run python eval.py ocrbench --model kimi/your-model-id \
    --thinking --think-mode kimi --max-tokens 16384 --stream

Thinking + reasoning effort (K3)

uv run python eval.py ocrbench \
    --model "${THINK_MODE}/${MODEL_NAME}" \
    --max-tokens 16384 \
    --thinking \
    --think-mode "$THINK_MODE" \
    --stream \
    --max-connections 50 \
    --temperature 1.0 \
    --top-p 0.95 \
    --thinking-effort high

MMMU Pro Vision

Non-Thinking

uv run python eval.py mmmu --model kimi/your-model-id \
    --think-mode kimi --max-tokens 65536 --stream

Thinking

uv run python eval.py mmmu --model kimi/your-model-id \
    --thinking --think-mode kimi --max-tokens 65536 --stream

Thinking + reasoning effort (K3)

uv run python eval.py mmmu \
    --model "$MODEL" \
    --max-tokens 98304 \
    --thinking \
    --think-mode "$THINK_MODE" \
    --stream \
    --max-connections 50 \
    --temperature 1.0 \
    --top-p 0.95 \
    --thinking-effort high

AIME 2025

Non-Thinking

uv run python eval.py aime2025 --model kimi/your-model-id \
    --think-mode kimi --max-tokens 98304 --stream

Thinking

uv run python eval.py aime2025 --model kimi/your-model-id \
    --thinking --think-mode kimi --max-tokens 98304 --stream

Run OCRBench first as a quick deployment sanity check, then run the full MMMU and other benchmark evaluations.

Agent Benchmark (DeepSWE)

To cover that capability, we also add an agent benchmark based on the open-source DeepSWE benchmark. DeepSWE contains 113 coding-agent tasks and is used here through the open-source Pier platform. The agent runtime uses Kimi-code 0.23.6 or later.

Before running the benchmark on Pier, register kimi-code on the Pier agent first. See the Kimi-code 0.23.6 setup reference: https://github.com/MoonshotAI/kimi-code/tree/%40moonshot-ai/kimi-code%400.23.6

Inspect-ai Parameters

Available Parameters

ParameterDescriptionDefault
benchmarkEvaluation task: ocrbench, mmmu, aime2025ocrbench
--modelModel identifier, for example kimi/your-model-idRequired
--max-tokensMaximum output tokens, see recommended benchmark parametersRequired
--thinkingEnable thinking modeOff
--think-modeThinking parameter format: none, kimi, or opensourcenone
--temperatureSampling temperatureOptional; determined by server or model defaults if omitted
--top-pTop-p samplingOptional; determined by server or model defaults if omitted
--streamEnable streaming, recommended for long reasoning requestsOff
--max-connectionsMaximum concurrent connectionsPer benchmark
--epochsNumber of sampling epochsPer benchmark
--client-timeoutHTTP timeout in seconds86400
--thinking-effortK3 reasoning effort parameter: low, high, maxNone

Thinking Mode Parameters

Model TypeParameter Combinationextra_body Sent
Kimi official + thinking off--think-mode kimi{"thinking": {"type": "disabled"}}
Kimi official + thinking on--thinking --think-mode kimi{"thinking": {"type": "enabled"}}
Kimi official + reasoning effort low--thinking --think-mode kimi --thinking-effort low{"thinking": {"type": "enabled", "keep": "all", "effort": "low"}}
Kimi official + reasoning effort high--thinking --think-mode kimi --thinking-effort high{"thinking": {"type": "enabled", "keep": "all", "effort": "high"}}
Kimi official + reasoning effort max--thinking --think-mode kimi --thinking-effort max{"thinking": {"type": "enabled", "keep": "all", "effort": "max"}}
Open-source framework + thinking off--think-mode opensource{"chat_template_kwargs": {"thinking": false}}
Open-source framework + thinking on--thinking --think-mode opensource{"chat_template_kwargs": {"thinking": true}}
Open-source framework + reasoning effort low--thinking --think-mode opensource --thinking-effort low{"chat_template_kwargs": {"thinking": true, "preserve_thinking": true, "thinking_effort": "low"}}
Open-source framework + reasoning effort high--thinking --think-mode opensource --thinking-effort high{"chat_template_kwargs": {"thinking": true, "preserve_thinking": true, "thinking_effort": "high"}}
Open-source framework + reasoning effort max--thinking --think-mode opensource --thinking-effort max{"chat_template_kwargs": {"thinking": true, "preserve_thinking": true, "thinking_effort": "max"}}

Viewing Results

# View logs with inspect view
uv run inspect view

# Logs are saved in the logs/ directory

Resuming an Interrupted Benchmark

uv run inspect eval-retry logs/<log-file>.eval

Notes

AIME 2025

AIME evaluation produces many output tokens; please note the following:

  1. Timeout settings: the client default is --client-timeout 86400 (24 hours), and the server, gateway, or proxy timeout should also be long enough.
  2. Streaming: strongly recommended with --stream; non-streaming requests are more prone to timeout in thinking mode.
  3. Concurrency control: if you see many 429s or RemoteProtocolErrors, lower --max-connections.
  4. Quick validation: first run all samples with --epochs 1, then run the full evaluation after configuration is confirmed.
# Step 1: Quick validation (30 samples x 1 epoch)
uv run python eval.py aime2025 --model kimi/your-model-id \
    --thinking --think-mode kimi --max-tokens 98304 --stream --epochs 1

# Step 2: Full evaluation (30 samples x 32 epochs)
uv run python eval.py aime2025 --model kimi/your-model-id \
    --thinking --think-mode kimi --max-tokens 98304 --stream

Automatic Retry Mechanism

The following network-related errors are automatically retried with exponential backoff (1-60 seconds), no manual configuration needed:

Error TypeDescription
RateLimitError / 429Server-side rate limiting
APIConnectionErrorConnection failure
ReadError / RemoteProtocolErrorNetwork read error

Non-network errors, such as model output format issues, are not retried and are logged directly for later analysis.

Project Structure

├── eval.py              # Main evaluation CLI entrypoint
├── tests/params/        # Pre-flight parameter validation
├── tests/tool_call_json_schema/ # walle tool-call schema validation
├── tests/prompt_tokens/ # usage.prompt_tokens accuracy validation
├── tests/k3_features/   # K3 feature contract validation
├── kimi_model.py        # Kimi Model API implementation
├── aime2025.py          # AIME 2025 evaluation task
├── mmmu_pro_vision.py   # MMMU Pro Vision evaluation task
├── ocr_bench.py         # OCRBench evaluation task
├── testdata/            # JSON schema and prompt-token test data
├── beam/                # BEAM 1M context data and scripts, see the directory README
├── logs/                # Evaluation logs
└── pyproject.toml       # Project configuration

Contact

If you have any questions or suggestions, please contact contact-kvv@kimi.com.

License

This project is licensed under the MIT License. See LICENSE for details.

Contributors

Zwysilence

11 commits

dijin-dev

5 commits

xiaochen-dev

3 commits

Languages

Python

100.0%