petmal/MindTrial

MindTrial: Evaluate and compare AI language models (LLMs) on text-based tasks with optional file/image attachments and tool use. Supports multiple providers (OpenAI, Google, Anthropic, DeepSeek, Mistral AI, xAI, Alibaba, Moonshot AI, OpenRouter), custom tasks in YAML, and HTML/CSV/JSON reports.

Go

21

129 commits

updated Oct 2, 2026

See the code

See what people are saying

SourceMessageScoreDate

Benchmark notes: Sonnet 5.5 jumps from 72 to 94/98; Opus 5.5 reaches 96/98 with much less request time (r/ClaudeAI)

I maintain [**MindTrial**](https://github.com/petmal/MindTrial) and tested **Sonnet 5.5** and **Opus 5.5** on the same 98-task suite as their predecessors: 39 text tasks and 59 visual tasks, with Python/scientific libraries available and a 10-call limit per task. All four Claude runs below use the…

9

Oct 4, 2026

README

MindTrial

Build License: MPL 2.0 Go Version Go Reference

MindTrial lets you test a single AI language model (LLM) or evaluate multiple models side-by-side. It supports providers like OpenAI, Google, Anthropic, DeepSeek, Mistral AI, xAI, Alibaba, Moonshot AI, and OpenRouter. You can create your own custom tasks with text prompts, plain text or structured JSON response formats, optional file attachments, and tool use for enhanced capabilities; validate responses through exact value matching, an LLM judge for semantic evaluation, or trusted Docker-backed custom validators; and get results in easy-to-read HTML, CSV, and JSON formats.

Quick Start Guide

  1. Install the tool:

    go install github.com/petmal/mindtrial/cmd/mindtrial@latest
    
  2. Run with default settings:

    mindtrial run
    

Prerequisites

  • Go 1.26
  • Docker (for tool execution; Docker Engine 25.0+ for task services)
  • API keys from your chosen AI providers

Key Features

  • Compare multiple AI models at once
  • Create custom evaluation tasks using simple YAML files
  • Attach files or images to prompts for visual tasks
  • Enable tool use for tasks with secure sandboxed execution
  • Use LLM judges for semantic validation of complex and creative tasks
  • Evaluate stateful tool use with task-scoped Docker services and trusted custom validators
  • Get results in HTML, CSV, and JSON formats
  • Merge and compare results from multiple runs
  • Easy to extend with new AI models
  • Smart rate limiting to prevent API overload
  • Interactive mode with terminal-based UI

Basic Usage

  1. Display available commands and options:

    mindtrial help
    
  2. Run with custom configuration and output options:

    mindtrial --config="custom-config.yaml" --tasks="custom-tasks.yaml" --output-dir="./results" --output-basename="custom-tasks-results" run
    
  3. Run with specific output formats (CSV only, no HTML):

    mindtrial --csv=true --html=false run
    
  4. Run in interactive mode to select models and tasks before starting:

    mindtrial --interactive run
    
  5. Merge results from multiple runs into a single output:

    mindtrial --input="results-1.json" --input="results-2.json" --html=true --csv=true --output-basename="merged" merge-results
    
  6. Compute derived statistics (pass rate, durations, token usage, and more) grouped by provider and run:

    mindtrial --input="results.json" stats
    

Merging Results

The merge-results command combines results from multiple trial runs into a single output. Input files are specified with the --input flag (can be repeated). Currently, only JSON is supported as the input format. Use the --json=true flag during trial runs to generate JSON output files that can later be merged. The merged output can be generated in any of the supported formats (HTML, CSV, JSON) using the corresponding flags.

[!TIP] You can also use merge-results with a single input file to convert between formats. For example, if you store results in JSON, you can convert them to HTML or CSV at any time:

mindtrial --input="results.json" --html=true --csv=true merge-results

[!TIP] If some results failed due to transient errors (e.g., network timeouts), you can re-run only the failed tasks and merge the new results into the original set. Because merge-results uses a last-in-wins strategy for duplicate entries (same provider, run, and task), the corrected results will replace the failed ones.

Computing Statistics

The stats command computes derived statistics (pass rate, accuracy, error rate, duration, token usage, tool calls, error diagnostics, and estimated candidate cost) over one or more result files, filtered and grouped by task/run metadata. Like merge-results, input files are specified with the --input flag (can be repeated); when multiple files are given, they are merged first using the same last-in-wins strategy before stats are computed. This is derived analytical output, not a canonical result format, so it is only written to stdout and is not persisted back into result artifacts. Progress messages are written to stderr, so stdout stays safe to redirect straight into a file or parser for every --stats-format.

mindtrial --input="results.json" --group-by="provider,run,model" --stats-format="csv" stats
  • --group-by: Comma-separated grouping dimensions: provider, run, model, suite, category, difficulty, tag (default: provider,run); each dimension may appear at most once. Grouping by tag is exploded: a task tagged with multiple tags contributes to each of those tag groups, so tag groups overlap and are not additive; duplicate tags on the same task count once. Results missing a value for a grouping dimension are grouped under (unspecified) rather than dropped.
  • --stats-format: Output format: text, csv, json, or jsonl (default: text).
  • --provider, --run, --model, --suite, --category, --difficulty, --status: Restrict stats to matching results; each can be specified multiple times (combined with OR). --status accepts passed, failed, error, or skipped. Pass (unspecified) to match results missing a value for that field.
  • --tag: Restrict stats to results carrying the given tag; can be specified multiple times. Combined according to --tag-mode. Pass (unspecified) to match untagged results.
  • --tag-mode: How multiple --tag filters combine: all (default, every tag must be present) or any (at least one tag must be present).

[!NOTE] Token, tool-call, and duration metrics reflect only the candidate answer and any subsequent error (not judge/validation usage), matching the HTML report's dynamic summary. Median*/Stddev* metrics require at least two contributing samples; otherwise they are omitted.

TotalInputTokens is normalized: providers that count cache reads/writes separately from input (currently Anthropic) have them added, and providers that already include them do not. TotalOutputTokens is normalized the same way: providers that count reasoning separately from output (currently Google and xAI) have it added, and providers that already include it do not. TotalReasoningTokens, TotalCacheReadTokens, and TotalCacheWriteTokens sum only the counts providers actually reported. EstimatedCandidateCost/CandidateCostCurrency are priced with the candidate run's own prices and never include judge/validation usage.

Example: given a result file with these three tasks:

ProviderRunStatusTags
openaigpt-4Passedvisual, spatial
openaigpt-4Failedvisual
anthropicclaudePassedtext

Grouping by the default provider,run treats each provider/run combination as one group:

$ mindtrial --input="results.json" --stats-format="csv" stats
provider,run,Count,Passed,Failed,...
anthropic,claude,1,1,0,...
openai,gpt-4,2,1,1,...

Grouping by tag instead explodes each task into every tag it carries, so tag groups overlap rather than partition the input (the first task counts toward both visual and spatial):

$ mindtrial --input="results.json" --group-by="tag" --stats-format="csv" stats
tag,Count,Passed,Failed,...
spatial,1,1,0,...
text,1,1,0,...
visual,2,1,1,...

Configuration Guide

MindTrial uses two simple YAML files to control everything:

1. config.yaml - Application Settings

Controls how MindTrial operates, including:

  • Where to save results
  • Which AI models to use
  • API settings and rate limits

2. tasks.yaml - Task Definitions

Defines what you want to evaluate, including:

  • Questions/prompts for the AI
  • Expected answers
  • Response format rules

[!TIP] New to MindTrial? Start with the example files provided and modify them for your needs.

[!TIP] Use interactive mode with the --interactive flag to select model configurations and tasks before running, without having to edit configuration files.

config.yaml

This file defines the tool's settings and target model configurations evaluated during the trial run. The main sections include:

  • output-dir: Path to the directory where results will be saved.
  • task-source: Path to the file with definitions of tasks to run.
  • providers: List of providers (i.e. target LLM configurations) to execute tasks during the trial run.
    • name: Name of the LLM provider (e.g. openai).
    • client-config: Configuration for this provider's client (e.g. API key).
    • max-parallel-requests-per-minute: Enables parallel execution of runs within this provider and limits the aggregate number of API requests per minute across all runs. Set to 0 or omit for sequential execution (default).
    • runs: List of runs (i.e. model configurations) for this provider. Unless disabled, all configurations will be trialed.
      • name: A unique display-friendly name to be shown in the results.
      • model: Model name must be exactly as defined by the backend service's API (e.g. gpt-4o-mini).
      • disable-structured-output: Disable structured JSON responses for this run and force plain-text answers. When enabled, MindTrial:
        • treats the model's entire response as the final answer (title and explanation are filled with placeholders)
        • forces the model to use plain-text response mode
        • skips tasks that require schema-based (response-result-format) JSON outputs
      • text-only: Skip tasks that require native file input to the model API. When enabled, only tasks without native file input will be executed. Tasks with files that are only available to local tools (access: [local]) are still executed. This is useful for text-only models that cannot process images or other files natively.

[!TIP] Use text-only for models that do not support vision capabilities, such as text-only language models hosted on platforms like OpenRouter.

[!TIP] If the model can output JSON as plain text but cannot follow a provider-enforced schema, prefer text-response-format. Use disable-structured-output only when the model cannot reliably output JSON at all.

[!IMPORTANT] For models that accept an explicit response-format parameter (e.g. OpenRouter), ensure the response format is unset or set to plain text when disable-structured-output is enabled; otherwise the run will fail.

[!IMPORTANT] The disable-structured-output flag cannot be used in judge configurations, as judges require structured responses for evaluation.

[!NOTE] The OpenAI provider routes GPT-5 and newer model families through the Responses API and currently relies on stored response state (previous_response_id) for multi-turn and tool-calling flows. As a result, OpenAI Zero Data Retention (ZDR) is not currently supported for those models. Legacy OpenAI models that still use the Chat Completions API are unaffected.

[!IMPORTANT] All provider names must match exactly:

  • openai: OpenAI GPT models
  • google: Google Gemini models
  • anthropic: Anthropic Claude models
  • deepseek: DeepSeek open-source models
  • mistralai: Mistral AI models
  • xai: xAI (Grok) models
  • alibaba: Alibaba (Qwen) models
  • moonshotai: Moonshot AI (Kimi) models
  • openrouter: OpenRouter-hosted models

[!TIP] Instead of a literal value, client-config.api-key can reference an environment variable using the {{.Env.NAME}} placeholder (e.g. "{{.Env.OPENAI_API_KEY}}"), so secrets don't need to be committed to the config file. If api-key is omitted entirely (or left blank), each provider falls back to its own default environment variable:

  • openai: OPENAI_API_KEY
  • google: GOOGLE_API_KEY
  • anthropic: ANTHROPIC_API_KEY
  • deepseek: DEEPSEEK_API_KEY
  • mistralai: MISTRAL_API_KEY
  • xai: XAI_API_KEY
  • alibaba: DASHSCOPE_API_KEY
  • moonshotai: MOONSHOT_API_KEY
  • openrouter: OPENROUTER_API_KEY

This fallback also applies to judge provider configurations under judges[].provider.client-config.

[!NOTE] Anthropic and DeepSeek providers support configurable request timeout in the client-config section:

  • request-timeout: Sets the timeout duration for API requests (i.e. thinking).

Alibaba and Moonshot AI providers support endpoint configuration in the client-config section:

  • endpoint: Specifies the network endpoint URL for the API. If not specified, defaults are:
    • Alibaba: Singapore endpoint (https://dashscope-intl.aliyuncs.com/compatible-mode/v1) for better international access. For China mainland, use https://dashscope.aliyuncs.com/compatible-mode/v1.
    • Moonshot AI: Public API endpoint (https://api.moonshot.ai/v1).

[!NOTE] Some models support additional model-specific runtime configuration parameters. These can be provided in the model-parameters section of the run configuration.

Currently supported parameters for OpenAI models include:

  • text-response-format: If true, use plain-text response format (less reliable) for compatibility with models that do not support JSON.
  • reasoning-effort: Controls effort on reasoning for reasoning models. (values: none, minimal, low, medium, high, xhigh, max). The max level is supported on GPT-5.6 and later reasoning models. Legacy models may not support all values.
  • reasoning-context: Controls which prior reasoning items are reused across conversation turns (persisted reasoning). (values: auto, current_turn, all_turns). Supported on GPT-5.6 and later reasoning models. auto uses the model's default; current_turn keeps reasoning from the active turn only, without rendering earlier turns' reasoning into the next call; all_turns renders compatible reasoning items from earlier turns into the next call as well, which only has an effect when prior response items are available (e.g. via previous_response_id, which MindTrial already relies on for its multi-turn tool-calling loop).
  • reasoning-mode: Selects an alternate reasoning execution mode. (values: pro). Pro mode performs additional model work before returning a single final answer, increasing latency and token usage; use selectively for demanding tasks. Supported on GPT-5.6 and later models.
  • verbosity: Controls how many output tokens are generated. (values: low, medium, high). May not be supported by legacy models.
  • temperature: Controls randomness/creativity of responses (range: 0.0 to 2.0, default: 1.0). Lower values produce more focused and deterministic outputs.
  • top-p: Controls diversity via nucleus sampling (range: 0.0 to 1.0, default: 1.0). Lower values produce more focused outputs.
  • presence-penalty: Penalizes new tokens based on their presence in text so far (range: -2.0 to 2.0, default: 0.0). Positive values encourage model to use new tokens.
  • frequency-penalty: Penalizes new tokens based on their frequency in text so far (range: -2.0 to 2.0, default: 0.0). Positive values encourage model to use less frequent tokens.
  • max-completion-tokens: Controls the maximum number of tokens available to the model for generating a response.

Currently supported parameters for OpenRouter models include:

  • response-format: Controls response format (values: json-schema, json-object, text, default: json-schema). MindTrial adjusts its parsing logic based on this value, so always use this typed parameter rather than passing response_format directly (see note below about extra parameters).
  • temperature: Controls randomness/creativity of responses (range: 0.0 to 2.0, default: 1.0). Lower values produce more focused and deterministic outputs.
  • top-p: Controls diversity via nucleus sampling (range: 0.0 to 1.0, default: 1.0). Lower values produce more focused outputs.
  • top-k: Limits token selection to top K candidates (range: 0 or above, default: 0). Value 0 disables this setting.
  • min-p: Filters tokens below minimum probability relative to most likely token (range: 0.0 to 1.0, default: 0.0).
  • top-a: Considers tokens with sufficiently high probability relative to most likely token (range: 0.0 to 1.0, default: 0.0).
  • presence-penalty: Penalizes token repetition based on presence in input (range: -2.0 to 2.0, default: 0.0).
  • frequency-penalty: Penalizes token repetition based on frequency in input (range: -2.0 to 2.0, default: 0.0).
  • repetition-penalty: Reduces repetition of tokens from input (range: 0.0 to 2.0, default: 1.0).
  • max-tokens: Sets upper limit on generated tokens (range: 1 or above, limited by model context).
  • max-completion-tokens: Sets the modern upper limit on generated tokens. It cannot be combined with max-tokens.
  • reasoning-effort: Controls reasoning depth (values: none, minimal, low, medium, high, xhigh, max). It maps to OpenRouter's native top-level reasoning_effort field.
  • seed: Enables deterministic sampling when supported.
  • parallel-tool-calls: Enables parallel function calling during tool use (default: true).
  • verbosity: Controls response verbosity (values: low, medium, high, default: medium).
  • server-tools: Injects provider-managed server-side tools into every request for this run. Each entry requires a type (the tool identifier, e.g. openrouter:fusion) and an optional parameters map whose fields are tool-specific. Server tools are appended to the request's tools array alongside any local Docker-based tools. Use an extra parameter (e.g. tool_choice: required) to control invocation behavior.

Any additional parameters not listed above can be specified directly in model-parameters and will be forwarded to the OpenRouter API. Use for provider-specific or OpenRouter-specific parameters. Prefer typed parameters where they exist. Note: if both a typed parameter and an equivalent extra parameter are specified (e.g., max-tokens: 100 and max_tokens: 500), the extra parameter takes precedence and the API will receive the extra parameter's value.

Currently supported parameters for Anthropic models include:

  • max-tokens: Controls the maximum number of tokens available to the model for generating a response.
  • thinking-budget-tokens: Enables extended thinking with a fixed token budget, giving the model more reasoning capacity on complex tasks. Must be at least 1024 and less than max-tokens. Ignored when effort is also set. Without either setting, thinking follows the model default. Deprecated: Claude Opus 4.7+ removed fixed thinking budgets; setting this returns a 400 error. Use effort instead.
  • effort: Enables adaptive extended thinking and guides how deeply the model reasons before responding, from quick answers (low) to thorough multi-step reasoning (max) (values: low, medium, high, xhigh, max). Without effort or thinking-budget-tokens, thinking follows the model default. When set, thinking-budget-tokens is ignored. Use max-tokens to cap total output (thinking + response text). The xhigh level is recommended for coding and agentic use cases on Claude Opus 4.7+.
  • thinking: Explicitly overrides the thinking mode (values: disabled, between_tools). disabled explicitly disables thinking and cannot be combined with thinking-budget-tokens or effort above high. between_tools, available on Claude Sonnet 5.5, disables up-front thinking while allowing progress thinking between tool calls; it supports low, medium, and high effort and cannot be combined with thinking-budget-tokens. When omitted, the model's default thinking behaviour is used unless effort or thinking-budget-tokens requests an explicit mode.
  • temperature: Controls randomness/creativity of responses (range: 0.0 to 1.0, default: 1.0). Lower values produce more focused and deterministic outputs. Deprecated: Claude Opus 4.7+ rejects non-default values with a 400 error.
  • top-p: Controls diversity via nucleus sampling (range: 0.0 to 1.0). Lower values produce more focused outputs. Deprecated: Claude Opus 4.7+ rejects non-default values with a 400 error.
  • top-k: Limits tokens considered for each position to top K options. Higher values allow more diverse outputs. Deprecated: Claude Opus 4.7+ rejects any value with a 400 error.
  • stream: If true, enables streaming mode for the API response. Streaming is recommended for requests with large max-tokens values, especially when extended thinking is enabled, to prevent HTTP timeouts on long-running requests. Responses are streamed incrementally and buffered internally before processing.
  • legacy-structured-output: If true, uses tool-based structured output instead of native JSON schema output. This is a workaround for models that have difficulty producing valid responses with native output_config.format constrained decoding when extended thinking is enabled. When set, the provider registers a submit_response tool and instructs the model to use it to submit its response.
  • prompt-cache-ttl: Enables Anthropic prompt caching and selects the cache lifetime. Supported values are 5m and 1h. When set, MindTrial enables top-level automatic caching and places one explicit cache breakpoint on the most reusable request prefix: the final configured local tool, the final system block, or the final cacheable initial user content block. When omitted, no prompt cache controls are added.

Currently supported parameters for Google models include:

  • text-response-format: If true, use plain-text response format (less reliable) for compatibility with models that do not support JSON. This setting applies to all tasks, including those with and without tools enabled.
  • text-response-format-with-tools: If true, forces plain-text response format when tools are enabled (required for pre-Gemini 3 models). If false or unset, uses JSON schema mode with tools (Gemini 3+ default behavior). This setting only applies to tasks with tools enabled.
  • thinking-level: Controls the maximum depth of the model's internal reasoning process (values: minimal, low, medium, high). The default is model-dependent — for example, Gemini 3 Pro defaults to high, while Gemini 3.5 Flash defaults to medium. minimal minimizes reasoning for lowest latency (does not guarantee thinking is disabled), low minimizes latency and cost for simple tasks, medium balances reasoning depth and latency, while high maximizes reasoning depth for complex tasks (the model may take longer but output is more carefully reasoned).
  • media-resolution: Controls the maximum number of tokens allocated per input image (values: low, medium, high). Higher resolutions improve fine text reading and small detail identification but increase token usage and latency. low uses 280 tokens; medium uses 560 tokens; high uses 1120 tokens. If unspecified, the model uses optimal defaults.
  • temperature: Controls randomness/creativity of responses (range: 0.0 to 2.0, default: 1.0). Lower values produce more focused and deterministic outputs. For Gemini 3, it's recommended to keep temperature at default 1.0 for optimal reasoning performance.
  • top-p: Controls diversity via nucleus sampling (range: 0.0 to 1.0). Lower values produce more focused outputs.
  • top-k: Limits tokens considered for each position to top K options. Higher values allow more diverse outputs.
  • presence-penalty: Penalizes new tokens based on whether they appear in the text so far. Positive values discourage reuse of tokens, increasing vocabulary. Negative values encourage token reuse.
  • frequency-penalty: Penalizes new tokens based on their frequency in the text so far. Positive values discourage frequent tokens proportionally. Negative values encourage token repetition.
  • seed: Seed used for deterministic generation. When set, the model attempts to provide consistent responses for identical inputs.

Currently supported parameters for DeepSeek models include:

  • temperature: Controls randomness/creativity of responses (range: 0.0 to 2.0, default: 1.0). Lower values produce more focused and deterministic outputs.
  • top-p: Controls diversity via nucleus sampling (range: 0.0 to 1.0). Lower values produce more focused outputs.
  • presence-penalty: Penalizes new tokens based on their presence in text so far (range: -2.0 to 2.0, default: 0.0). Positive values encourage model to use new tokens.
  • frequency-penalty: Penalizes new tokens based on their frequency in text so far (range: -2.0 to 2.0, default: 0.0). Positive values encourage model to use less frequent tokens.
  • thinking: Toggles reasoning (thinking) mode for V4 and newer thinking-capable models. When enabled, the model produces chain-of-thought reasoning before the final answer (values: enabled, disabled; default: enabled for V4 models). Note: temperature, top-p, presence-penalty, and frequency-penalty are silently ignored in thinking mode.
  • reasoning-effort: Controls how deeply the model reasons in thinking mode (values: low, medium, high, xhigh, max). Default is high. Currently only high and max are distinct effective values — low and medium are mapped to high, and xhigh is mapped to max.
  • max-tokens: Controls the maximum number of tokens generated by the model.

Currently supported parameters for Mistral AI models include:

  • temperature: Controls randomness/creativity of responses (range: 0.0 to 1.5). Lower values produce more focused and deterministic outputs.
  • top-p: Controls diversity via nucleus sampling (range: 0.0 to 1.0). Lower values produce more focused outputs.
  • max-tokens: Controls the maximum number of tokens available to the model for generating a response.
  • presence-penalty: Penalizes new tokens based on their presence in text so far (range: -2.0 to 2.0, default: 0.0). Positive values encourage model to use new tokens.
  • frequency-penalty: Penalizes new tokens based on their frequency in text so far (range: -2.0 to 2.0, default: 0.0). Positive values encourage model to use less frequent tokens.
  • random-seed: Provides the seed to use for random sampling. If set, requests will generate deterministic results.
  • prompt-mode: When set to "reasoning", instructs the model to reason if supported.
  • reasoning-effort: Controls reasoning depth for current models (values: none, minimal, low, medium, high, xhigh). This is independent of the legacy prompt-mode; high can increase output-token cost.
  • safe-prompt: Enables content filtering to ensure outputs comply with usage policies.

Currently supported parameters for xAI models include:

  • temperature: Controls randomness/creativity of responses (range: 0.0 to 2.0, default: 1.0). Lower values produce more focused and deterministic outputs.
  • top-p: Controls diversity via nucleus sampling (range: 0.0 to 1.0, default: 1.0). Lower values produce more focused outputs.
  • max-completion-tokens: Controls the maximum number of tokens available to the model for generating a response.
  • presence-penalty: Penalizes new tokens based on their presence in text so far (range: -2.0 to 2.0, default: 0.0). Positive values encourage model to use new tokens.
  • frequency-penalty: Penalizes new tokens based on their frequency in text so far (range: -2.0 to 2.0, default: 0.0). Positive values encourage model to use less frequent tokens.
  • reasoning-effort: Controls effort on reasoning for supported reasoning-capable models (values: low, medium, high, xhigh). Not all xAI reasoning models (i.e. Grok 4) accept this parameter; Grok 4.5 defaults to high when unset. The xhigh level is supported on Grok 4.6 and later.
  • seed: Integer seed to request deterministic sampling when possible. Determinism is best-effort. xAI makes a best-effort to return repeatable outputs for identical inputs when seed and other parameters are the same.

Currently supported parameters for Alibaba models include:

  • response-format: Selects json-schema, json-object, or text. MindTrial applies its legacy schema-instruction behavior when this is omitted and structured output is not disabled. Cannot be combined with the deprecated text-response-format or disable-legacy-json-mode properties.
  • stream: If true, enables streaming mode for the API response. Some models (e.g. QwQ, QVQ, and Qwen-Omni) require streaming to be enabled. Responses are streamed incrementally and buffered internally before processing.
  • enable-thinking: Enables hybrid thinking on supported Qwen models.
  • preserve-thinking: Preserves reasoning_content across tool-call turns. Preserved reasoning is included in later input-token counts and billing.
  • thinking-budget: Optional positive token budget for thinking. This is distinct from max-tokens/max-completion-tokens, which limit the complete generated response. Mutually exclusive with reasoning-effort.
  • reasoning-effort: Controls reasoning depth for supported Qwen models (e.g. Qwen 3.8 Max) (values: none, minimal, low, medium, high, xhigh, max). Mutually exclusive with thinking-budget.
  • temperature: Controls randomness/creativity of responses (range: 0.0 to 2.0, default: 1.0). Lower values produce more focused and deterministic outputs.
  • top-p: Controls diversity via nucleus sampling (range: 0.0 to 1.0). Lower values produce more focused outputs.
  • max-tokens: Controls the maximum number of tokens available to the model for generating a response. Deprecated: use max-completion-tokens instead. Mutually exclusive with max-completion-tokens.
  • max-completion-tokens: Controls the maximum number of tokens available to the model for generating a response, including reasoning tokens for thinking models. Mutually exclusive with max-tokens.
  • presence-penalty: Penalizes new tokens based on whether they appear in the text so far (range: -2.0 to 2.0, default: 0.0). Positive values encourage introducing new topics.
  • frequency-penalty: Penalizes new tokens based on their frequency in text so far (range: -2.0 to 2.0, default: 0.0). Positive values encourage model to use less frequent tokens.
  • seed: Makes text generation more deterministic by using the same seed value. When using the same seed and keeping other parameters unchanged, the model makes best-effort to return consistent outputs for identical inputs.
  • text-response-format: If true, use plain-text response format (less reliable) for compatibility with models that do not support JSON (for example, when thinking is enabled on certain Qwen models). Deprecated: use response-format: text instead. Cannot be combined with response-format.
  • disable-legacy-json-mode: Compatibility toggle that controls legacy prompt injection for JSON formatting. Default: false (legacy mode on), which adds an explicit JSON formatting instruction to the prompt for improved compatibility with most Qwen models. Setting this to true disables the legacy prompt injection. For best compatibility and reliable JSON responses, keep this set to false unless you are certain the target model works correctly without legacy prompt injection. Deprecated: use response-format: json-schema instead. Cannot be combined with response-format.

Currently supported parameters for Moonshot AI models include:

  • temperature: Controls randomness/creativity of responses (range: 0.0 to 1.0, default: 0.0). Higher values make output more random, while lower values make it more focused and deterministic. Moonshot AI recommends 0.6 for kimi-k2 models and 1.0 for kimi-k2-thinking models.
  • top-p: Controls diversity via nucleus sampling (range: 0.0 to 1.0, default: 1.0). Lower values produce more focused outputs. Generally, change either this or temperature, but not both at the same time.
  • max-tokens: Controls the maximum number of tokens to generate for the chat completion.
  • max-completion-tokens: Controls the modern maximum number of generated tokens. It cannot be combined with deprecated max-tokens.
  • reasoning-effort: Controls Kimi K3 reasoning depth (values: low, high, max).
  • response-format: Selects json-schema, json-object, or text. MindTrial applies json-object when this is omitted (unless structured output is disabled). Kimi K3 can use native json-schema.
  • stream: Enables streaming and usage accumulation for long-running responses.
  • presence-penalty: Penalizes new tokens based on whether they appear in the text (range: -2.0 to 2.0, default: 0.0). Positive values increase the likelihood of the model discussing new topics.
  • frequency-penalty: Penalizes new tokens based on their existing frequency in the text (range: -2.0 to 2.0, default: 0.0). Positive values reduce the likelihood of the model repeating the same phrases verbatim.
  • thinking: Toggles the reasoning (thinking) capability for thinking-capable models such as kimi-k2.6. Accepted values are enabled (default for kimi-k2.6) and disabled. Older Kimi models that do not support this parameter should omit it.
  • preserve-thinking: Enables Moonshot's Preserved Thinking feature for kimi-k2.6, which preserves the model's chain-of-thought across model calls that share the same conversation context (e.g. successive calls in a tool-using task), so the model can build on its earlier reasoning. Accepted value: all; when omitted, prior reasoning is dropped between calls — reducing token cost at the expense of chain-of-thought continuity. Older Kimi models do not support this parameter and should omit it.

For kimi-k2.5 and kimi-k2.6, Moonshot AI fixes temperature, top-p, presence-penalty, and frequency-penalty to model-specific defaults — supplying any of these parameters will cause the API to reject the request.

[!NOTE] The results will be saved to <output-dir>/<output-basename>.<format>. If the result output file already exists, it will be replaced. If the log file already exists, it will be appended to.

[!TIP] The following placeholders are available for output paths and names:

  • {{.Year}}: Current year
  • {{.Month}}: Current month
  • {{.Day}}: Current day
  • {{.Hour}}: Current hour
  • {{.Minute}}: Current minute
  • {{.Second}}: Current second

[!TIP] If log-file and/or output-basename is blank, the log and/or output will be written to the stdout.

[!NOTE] MindTrial processes tasks across different AI providers simultaneously (in parallel). However, when running multiple configurations from the same provider (e.g. different OpenAI models), these are processed one after another (sequentially) by default.

[!TIP] To run multiple configurations from the same provider in parallel, set max-parallel-requests-per-minute on the provider. This enables parallel execution of all runs within that provider, while limiting the aggregate number of API requests per minute across all runs to the specified value. When set to 0 (or omitted), runs execute sequentially (the default behavior).

[!TIP] Models can use the max-requests-per-minute property in their run configurations to limit the number of requests made per minute.

[!TIP] To automatically retry failed requests due to rate limiting or other transient errors, set retry-policy at the provider level to apply to all runs. An individual run configuration can override this by setting its own retry-policy:

  • max-retry-attempts: Maximum number of retry attempts (default: 0 means no retry).
  • initial-delay-seconds: Initial delay before the first retry in seconds.

Retries use exponential backoff starting with the initial delay.

[!TIP] To estimate what a trial run costs, set pricing at the application, provider, or run level. Rates are per million tokens:

  • currency: ISO 4217 code the rates are expressed in (default: USD).
  • input-per-million: Price per million uncached input tokens.
  • output-per-million: Price per million generated output tokens.
  • cache-read-per-million: Price per million input tokens read from a prompt cache (default: the input price).
  • cache-write-per-million: Price per million input tokens written to a prompt cache (default: the input price).
  • reasoning-per-million: Price per million reasoning tokens (default: the output price).

pricing is inherited as a whole, not field-by-field: the nearest level (run, then provider, then application) that sets any field is used in its entirety, replacing rather than merging with whatever a less specific level configured. To override a single rate, restate the full price list at that level:

config:
  pricing:
    currency: USD
    input-per-million: 1.25
    output-per-million: 10.00
  providers:
    - name: openai
      runs:
        - name: "GPT-5.2"
          model: "gpt-5.2"
          # Overriding output-per-million requires restating the rest of the list.
          pricing:
            currency: USD
            input-per-million: 1.25
            output-per-million: 12.00

Estimated costs are reported by the stats command and are always estimates derived from these configured rates, never billed amounts. A rate left unset is treated as unknown rather than free, so any estimate depending on it is omitted instead of being understated. The effective prices are also recorded in JSON results, so historical estimates stay reproducible.

[!TIP] To disable all run configurations for a given provider, set disabled: true on that provider. An individual run configuration can override this by setting disabled: false (e.g. to enable just that one configuration).

Example snippet from config.yaml:

# config.yaml
config:
  log-file: ""
  output-dir: "./results/{{.Year}}-{{.Month}}-{{.Day}}/"
  output-basename: "{{.Hour}}-{{.Minute}}-{{.Second}}"
  task-source: "./tasks.yaml"
  providers:
    - name: openai
      disabled: true
      client-config:
        # Resolved from the OPENAI_API_KEY environment variable at load time.
        api-key: "{{.Env.OPENAI_API_KEY}}"
      retry-policy:
        max-retry-attempts: 5
        initial-delay-seconds: 30
      runs:
        - name: "4o-mini - latest"
          disabled: false
          model: "gpt-4o-mini"
          max-requests-per-minute: 3
        - name: "o1-mini - latest"
          model: "o1-mini"
          max-requests-per-minute: 3
          model-parameters:
            text-response-format: true
        - name: "o3-mini - latest (high reasoning)"
          model: "o3-mini"
          max-requests-per-minute: 3
          model-parameters:
            reasoning-effort: "high"
    - name: openrouter
      retry-policy:
        max-retry-attempts: 5
        initial-delay-seconds: 30
      max-parallel-requests-per-minute: 30
      client-config:
        api-key: "<your-api-key>"
      runs:
        - name: "OpenAI GPT-5.2 (xhigh reasoning)"
          model: "openai/gpt-5.2"
          max-requests-per-minute: 20
          model-parameters:
            verbosity: "medium"
            # Pass-through parameters use OpenAI API naming (underscores).
            reasoning_effort: "xhigh"
        - name: "GPT via Fusion (xhigh panel + judge, high outer)"
          # Outer model.
          model: "~openai/gpt-latest"
          max-requests-per-minute: 2
          model-parameters:
            # Outer model reasoning effort.
            reasoning:
              effort: "high"
            server-tools:
              - type: openrouter:fusion
                parameters:
                  # Inner panel models.
                  analysis_models:
                    - "~anthropic/claude-opus-latest"
                    - "~openai/gpt-latest"
                    - "~google/gemini-pro-latest"
                  # Judge model.
                  model: "~anthropic/claude-opus-latest"
                  # Inner panel + judge parameters.
                  reasoning:
                    effort: "xhigh"
                  max_completion_tokens: 65536
                  max_tool_calls: 16
        - name: "Google Gemma 3 27B IT (free)"
          model: "google/gemma-3-27b-it:free"
          max-requests-per-minute: 3
          model-parameters:
            response-format: "text"
    - name: google
      client-config:
        api-key: "<your-api-key>"
      runs:
        - name: "Gemini 2.5 Pro - latest"
          model: "gemini-2.5-pro"
          max-requests-per-minute: 3
          model-parameters:
            text-response-format-with-tools: true
        - name: "Gemini 3 Pro - latest"
          model: "gemini-3-pro-preview"
          max-requests-per-minute: 3
          model-parameters:
            thinking-level: "high"
            media-resolution: "high"
    - name: anthropic
      client-config:
        api-key: "<your-api-key>"
      runs:
        - name: "Claude 3.7 Sonnet - latest"
          model: "claude-3-7-sonnet-latest"
          max-requests-per-minute: 5
          model-parameters:
            max-tokens: 4096
        - name: "Claude 3.7 Sonnet - latest (extended thinking)"
          model: "claude-3-7-sonnet-latest"
          max-requests-per-minute: 5
          model-parameters:
            max-tokens: 8192
            thinking-budget-tokens: 2048
            stream: true
        - name: "Claude 4.6 Opus - latest (max adaptive thinking)"
          model: "claude-opus-4-6"
          max-requests-per-minute: 5
          model-parameters:
            max-tokens: 65536
            effort: max
            stream: true
        - name: "Claude Opus 4.7 (xhigh adaptive thinking)"
          model: "claude-opus-4-7"
          max-requests-per-minute: 5
          model-parameters:
            max-tokens: 65536
            effort: xhigh
            stream: true
        - name: "Claude Opus 5 (xhigh adaptive thinking with prompt caching)"
          model: "claude-opus-5"
          max-requests-per-minute: 5
          model-parameters:
            max-tokens: 65536
            effort: xhigh
            stream: true
            prompt-cache-ttl: 5m
    - name: deepseek
      client-config:
        api-key: "<your-api-key>"
        request-timeout: 10m
      runs:
        - name: "DeepSeek-V3.1 - latest (thinking mode)"
          model: "deepseek-reasoner"
          max-requests-per-minute: 15
    - name: mistralai
      client-config:
        api-key: "<your-api-key>"
      runs:
        - name: "Mistral Large - latest"
          model: "mistral-large-latest"
          max-requests-per-minute: 5
          retry-policy:
            max-retry-attempts: 5
            initial-delay-seconds: 30
        - name: "Mistral Medium 3.5 - latest (high reasoning)"
          model: "mistral-medium-3-5"
          max-requests-per-minute: 5
          model-parameters:
            max-tokens: 65536
            reasoning-effort: "high"
    - name: alibaba
      client-config:
        api-key: "<your-api-key>"
        endpoint: "https://dashscope-intl.aliyuncs.com/compatible-mode/v1"  # Singapore region
      retry-policy:
        max-retry-attempts: 5
        initial-delay-seconds: 30
      runs:
        - name: "Qwen3-Max-Preview"
          model: "qwen3-max-preview"
          max-requests-per-minute: 30
        - name: "Qwen3-Max-Preview - unstructured"
          model: "qwen3-max-preview"
          disable-structured-output: true
          max-requests-per-minute: 30
        - name: "Qwen-VL-Max-Latest"
          model: "qwen-vl-max-latest"
          max-requests-per-minute: 30
          model-parameters:
            disable-legacy-json-mode: true
        - name: "Qwen3-Next-80B-A3B-Thinking"
          model: "qwen3-next-80b-a3b-thinking"
          max-requests-per-minute: 30
          model-parameters:
            text-response-format: true
        - name: "QVQ-Max (vision reasoning)"
          model: "qvq-max"
          max-requests-per-minute: 30
          model-parameters:
            stream: true  # Required for QvQ models
        - name: "Qwen3.7 Plus - latest (thinking)"
          model: "qwen3.7-plus"
          max-requests-per-minute: 30
          model-parameters:
            enable-thinking: true
            preserve-thinking: true
            max-tokens: 65536
            stream: true
        - name: "Qwen3.8 Max - latest (xhigh reasoning)"
          model: "qwen3.8-max"
          max-requests-per-minute: 30
          model-parameters:
            reasoning-effort: "xhigh"
            preserve-thinking: true
            max-completion-tokens: 65536
            response-format: "json-object"
            stream: true
    - name: moonshotai
      client-config:
        api-key: "<your-api-key>"
      runs:
        - name: "Kimi K2 - latest (thinking)"
          model: "kimi-k2-thinking"
          text-only: true  # Skip tasks that require native file input
          max-requests-per-minute: 3
          model-parameters:
            temperature: 1.0
            max-tokens: 16000
        - name: "Kimi K2.6 (thinking)"
          model: "kimi-k2.6"
          max-requests-per-minute: 3
          model-parameters:
            max-tokens: 32000
            thinking: enabled     # default for kimi-k2.6; "disabled" turns off reasoning
            preserve-thinking: all  # preserve chain-of-thought across model calls (Preserved Thinking)
        - name: "Kimi K3 - latest (max reasoning)"
          model: "kimi-k3"
          max-requests-per-minute: 3
          model-parameters:
            reasoning-effort: "max"
            max-completion-tokens: 65536
            response-format: "json-schema"
            stream: true

tasks.yaml

This file defines the tasks to be executed on all enabled run configurations. Each task defines the following properties; name, prompt, and response-result-format are always required, and expected-result is required unless the task uses a custom validator:

  • name: A unique display-friendly name to be shown in the results.
  • prompt: The prompt (i.e. task) that will be sent to the AI model.
  • response-result-format: Defines how the AI should format the final answer to the prompt. This can be either:
    • Plain text format: A string instruction describing the expected answer format (e.g., "single number", "list of words separated by commas").
    • Structured schema format: A JSON schema object defining the structure of the expected response for complex data (e.g., objects with specific fields and types).
  • expected-result: Defines the accepted valid answer(s) to the prompt. The format depends on the response-result-format type:
    • For plain text format: A string value or list of string values that follow the format instruction precisely.
    • For structured schema format: An object value or list of object values that conform to the JSON schema definition. Only one expected result needs to match for the response to be considered correct. With a custom validator, expected-result is optional trusted reference data that is passed to the validator unchanged and does not need to match response-result-format.

Optionally, a task can include a list of files to be sent along with the prompt:

  • files: A list of files to attach to the prompt. Each file entry defines the following properties:

    • name: A unique name for the file. Every attached file is announced to the model in the prompt as [file: <name>], regardless of its access setting, and local tools mount the file under this exact name.
    • uri: The path or URI to the file. Local file paths and remote HTTP/HTTPS URLs are supported. The content is loaded on demand: it is sent with the request for files with native access, and copied to the tool auxiliary directory for files with local access.
    • type: The MIME type of the file (e.g., image/png, image/jpeg). If omitted, the tool will attempt to infer the type based on the file extension or content.
    • options: Optional per-file processing options that override the file-options defaults from the task-config section.
      • image-detail: Controls the fidelity level at which the model processes input images (values: auto, low, medium, high, original). If the provider does not natively support the requested level, the next higher level or the highest available level is selected. If not set or unknown, the provider uses its own default behavior. Currently only the OpenAI provider honors this setting. It also controls how PDF pages are rendered as images, but only on the OpenAI Responses API; the Chat Completions API does not accept a detail level for file inputs. The original level has no PDF equivalent and is treated as high.
      • access: Controls how the file is exposed. Supported values are native (provider native file/multimodal input) and local (local Docker tools). When omitted, the stable default is [native, local]. A per-file access replaces the inherited file-options value entirely (no union). An explicit empty list [] is invalid. The setting applies to every file type including images, so an image with access: [local] is never sent to the model and can only be inspected through a tool. Tasks with files that are only available to local tools (access: [local]) do not require provider file support and are not skipped by text-only.
  • file-options: Default file processing options for all task files in task-config. Individual files can override via files[].options.

    • image-detail: Default image fidelity (same values as above).
    • access: Default access list (same semantics as above). Example: file-options: { access: [native, local] }.

[!NOTE] If a task requires native file input (any file with native access), it will be skipped for provider configurations that do not support file uploads or the specific file type. Tasks with only local access never require native file support.

[!IMPORTANT] A file with local access is only readable if the task also enables at least one tool that defines auxiliary-dir. Otherwise the model sees the [file: <name>] reference but has no way to read the contents.

[!NOTE] Currently supported image types include: image/jpeg, image/jpg, image/png, image/gif, image/webp. Support may vary by provider. Native non-image (document) input is additionally supported by:

  • OpenAI: the complete accepted file types list — PDF; Word, Excel, PowerPoint, Pages, Keynote, Google Docs/Sheets/Slides, RTF and OpenDocument text; CSV, TSV and IIF; and a broad set of text and code formats including plain text, Markdown, HTML, XML, CSS, JSON, YAML, TOML, calendar, vCard, subtitles, email and most programming languages
  • Google: application/pdf, application/json, text/plain, text/html, text/css, text/xml, text/csv, text/rtf and text/javascript, per the supported content types. Only PDF is read with document vision; the other types are extracted as plain text, so charts and formatting are lost.
  • Anthropic: application/pdf and text/plain, per the citations documentation. A plain text document is sent verbatim rather than base64-encoded; other text formats such as CSV or Markdown must declare type: "text/plain" explicitly to use this path.
  • OpenRouter, Mistral AI: application/pdf
  • Alibaba, Moonshot AI, xAI: images only
  • DeepSeek: images only on vision-capable models

A file with native access whose type is not supported by the selected provider makes the task unsupported; it is never silently downgraded to local access.

Example file access configuration:

task-config:
  file-options:
    access: [native, local]  # default for all files
  tasks:
    - name: "vision task"
      prompt: "Describe the image."
      response-result-format: "single sentence"
      expected-result: "A cat."
      files:
        - name: "picture"
          uri: "./taskdata/cat.png"
          type: "image/png"
          # inherits access: [native, local]
    - name: "local-only task"
      prompt: "Use the data file to answer."
      response-result-format: "single number"
      expected-result: "42"
      files:
        - name: "data"
          uri: "./taskdata/data.csv"
          type: "text/csv"
          options:
            access: [local]  # mounted to tool auxiliary-dir only, never sent natively

[!TIP] To disable all tasks by default, set disabled: true in the task-config section. An individual task can override this by setting disabled: false (e.g. to enable just that one task).

Optionally, a task can also carry descriptive metadata used for filtering and grouping:

  • suite: A grouping label for organizing related tasks (e.g. a benchmark suite name).
  • category: A classification label for the task (e.g. "math", "coding").
  • difficulty: A free-form difficulty label for the task (e.g. "easy", "hard").
  • tags: A list of free-form labels for filtering and grouping tasks.

These fields are optional and have no effect on task execution or validation. When present, they are included in the JSON results (TaskMetadata), can be filtered on in the HTML report alongside the existing status/task filters, and are used for the suite/category/difficulty cycling hotkeys (s/c/d) and tag search (/) in the interactive task picker's checklist.

- name: "math problem"
  suite: "arithmetic-basics"
  category: "math"
  difficulty: "easy"
  tags: ["smoke", "regression"]
  prompt: "What is 2 + 2?"
  response-result-format: "single number"
  expected-result: "4"

Structured Response Formats

MindTrial supports two types of response formats for tasks:

Plain Text Format

For tasks where the final answer can be represented as a text value:

- name: "math problem"
  prompt: "What is 2 + 2?"
  response-result-format: "single number"
  expected-result: "4"
Structured Schema Format

For tasks requiring complex structured answers, you can define a JSON schema that describes the expected response format:

- name: "perfect square check"
  prompt: "For each number in [4, 9, 10, 16], determine if it's a perfect square and if so, provide the square root."
  response-result-format:
    type: array
    items:
      type: object
      additionalProperties: false
      properties:
        number:
          type: integer
        is_perfect_square:
          type: boolean
        square_root:
          type: integer
      required: ["number", "is_perfect_square"]
  expected-result:
    - - number: 4
        is_perfect_square: true
        square_root: 2
      - number: 9
        is_perfect_square: true
        square_root: 3
      - number: 10
        is_perfect_square: false
      - number: 16
        is_perfect_square: true
        square_root: 4

[!IMPORTANT] Structured schema format caveats:

  • Semantic validation (LLM judges) cannot be used with structured schema-based response formats.
  • All expected results must be objects that conform to the same schema. For array schemas, the entire expected array must be wrapped in a single list item under expected-result to avoid treating each array element as a separate expected answer.
  • Models must support structured JSON response generation for reliable results.
  • The OpenAI provider requires JSON schemas to have additionalProperties: false and all fields must be required (no optional fields allowed). Other providers may be more flexible.

System Prompt

The system prompt controls how the response format instruction is presented to the AI model.

You can customize this template globally for all tasks in the task-config section, and override it for individual tasks if needed. The template uses Go's template syntax and can reference {{.ResponseResultFormat}} to include the task's response-result-format.

Default system prompt for all tasks is:

Provide the final answer in exactly this format: {{.ResponseResultFormat}}

  • system-prompt: A configuration section for the system prompt.
    • template: The template string for the system prompt instruction. If not specified, uses the default.
    • enable-for: Controls when system prompt should be sent to AI models. Options:
      • "all": Send system prompt for all tasks (both plain text and structured schema formats).
      • "text": Send system prompt only for tasks with plain text response format (default).
      • "none": Do not send system prompt.

[!NOTE] For structured schema response formats, the JSON schema is automatically passed to the AI model through the provider's structured response mechanism, making explicit format instructions in the system prompt optional.

Validation Rules

These rules control how the validator compares the model's answer to the expected results. By default, comparisons are case-insensitive and only trim leading and trailing whitespace.

You can set validation rules globally for all tasks in the task-config section, and override them for individual tasks if needed; any option not specified at the task level will inherit the global setting from task-config:

  • validation-rules: Controls how model responses are validated against expected results.
    • case-sensitive: If true, comparison is case-sensitive. If false (default), comparison ignores case.
    • ignore-whitespace: If true, all whitespace (spaces, tabs, newlines) is removed before comparison. If false (default), only leading/trailing whitespace is trimmed, and internal whitespace is preserved.
    • trim-lines: If true, trims leading and trailing whitespace from each line before comparison while preserving internal spaces within lines. CRLF line endings are normalized to LF. This option is ignored when ignore-whitespace is enabled. If false (default), lines are not individually trimmed.
    • schema-validation: If true, validates the raw candidate answer against the single expected-result JSON Schema. See Schema-Based Validation for details. Mutually exclusive with judge and custom-validator.
    • judge: Optional LLM-based semantic validation. See Judge-Based Validation for details. Mutually exclusive with schema-validation and custom-validator.
      • enabled: If true, uses an LLM judge to evaluate semantic equivalence. If false (default), uses exact value matching.
      • name: The name of the judge configuration defined in the config.yaml file.
      • variant: The specific run variant from the judge's provider to use.
    • custom-validator: Name of a trusted Docker-backed validator defined in the config.yaml file that decides whether the answer is correct. See Custom Validators for details. Mutually exclusive with schema-validation and judge. Set it to "" in a task to opt out of an inherited custom validator.

Judge-Based Validation

For complex or open-ended tasks where exact value matching is insufficient, you can configure LLM judges to evaluate responses semantically. This is particularly useful for creative writing, reasoning tasks, or when multiple valid answer formats exist.

How it works: Instead of comparing text exactly, an LLM judge evaluates whether the model's response semantically matches the expected result, considering meaning and intent rather than exact wording.

To use judge validation:

  1. Define and configure judge models in config.yaml:

    config:
      # ... existing configuration ...
      judges:
        - name: "mistral-judge"  # A unique name for the judge configuration.
          provider:
            name: "mistralai"
            client-config:
              api-key: "<your-api-key>"
            runs:
              - name: "fast"
                model: "mistral-medium-latest"
                max-requests-per-minute: 30
                model-parameters:
                  temperature: 0.20
                  random-seed: 847629
              - name: "reasoning"
                model: "magistral-medium-latest"
                max-requests-per-minute: 30
                model-parameters:
                  prompt-mode: "reasoning"
                  temperature: 0.20
                  random-seed: 847629
        - name: "deepseek-judge"
          provider:
            name: "deepseek"
            client-config:
              api-key: "<your-api-key>"
            runs:
              - name: "fast"
                model: "deepseek-chat"
                max-requests-per-minute: 30
                model-parameters:
                  temperature: 0.20
              - name: "reasoning"
                model: "deepseek-reasoner"
                max-requests-per-minute: 30
    
  2. Enable judge validation in tasks.yaml:

    # Enable globally for all tasks.
    task-config:
      validation-rules:
        judge:
          enabled: true
          name: "mistral-judge"
          variant: "fast"
      tasks:
        # ... tasks will use judge validation by default ...
    
    # Override per-task (inherit global settings and override specific options).
    task-config:
      validation-rules:
        judge:
          enabled: false  # Default: use exact value matching.
          name: "mistral-judge"
          variant: "fast"
      tasks:
        - name: "exact matching task"
          prompt: "What is 2+2?"
          response-result-format: "single number"
          expected-result: "4"
          # Inherits global validation-rules (exact value matching).
        
        - name: "creative writing task"
          prompt: "Write a short story about..."
          response-result-format: "short story narrative"
          expected-result: "A creative and engaging short story"
          validation-rules:
            judge:
              enabled: true  # Override: enable judge validation for this task.
              # Inherits name: "mistral-judge" and variant: "fast" from global config.
        
        - name: "complex reasoning task"
          prompt: "Analyze this philosophical argument..."
          response-result-format: "structured analysis with reasoning"
          expected-result: "A thoughtful analysis with logical reasoning"
          validation-rules:
            judge:
              enabled: true
              variant: "reasoning"  # Override: use reasoning run variant instead of fast.
              # Inherits name: "mistral-judge" from global config.
    

Schema-Based Validation

For tasks where exact value matching is too strict but an LLM judge is unnecessary, you can validate responses against a JSON Schema. This is particularly useful for numeric ranges, tolerances, bounding boxes, or any structured output where the set of acceptable answers is better expressed as constraints.

How it works: Instead of comparing the candidate answer to literal expected values, the raw candidate answer is validated directly against the single expected-result JSON Schema without canonicalization, normalization, or type coercion. The schema's $schema is optional and defaults to Draft 2020-12; normalization flags (case-sensitive, ignore-whitespace, trim-lines) are ignored and the mode is mutually exclusive with judge.

[!IMPORTANT] Schema validation performs no type coercion. A plain-text response-result-format always produces a string, so its expected-result schema must accept strings (e.g., type: string with pattern). A schema with type: number will not match the string "10" — use a structured response-result-format such as type: number when you need numeric validation.

To use schema validation, enable it in tasks.yaml:

# Simple range — any number between 9.9 and 10.1 is accepted.
task-config:
  tasks:
    - name: "numeric range"
      prompt: "Pick a number between 9.9 and 10.1"
      response-result-format:
        type: number
      validation-rules:
        schema-validation: true
      expected-result:
        $schema: "https://json-schema.org/draft/2020-12/schema"
        type: number
        minimum: 9.9
        maximum: 10.1

Plain-text responses can still use schema validation — the schema just needs to describe a string:

- name: "hex colour"
  prompt: "Return a six-digit hexadecimal colour starting with #"
  response-result-format: "a six-digit hexadecimal colour starting with #"
  validation-rules:
    schema-validation: true
  expected-result:
    $schema: "https://json-schema.org/draft/2020-12/schema"
    type: string
    pattern: "^#[0-9A-Fa-f]{6}$"

When the model is asked to produce structured JSON, response-result-format and the validator schema are distinct — the former is sent to the model, the latter is evaluator-only:

- name: locate-object
  prompt: >
    Locate the red object in the image and return its normalized
    bounding box.

  response-result-format:
    type: object
    properties:
      x:
        type: number
        minimum: 0
        maximum: 1
      y:
        type: number
        minimum: 0
        maximum: 1
      width:
        type: number
        minimum: 0
        maximum: 1
      height:
        type: number
        minimum: 0
        maximum: 1
    required: [x, y, width, height]
    additionalProperties: false

  expected-result:
    $schema: "https://json-schema.org/draft/2020-12/schema"
    type: object
    properties:
      x:
        type: number
        minimum: 0.31
        maximum: 0.35
      y:
        type: number
        minimum: 0.42
        maximum: 0.46
      width:
        type: number
        minimum: 0.19
        maximum: 0.23
      height:
        type: number
        minimum: 0.27
        maximum: 0.31
    required: [x, y, width, height]
    additionalProperties: false

  validation-rules:
    schema-validation: true

[!TIP] Use anyOf/oneOf/allOf/enum inside the single schema to express alternatives instead of multiple outer expected-result values. A literal object containing $schema without schema-validation: true is still compared via standard value matching — the canonical values are compared exactly (e.g., when the model was asked to generate a schema).

Judge Prompt Customization

MindTrial automatically applies a built-in semantic evaluation template that compares candidate responses against expected answers. For advanced use cases, you can customize the judge prompt template, response format, and acceptance criteria.

Judge prompts can be customized in the validation-rules.judge.prompt section of your tasks.yaml file, either globally in task-config or individually per task.

Customization Fields:

  • template: Custom prompt template for the judge (supports template variables listed below).
  • verdict-format: Expected response format from the judge (plain text instruction or JSON schema).
  • passing-verdicts: Set of verdict values that indicate a passing evaluation, or a single explicit JSON Schema object (identified by a $schema field) for threshold-style criteria (e.g. a minimum score).

[!IMPORTANT]

  • template is independently optional: you can customize it without also overriding verdict-format/passing-verdicts, and vice versa.
  • verdict-format and passing-verdicts must either both be specified together, or both left unset to fall back to the built-in defaults; specifying only one leaves the other ambiguous and is rejected.
  • When passing-verdicts is a set of literal value(s) (the common case), each value must conform to the verdict-format structure.
  • When passing-verdicts is an explicit JSON Schema, it is validated as a standalone schema and matched directly against the judge's raw verdict, without any case/whitespace normalization.

[!TIP] The following template variables are available for judge prompts:

  • {{.OriginalTask.Prompt}}: The original task prompt
  • {{.OriginalTask.ResponseResultFormat}}: Format instruction from the task
  • {{.OriginalTask.ExpectedResults}}: Array of expected answers
  • {{.Candidate.Response}}: The model's response being evaluated
  • {{.Rules.CaseSensitive}}: Boolean case-sensitive validation flag
  • {{.Rules.IgnoreWhitespace}}: Boolean ignore whitespace flag
  • {{.Rules.TrimLines}}: Boolean trim lines flag
  • {{.Verdict.Format}}: The resolved verdict format, rendered as plain text or pretty-printed JSON schema

A sample task from tasks.yaml:

# tasks.yaml
task-config:
  disabled: true
  file-options:
    image-detail: high
  system-prompt:
    enable-for: "text"
    template: |
      Provide the final answer in exactly this format: {{.ResponseResultFormat}}
      Treat every substring enclosed in `<` and `>` as a variable placeholder.
      Substitute only the raw value in place of `<variable name>`, removing the `<` and `>` characters.
      Do not add any extra words, punctuation, quotes, or whitespace beyond what the format string shows.
  validation-rules:
    case-sensitive: false
    ignore-whitespace: false
  tasks:
    - name: "riddle - split words - v1"
      disabled: false
      prompt: |-
        There are four 8-letter words (animals) that have been split into 2-letter pieces.
        Find these four words by putting appropriate pieces back together:

        RR TE KA DG EH AN SQ EL UI OO HE LO AR PE NG OG
      response-result-format: |-
        list of words in alphabetical order separated by ", "
      system-prompt:
        template: "Provide the final answer in exactly this format: {{.ResponseResultFormat}}"
      expected-result: |-
        ANTELOPE, HEDGEHOG, KANGAROO, SQUIRREL
    - name: "visual - shapes - v1"
      prompt: |-
        The attached picture contains various shapes marked by letters.
        It also contains a set of same shapes that have been rotated marked by numbers.
        Your task is to find all matching pairs.
      response-result-format: |-
        <shape number>: <shape letter> pairs separated by ", " and ordered by shape number
      expected-result: |-
        1: G, 2: F, 3: B, 4: A, 5: C, 6: D, 7: E
      validation-rules:
        ignore-whitespace: true
      files:
        - name: "picture"
          uri: "./taskdata/visual-shapes-v1.png"
          type: "image/png"
          options:
            image-detail: original
    - name: "riddle - anagram - v3"
      prompt: |-
        Two words (each individual word is a fruit) have been combined and their letters arranged in alphabetical order forming a single group.
        Find the original words for each of these 2 groups:

        1. AACEEGHPPR
        2. ACEILMNOOPRT
      response-result-format: |-
        1. <word>, <word>
        2. <word>, <word>
        (words in each group must be alphabetically ordered)
      expected-result:
        - |
          1. GRAPE, PEACH
          2. APRICOT, MELON
        - |
          1. GRAPE, PEACH
          2. APRICOT, LEMON
    - name: "chemistry - observable phenomena - v1"
      disabled: true
      prompt: |-
        What are the primary observable results of mixing household vinegar (an aqueous solution of acetic acid, $CH_3COOH$) with baking soda (sodium bicarbonate, $NaHCO_3$)?
      response-result-format: |-
        Provide a bulleted list of the main, directly observable phenomena. Focus on what one would see and hear. Do not include the chemical equation.
      expected-result: |-
        The response must correctly identify the two main observable results of the chemical reaction.
        Crucially, it must mention the production of a gas, described as fizzing, bubbling, or effervescence.
        It should also note that the solid baking soda dissolves or disappears as it reacts with the vinegar.
      validation-rules:
        judge:
          enabled: true
          name: "mistral-judge"
          variant: "reasoning"
          # Uses default judge prompt configuration for semantic evaluation.
    - name: "code quality - custom judge"
      prompt: |-
        Write a Python function that finds the maximum value in a list.
      response-result-format: |-
        complete Python function with proper naming and structure
      expected-result: |-
        A well-written Python function that correctly finds the maximum value with good practices
      validation-rules:
        judge:
          enabled: true
          name: "mistral-judge"
          variant: "reasoning"
          prompt:
            template: |-
              Evaluate this Python code for both correctness and quality:
              {{.Candidate.Response}}

              Criteria: 1) Correctly finds max value, 2) Proper function name/structure, 3) Handles edge cases, 4) Good Python style
            verdict-format:
              type: object
              properties:
                quality_score:
                  type: string
                  enum: ["excellent", "good", "poor"]
              required: ["quality_score"]
              additionalProperties: false
            passing-verdicts:
              - quality_score: "excellent"
              - quality_score: "good"
    - name: "essay quality - score threshold"
      prompt: |-
        Write a short essay explaining photosynthesis for a middle school audience.
      response-result-format: |-
        a clear, well-organized short essay
      expected-result: |-
        A clear and accurate explanation of photosynthesis appropriate for the target audience
      validation-rules:
        judge:
          enabled: true
          name: "mistral-judge"
          variant: "reasoning"
          prompt:
            template: |-
              Score this essay from 0 to 100 for clarity, accuracy, and audience appropriateness:
              {{.Candidate.Response}}
            verdict-format:
              type: object
              properties:
                score:
                  type: integer
                  minimum: 0
                  maximum: 100
              required: ["score"]
              additionalProperties: false
            passing-verdicts:
              # An explicit JSON Schema (identified by "$schema") is matched directly against
              # the judge's raw verdict, enabling threshold-style criteria.
              $schema: "https://json-schema.org/draft/2020-12/schema"
              type: object
              properties:
                score:
                  type: integer
                  exclusiveMinimum: 80
              required: ["score"]
    - name: "structured response - log parsing"
      prompt: |-
        Parse the following log lines and extract the timestamp, log level, and message for each. If a user ID is present, extract that as well.
        Log lines:
        [2025-09-14 10:30:00] INFO: User 'admin' logged in successfully.
        [2025-09-14 10:31:15] WARN: System memory usage is high.
      response-result-format:
        type: array
        items:
          type: object
          additionalProperties: false
          properties:
            timestamp:
              type: string
              format: "date-time"
            level:
              type: string
              enum: ["INFO", "WARN", "ERROR"]
            message:
              type: string
            user_id:
              type: string
          required: ["timestamp", "level", "message", "user_id"]
      expected-result:
        - - timestamp: "2025-09-14T10:30:00Z"
            level: "INFO"
            message: "User 'admin' logged in successfully."
            user_id: "admin"
          - timestamp: "2025-09-14T10:31:15Z"
            level: "WARN"
            message: "System memory usage is high."
            user_id: ""

Tools

MindTrial supports tool use for tasks, allowing AI models to execute external tools during task solving. Tools are executed in sandboxed Docker containers with resource limits and network isolation.

Tool Definitions

Tools must be defined in config.yaml under the tools section. Each tool defines how to execute a specific capability:

  • name: A unique name for the tool.
  • image: Docker image to use for the tool execution.
  • description: A detailed description of what the tool does and how to use it. This description is provided to the LLM to help it understand when and how to use the tool. Be specific and avoid ambiguity to help the LLM choose the correct tool and provide appropriate parameters.
  • parameters: JSON schema defining the tool's input parameters. The LLM will generate the actual parameter values based on this schema. Provide comprehensive descriptions that explain parameter purpose and format.
  • parameter-files: Mapping of parameter names to container file paths where argument values should be written. Argument values are converted to strings, non-string values are marshaled to JSON. The tool's command should read these files as needed.
  • auxiliary-dir: Directory path inside the container where task files with local access will be automatically mounted. If specified, files with local access attached to the task will be mounted to this directory using each file's unique reference name exactly as provided. Files in this directory are reset between tool calls.
  • shared-dir: Directory path inside the container that persists across all tool calls within a single task. If specified, files created in this directory will be available for any subsequent tool calls but will be removed when the task completes.
  • command: Command to run inside the container. The standard output of the command execution is captured and passed back to the LLM as is.
  • env: Environment variables to set in the container.
  • dependencies: Task services the tool can access, each with a service name and optional env templates.

[!IMPORTANT] Tool use requires Docker to be installed and running on the system. Tools are executed in isolated containers with no network access, unless they depend on task services.

Example tool definition in config.yaml:

config:
  tools:
    - name: python-code-executor
      image: python:latest
      description: |
        Executes Python 3 code in a secure, sandboxed environment to perform calculations, data manipulation, or algorithmic tasks.
        IMPORTANT:
        - Only the Python standard library is available. No third-party packages (like pandas or numpy) can be imported.
        - The environment has no network access.
        - Task files made available to this tool are mounted under /app/data/ using their [file: filename] names.
        - Use standard file operations like open('/app/data/filename', 'r') to read attached files, where 'filename' matches the name shown in [file: filename] references.
        - A persistent shared directory is available at /app/shared/ that persists across ALL tool calls within the same task (regardless of which tool is being called). Files created in this directory will be available in any subsequent tool call.
        - Any files or changes outside of /app/shared/ are ephemeral and will be reset between tool calls.
        - The code must print its final result to standard output to be returned.
      parameters:
        type: object
        properties:
          code:
            type: string
            description: "A string containing a self-contained Python 3 script. The script must use the `print()` function to return a final result. Example: `print(sum([i for i in range(101) if i % 2 == 0]))`. To read attached files, use open('/app/data/filename', 'r') where 'filename' matches what appears in [file: filename] references."
        required:
          - code
        additionalProperties: false
      parameter-files:
        code: /app/main.py
      auxiliary-dir: /app/data
      shared-dir: /app/shared
      command:
        - python
        - /app/main.py
      env:
        PYTHONIOENCODING: "UTF-8"
        PYTHONUNBUFFERED: "1"
        PYTHONHASHSEED: "847629"
Tool Selection

You can configure tool selection globally for all tasks in the task-config section, and override it for individual tasks if needed. Tools must be defined in config.yaml first.

  • tool-selector: Configuration for tool availability during task execution.
    • disabled: If true, no tools are available for tasks (default: false).
    • tools: List of tools to make available, with per-tool limits.
      • name: Name of the tool as defined in config.yaml.
      • disabled: If true, this tool is not available (default: false).
      • max-calls: Maximum number of times this tool can be called per task (optional).
      • timeout: Maximum execution time per tool call (e.g., 60s, optional).
      • max-memory-mb: Maximum memory usage in MB per tool call (optional).
      • cpu-percent: Maximum CPU usage as percentage per tool call (optional).
    • service-inputs: Inputs for the task services that the task uses, keyed by service name and then by input name (optional). Inputs set in task-config apply to every task, and a task can override single inputs. Inputs only take effect for tasks that use the service.

Example tool configuration in tasks.yaml:

task-config:
  tool-selector:
    disabled: false
    tools:
      - name: python-code-executor
        disabled: false
        max-calls: 10
        timeout: 60s
        max-memory-mb: 512
        cpu-percent: 25
  tasks:
    - name: "math calculation"
      prompt: "Calculate the sum of even numbers from 1 to 100."
      response-result-format: "single number"
      expected-result: "2550"
      # Inherits global tool-selector configuration.
    - name: "simple math"
      prompt: "What is 2 + 2?"
      response-result-format: "single number"
      expected-result: "4"
      tool-selector:
        tools:
          - name: python-code-executor
            disabled: true  # Selectively disable tool for this simple task.
Conversation Turn Limit

You can set a maximum number of conversation turns per task to act as a safety net against infinite conversation loops (e.g., when a model repeatedly requests exhausted tools). The limit can be configured globally in the task-config section, and overridden for individual tasks if needed. A value of 0 means unlimited.

  • max-turns: Maximum number of conversation turns allowed per task (default: 0, unlimited).

Example configuration in tasks.yaml:

task-config:
  max-turns: 100  # Default limit for all tasks.
  tasks:
    - name: "trivia - geography - Asia"
      prompt: "What is the capital of Japan?"
      response-result-format: "city name"
      expected-result: "Tokyo"
      max-turns: 200  # Override: allow more turns for this task.
    - name: "trivia - geography"
      prompt: "What is the capital of Australia?"
      response-result-format: "city name"
      expected-result: "Canberra"
      # Inherits the global limit of 100 turns.

Task Services

A task service is a Docker container that keeps state while the model works on a task, such as a simulated environment, a database, or a web shop. Services are used together with:

  • tools, which let the model read and change the service state, and
  • custom validators, which can check the final service state after the model answers.

Services and custom validators are independent of each other. Tools that use services work with any validation method (exact match, schema, judge, or custom validator), and a custom validator does not need any services.

[!IMPORTANT] Task services require Docker Engine 25.0 or newer (Engine API 1.44+). Before an evaluation that uses services starts, MindTrial checks the Docker daemon and stops with an error if it is too old.

Services are defined in config.yaml under the services section:

  • name: A unique name for the service, used in dependencies and service-inputs.
  • image: Docker image used to run the service.
  • command: Command overriding the image's default command (optional).
  • env: Environment variables set for every instance of the service (optional).
  • endpoint: Where tools and validators connect to the service.
    • port: Container port on which the service accepts connections.
    • scheme: URL scheme, http (default) or https.
  • input-env: Inputs that tasks can set, mapping each input name to the environment variable that receives its value in the service container (optional). An input must not set a different value for a variable that env already defines.
  • healthcheck: Command that checks whether the service is ready, run inside the service container without a shell (optional). If set, the service is ready once the command succeeds. If not set, the service is ready as soon as its container is running.
  • startup-timeout: Maximum time a service with a healthcheck may take to become ready (optional, default 15s). The check is repeated until it succeeds or this time runs out, and a single check may also take up to this long.
  • max-memory-mb: Maximum memory available to the service container in MB (optional).
  • cpu-percent: Maximum CPU usage as a percentage of total host CPU (optional).

A tool or custom validator gets access to a service by listing it in dependencies:

  • dependencies: List of services the tool or validator can access.
    • service: Name of the service.
    • env: Environment variables set in the tool or validator container (optional). Each value is a template filled in from the running service, for example CART_URL: "{{ .Endpoint }}".

A task sets service inputs with service-inputs in its tool selector:

tool-selector:
  service-inputs:
    cart-service:        # service name
      customer_id: "42"  # input name declared in the service's input-env
  • Input values must be strings, numbers, or booleans.
  • String values are templates, so a task can derive an input from the evaluation seed, e.g. "{{ hash .Evaluation.Seed .Task.Name }}". MindTrial generates a new evaluation seed for each evaluation, writes it to the log, and records it in every result (Evaluation.Seed in the JSON output); pass the same seed with --evaluation-seed to get the same inputs again.
  • service-inputs set in task-config apply to every task, and a task can override single inputs. Inputs only take effect for tasks that use the service.
  • When an input is not set, the service uses its own default.

How services run:

  • Which services start: The services that the task's enabled tools and its custom validator depend on. Tasks that need no services start none.
  • When services start: Before the model receives the prompt. Every service must be ready within its startup-timeout (15 seconds by default); otherwise the attempt fails and its services are removed.
  • One set of services per attempt: An attempt is one try of one task by one run configuration. Each attempt gets its own service instances, so runs never share state. A retry is a new attempt with fresh services started from the same inputs. If a service does not derive its initial state from its inputs (for example, from a seed input), a retry may start from a different state.
  • Validation: After a successful attempt, a custom validator that depends on a service connects to the same instance that the model's tools used. The services are removed after validation.
  • Network isolation: Each service has its own internal Docker network without internet access. A tool or validator joins only the networks of the services in its dependencies, and one without dependencies has no network access. Services cannot reach each other.
  • Service failure: A service that stops during an attempt is not restarted, because that would silently reset its state. The next tool call or validation that needs it fails, and the task result is an error.

Custom Validators

A custom validator is a Docker container that you provide to decide whether the model's answer is correct. Use it when an answer can only be checked by running code, for example to run generated code against hidden tests or to check the final state of a simulated environment.

A custom validator does not need task services. Without dependencies, it runs in an isolated container with no network access. With dependencies, it can inspect the same service instances that the model's tools used.

How it works: After a successful model attempt, MindTrial runs the validator container once. It passes the task and the model's answer to the validator through templated command arguments, environment variables, and files. The validator prints its verdict as JSON on standard output.

[!IMPORTANT] Custom validators are trusted: they are part of your evaluation setup, not tools for the model, and they decide task outcomes. Only use validator images you control.

Validators are defined in config.yaml under the validators section:

  • name: A unique name for the validator, used by custom-validator in task validation rules.
  • image: Docker image used to run the validator.
  • command: Command overriding the image's default command (optional). Each argument is a template and is passed directly to Docker, without a shell.
  • env: Environment variables set in the validator container (optional). Each value is a template.
  • template-files: Files created for each validation and mounted read-only into the validator container (optional).
    • path: Absolute path of the file inside the container. Each path must be unique.
    • template: Template that produces the file content.
  • dependencies: Task services the validator can access (optional).
  • timeout: Maximum duration of one validator run (e.g., 60s). If not set, there is no timeout.
  • max-memory-mb: Maximum memory available to the validator container in MB (optional).
  • cpu-percent: Maximum CPU usage as a percentage of total host CPU (optional).

[!TIP] Pass the model's answer through template-files. Use command and env only for short values you control, such as names, seeds, and flags. The model's answer can be of any size and may contain characters that command arguments and environment variables cannot carry; if the validator cannot start, the task result is an error rather than a failed answer.

A task selects a validator with the custom-validator validation rule:

  • response-result-format is still required: it tells the model how to write its final answer.
  • expected-result is optional. When set, it is passed to the validator unchanged as reference data and does not need to match response-result-format.
  • case-sensitive, ignore-whitespace, and trim-lines are passed to the validator, which decides whether to use them. MindTrial does not change the answer.

Example of a validator that runs hidden tests without any services:

# config.yaml
config:
  validators:
    - name: hidden-tests
      image: example/code-checker:latest
      command: [check, --solution, /input/solution.py]
      template-files:
        - path: /input/solution.py
          template: "{{ .Candidate.Response }}"
      timeout: 60s
# tasks.yaml
task-config:
  tasks:
    - name: fizzbuzz
      prompt: "Write a Python function fizzbuzz(n) that returns the FizzBuzz sequence from 1 to n."
      response-result-format: "complete Python source code"
      validation-rules:
        custom-validator: hidden-tests

The validator must exit with code 0 and print exactly one JSON object to standard output, for example:

{"correct": false, "title": "2 of 5 tests failed", "explanation": "fizzbuzz(15) returned '15' instead of 'FizzBuzz'."}

The object must match this schema:

{
  "type": "object",
  "properties": {
    "correct": {"type": "boolean"},
    "title": {"type": "string", "pattern": "\\S"},
    "explanation": {"type": "string", "pattern": "\\S"}
  },
  "required": ["correct", "title", "explanation"],
  "additionalProperties": false
}
  • correct: true marks the answer as correct, and correct: false marks it as failed.
  • title and explanation must contain non-whitespace text.
  • Anything else makes the task result an error, not a failed answer: a non-zero exit code, a timeout, a Docker error, missing or unknown fields, or any output after the JSON object.

Results checked by a custom validator record custom as their validation method. Failed answers are shown exactly as the model returned them, without a comparison against expected-result. Validator runs are not counted as model tool calls.

Templates

Service inputs, dependency environment variables, and custom validator settings use Go template syntax. MindTrial checks template syntax before any task runs, and using a field that does not exist is an error.

Each kind of template has its own fields:

TemplateAvailable fields
service-inputs string values.Evaluation.Seed, .Task.Name, .Provider.Name, .Run.Name
dependencies[].env values.Name, .Host, .Port, .Endpoint
Validator command, env, and template-files.OriginalTask.*, .Candidate.Response, .Rules.*, .Evaluation.Seed, .Task.Name, .Provider.Name, .Run.Name

Fields:

  • {{ .Evaluation.Seed }}: Evaluation seed shared by all providers, runs, tasks, and attempts of one evaluation.
  • {{ .Task.Name }}, {{ .Provider.Name }}, {{ .Run.Name }}: Names of the task, the provider, and the run configuration.
  • {{ .Name }}: Name of the service.
  • {{ .Host }}: Network host name of the service. MindTrial generates host names, so use .Host or .Endpoint instead of the service name.
  • {{ .Port }}: Endpoint port of the service.
  • {{ .Endpoint }}: Endpoint URL of the service, e.g. http://<host>:8080.
  • {{ .OriginalTask.Prompt }}: The task prompt.
  • {{ .OriginalTask.ResponseResultFormat }}: The task's response-result-format.
  • {{ .OriginalTask.ExpectedResults }}: List of the task's expected results. It is an empty list (not null) when the task has no expected-result.
  • {{ .Candidate.Response }}: The model's final answer. For structured response formats, this is a structured value.
  • {{ .Rules.CaseSensitive }}, {{ .Rules.IgnoreWhitespace }}, {{ .Rules.TrimLines }}: The task's validation flags.

The validator fields use the same names as judge prompt templates.

Helper functions:

  • hash: Returns a stable unsigned 64-bit number derived from all of its arguments, e.g. {{ hash .Evaluation.Seed .Task.Name }}.
  • json: Encodes its argument as JSON, e.g. {{ json .OriginalTask.ExpectedResults }}.

[!IMPORTANT] Go prints structured values in its own format, which is not JSON. Use json to pass structured data, e.g. {{ json .Candidate.Response }}.

Example: Stateful Shopping Cart

In this example, the model uses the cart tool to change a shopping cart held by the cart-service service. After the model answers, the cart-state validator checks the same cart:

# config.yaml
config:
  services:
    - name: cart-service
      image: example/cart-service:latest
      command: [cart-server]
      env:
        CART_CURRENCY: CAD
      endpoint:
        port: 8080
      input-env:
        customer_id: CART_CUSTOMER_ID
      healthcheck: [cart-server, --check]

  tools:
    - name: cart
      image: example/cart-client:latest
      command: [cart-client, --request, /input/request.json]
      description: >
        Modify the current customer's shopping cart by adding or removing items.
      parameters:
        type: object
        properties:
          request:
            type: object
            properties:
              action:
                type: string
                enum: [ADD, REMOVE]
                description: Operation to perform on the cart.
              item:
                type: string
                enum: [apple, orange, bread]
                description: Item to add or remove.
              quantity:
                type: integer
                minimum: 1
                description: Number of items to add or remove.
            required: [action, item, quantity]
            additionalProperties: false
        required: [request]
        additionalProperties: false
      parameter-files:
        request: /input/request.json
      dependencies:
        - service: cart-service
          env:
            CART_URL: "{{ .Endpoint }}"

  validators:
    - name: cart-state
      image: example/cart-validator:latest
      command: [cart-validator, --expected, /input/expected.json]
      template-files:
        - path: /input/expected.json
          template: "{{ json .OriginalTask.ExpectedResults }}"
      timeout: 60s
      dependencies:
        - service: cart-service
          env:
            CART_URL: "{{ .Endpoint }}"

CART_CURRENCY is a fixed service setting. customer_id is an input that each task can set; the service receives it as CART_CUSTOMER_ID. The task below derives the customer ID from the evaluation seed, so all providers and runs in one evaluation start with the same customer. Passing the same --evaluation-seed again gives the same customer:

# tasks.yaml
task-config:
  tasks:
    - name: update-shopping-cart
      prompt: >
        Use the cart tool to add two apples and one loaf of bread.
      response-result-format: "short summary of the final cart contents"
      expected-result:
        items:
          apple: 2
          bread: 1
      tool-selector:
        tools:
          - name: cart
        service-inputs:
          cart-service:
            customer_id: "{{ hash .Evaluation.Seed .Task.Name }}"
      validation-rules:
        custom-validator: cart-state

The expected-result object is reference data for the validator; it does not have to match the plain-text response-result-format. The validator reads it from /input/expected.json as [{"items":{"apple":2,"bread":1}}] and compares it with the cart state it gets from CART_URL.

Command Reference

mindtrial [options] [command]

Commands:
  run                       Start the trials
  merge-results             Merge results from multiple runs
  stats                     Compute derived statistics from result files
  help                      Show help
  version                   Show version

Options:
  --config string           Configuration file path (default: config.yaml)
  --tasks string            Task definitions file path
  --output-dir string       Results output directory
  --output-basename string  Base filename for results; replace if exists; blank = stdout
  --html                    Generate HTML output (default: true)
  --csv                     Generate CSV output (default: false)
  --json                    Generate JSON output (default: false)
  --input string            Input result file path for merge-results/stats; can be specified multiple times
  --log string              Log file path; append if exists; blank = stdout
  --verbose                 Enable detailed logging
  --debug                   Enable low-level debug logging (implies --verbose)
  --interactive             Enable interactive interface for run configuration, and real-time progress monitoring (default: false)
  --evaluation-seed string  Seed for reproducible evaluation behavior (e.g., derived task service inputs); generated for each evaluation when omitted
  --group-by string         Comma-separated stats grouping dimensions: provider, run, model, suite, category, difficulty, tag (default: provider,run)
  --stats-format string     Stats output format: text, csv, json, or jsonl (default: text)
  --provider string         Filter stats to this provider; can be specified multiple times
  --run string              Filter stats to this run configuration; can be specified multiple times
  --model string            Filter stats to this model; can be specified multiple times
  --suite string            Filter stats to this task suite; can be specified multiple times
  --category string         Filter stats to this task category; can be specified multiple times
  --difficulty string       Filter stats to this task difficulty; can be specified multiple times
  --status string           Filter stats to this result status (passed, failed, error, skipped); can be specified multiple times
  --tag string              Filter stats to results tagged with this value; can be specified multiple times
  --tag-mode string         How multiple --tag filters combine for stats: all or any (default: all)

Contributing

Contributions are welcome! Please review our CONTRIBUTING.md guidelines for more details.

Getting the Source Code

Clone the repository and install dependencies:

git clone https://github.com/petmal/mindtrial.git
cd mindtrial
go mod download

Running Tests

Execute the unit tests with:

go test -tags=test -race -v ./...

Project Details

/
├── cmd/
│   └── mindtrial/       # Command-line interface and main entry point
│       └── tui/         # Terminal-based UI and interactive mode functionality
├── config/              # Data models and management for configuration and task definitions
├── formatters/          # Output formatting for results
├── pkg/                 # Shared packages and utilities
├── providers/           # AI model service provider connectors
│   ├── execution/       # Provider run execution utilities and coordination
│   └── tools/           # Execution engine for external tools used by models
├── runners/             # Task execution and result aggregation
├── stats/               # Derived statistics (filtering, grouping, aggregation) for the stats command
├── taskdata/            # Auxiliary files referenced by tasks in tasks.yaml
├── validators/          # Result validation logic
└── version/             # Application metadata

License

This project is licensed under the Mozilla Public License 2.0 - see the LICENSE file for details.

ai-benchmark
ai-evaluation-tools
ai-model-comparison
ai-tool
anthropic
artificial-intelligence-projects
deepseek
google-gemini-ai
grok-ai
language-models-ai
llm-benchmarking
llm-comparison
llm-evaluation-framework
mistral-ai
moonshot-ai
openai
openrouter
opensource
qwen
xai

Significant stargazers

mitchell

88 followers · starred May 2026

petmal/MindTrial

MindTrial: Evaluate and compare AI language models (LLMs) on text-based tasks with optional file/image attachments and tool use. Supports multiple providers (OpenAI, Google, Anthropic, DeepSeek, Mistral AI, xAI, Alibaba, Moonshot AI, OpenRouter), custom tasks in YAML, and HTML/CSV/JSON reports.

Go

21

129 commits

updated Oct 2, 2026

See the code

See what people are saying

SourceMessageScoreDate

Benchmark notes: Sonnet 5.5 jumps from 72 to 94/98; Opus 5.5 reaches 96/98 with much less request time (r/ClaudeAI)

I maintain [**MindTrial**](https://github.com/petmal/MindTrial) and tested **Sonnet 5.5** and **Opus 5.5** on the same 98-task suite as their predecessors: 39 text tasks and 59 visual tasks, with Python/scientific libraries available and a 10-call limit per task. All four Claude runs below use the…

9

Oct 4, 2026

README

MindTrial

Build License: MPL 2.0 Go Version Go Reference

MindTrial lets you test a single AI language model (LLM) or evaluate multiple models side-by-side. It supports providers like OpenAI, Google, Anthropic, DeepSeek, Mistral AI, xAI, Alibaba, Moonshot AI, and OpenRouter. You can create your own custom tasks with text prompts, plain text or structured JSON response formats, optional file attachments, and tool use for enhanced capabilities; validate responses through exact value matching, an LLM judge for semantic evaluation, or trusted Docker-backed custom validators; and get results in easy-to-read HTML, CSV, and JSON formats.

Quick Start Guide

  1. Install the tool:

    go install github.com/petmal/mindtrial/cmd/mindtrial@latest
    
  2. Run with default settings:

    mindtrial run
    

Prerequisites

  • Go 1.26
  • Docker (for tool execution; Docker Engine 25.0+ for task services)
  • API keys from your chosen AI providers

Key Features

  • Compare multiple AI models at once
  • Create custom evaluation tasks using simple YAML files
  • Attach files or images to prompts for visual tasks
  • Enable tool use for tasks with secure sandboxed execution
  • Use LLM judges for semantic validation of complex and creative tasks
  • Evaluate stateful tool use with task-scoped Docker services and trusted custom validators
  • Get results in HTML, CSV, and JSON formats
  • Merge and compare results from multiple runs
  • Easy to extend with new AI models
  • Smart rate limiting to prevent API overload
  • Interactive mode with terminal-based UI

Basic Usage

  1. Display available commands and options:

    mindtrial help
    
  2. Run with custom configuration and output options:

    mindtrial --config="custom-config.yaml" --tasks="custom-tasks.yaml" --output-dir="./results" --output-basename="custom-tasks-results" run
    
  3. Run with specific output formats (CSV only, no HTML):

    mindtrial --csv=true --html=false run
    
  4. Run in interactive mode to select models and tasks before starting:

    mindtrial --interactive run
    
  5. Merge results from multiple runs into a single output:

    mindtrial --input="results-1.json" --input="results-2.json" --html=true --csv=true --output-basename="merged" merge-results
    
  6. Compute derived statistics (pass rate, durations, token usage, and more) grouped by provider and run:

    mindtrial --input="results.json" stats
    

Merging Results

The merge-results command combines results from multiple trial runs into a single output. Input files are specified with the --input flag (can be repeated). Currently, only JSON is supported as the input format. Use the --json=true flag during trial runs to generate JSON output files that can later be merged. The merged output can be generated in any of the supported formats (HTML, CSV, JSON) using the corresponding flags.

[!TIP] You can also use merge-results with a single input file to convert between formats. For example, if you store results in JSON, you can convert them to HTML or CSV at any time:

mindtrial --input="results.json" --html=true --csv=true merge-results

[!TIP] If some results failed due to transient errors (e.g., network timeouts), you can re-run only the failed tasks and merge the new results into the original set. Because merge-results uses a last-in-wins strategy for duplicate entries (same provider, run, and task), the corrected results will replace the failed ones.

Computing Statistics

The stats command computes derived statistics (pass rate, accuracy, error rate, duration, token usage, tool calls, error diagnostics, and estimated candidate cost) over one or more result files, filtered and grouped by task/run metadata. Like merge-results, input files are specified with the --input flag (can be repeated); when multiple files are given, they are merged first using the same last-in-wins strategy before stats are computed. This is derived analytical output, not a canonical result format, so it is only written to stdout and is not persisted back into result artifacts. Progress messages are written to stderr, so stdout stays safe to redirect straight into a file or parser for every --stats-format.

mindtrial --input="results.json" --group-by="provider,run,model" --stats-format="csv" stats
  • --group-by: Comma-separated grouping dimensions: provider, run, model, suite, category, difficulty, tag (default: provider,run); each dimension may appear at most once. Grouping by tag is exploded: a task tagged with multiple tags contributes to each of those tag groups, so tag groups overlap and are not additive; duplicate tags on the same task count once. Results missing a value for a grouping dimension are grouped under (unspecified) rather than dropped.
  • --stats-format: Output format: text, csv, json, or jsonl (default: text).
  • --provider, --run, --model, --suite, --category, --difficulty, --status: Restrict stats to matching results; each can be specified multiple times (combined with OR). --status accepts passed, failed, error, or skipped. Pass (unspecified) to match results missing a value for that field.
  • --tag: Restrict stats to results carrying the given tag; can be specified multiple times. Combined according to --tag-mode. Pass (unspecified) to match untagged results.
  • --tag-mode: How multiple --tag filters combine: all (default, every tag must be present) or any (at least one tag must be present).

[!NOTE] Token, tool-call, and duration metrics reflect only the candidate answer and any subsequent error (not judge/validation usage), matching the HTML report's dynamic summary. Median*/Stddev* metrics require at least two contributing samples; otherwise they are omitted.

TotalInputTokens is normalized: providers that count cache reads/writes separately from input (currently Anthropic) have them added, and providers that already include them do not. TotalOutputTokens is normalized the same way: providers that count reasoning separately from output (currently Google and xAI) have it added, and providers that already include it do not. TotalReasoningTokens, TotalCacheReadTokens, and TotalCacheWriteTokens sum only the counts providers actually reported. EstimatedCandidateCost/CandidateCostCurrency are priced with the candidate run's own prices and never include judge/validation usage.

Example: given a result file with these three tasks:

ProviderRunStatusTags
openaigpt-4Passedvisual, spatial
openaigpt-4Failedvisual
anthropicclaudePassedtext

Grouping by the default provider,run treats each provider/run combination as one group:

$ mindtrial --input="results.json" --stats-format="csv" stats
provider,run,Count,Passed,Failed,...
anthropic,claude,1,1,0,...
openai,gpt-4,2,1,1,...

Grouping by tag instead explodes each task into every tag it carries, so tag groups overlap rather than partition the input (the first task counts toward both visual and spatial):

$ mindtrial --input="results.json" --group-by="tag" --stats-format="csv" stats
tag,Count,Passed,Failed,...
spatial,1,1,0,...
text,1,1,0,...
visual,2,1,1,...

Configuration Guide

MindTrial uses two simple YAML files to control everything:

1. config.yaml - Application Settings

Controls how MindTrial operates, including:

  • Where to save results
  • Which AI models to use
  • API settings and rate limits

2. tasks.yaml - Task Definitions

Defines what you want to evaluate, including:

  • Questions/prompts for the AI
  • Expected answers
  • Response format rules

[!TIP] New to MindTrial? Start with the example files provided and modify them for your needs.

[!TIP] Use interactive mode with the --interactive flag to select model configurations and tasks before running, without having to edit configuration files.

config.yaml

This file defines the tool's settings and target model configurations evaluated during the trial run. The main sections include:

  • output-dir: Path to the directory where results will be saved.
  • task-source: Path to the file with definitions of tasks to run.
  • providers: List of providers (i.e. target LLM configurations) to execute tasks during the trial run.
    • name: Name of the LLM provider (e.g. openai).
    • client-config: Configuration for this provider's client (e.g. API key).
    • max-parallel-requests-per-minute: Enables parallel execution of runs within this provider and limits the aggregate number of API requests per minute across all runs. Set to 0 or omit for sequential execution (default).
    • runs: List of runs (i.e. model configurations) for this provider. Unless disabled, all configurations will be trialed.
      • name: A unique display-friendly name to be shown in the results.
      • model: Model name must be exactly as defined by the backend service's API (e.g. gpt-4o-mini).
      • disable-structured-output: Disable structured JSON responses for this run and force plain-text answers. When enabled, MindTrial:
        • treats the model's entire response as the final answer (title and explanation are filled with placeholders)
        • forces the model to use plain-text response mode
        • skips tasks that require schema-based (response-result-format) JSON outputs
      • text-only: Skip tasks that require native file input to the model API. When enabled, only tasks without native file input will be executed. Tasks with files that are only available to local tools (access: [local]) are still executed. This is useful for text-only models that cannot process images or other files natively.

[!TIP] Use text-only for models that do not support vision capabilities, such as text-only language models hosted on platforms like OpenRouter.

[!TIP] If the model can output JSON as plain text but cannot follow a provider-enforced schema, prefer text-response-format. Use disable-structured-output only when the model cannot reliably output JSON at all.

[!IMPORTANT] For models that accept an explicit response-format parameter (e.g. OpenRouter), ensure the response format is unset or set to plain text when disable-structured-output is enabled; otherwise the run will fail.

[!IMPORTANT] The disable-structured-output flag cannot be used in judge configurations, as judges require structured responses for evaluation.

[!NOTE] The OpenAI provider routes GPT-5 and newer model families through the Responses API and currently relies on stored response state (previous_response_id) for multi-turn and tool-calling flows. As a result, OpenAI Zero Data Retention (ZDR) is not currently supported for those models. Legacy OpenAI models that still use the Chat Completions API are unaffected.

[!IMPORTANT] All provider names must match exactly:

  • openai: OpenAI GPT models
  • google: Google Gemini models
  • anthropic: Anthropic Claude models
  • deepseek: DeepSeek open-source models
  • mistralai: Mistral AI models
  • xai: xAI (Grok) models
  • alibaba: Alibaba (Qwen) models
  • moonshotai: Moonshot AI (Kimi) models
  • openrouter: OpenRouter-hosted models

[!TIP] Instead of a literal value, client-config.api-key can reference an environment variable using the {{.Env.NAME}} placeholder (e.g. "{{.Env.OPENAI_API_KEY}}"), so secrets don't need to be committed to the config file. If api-key is omitted entirely (or left blank), each provider falls back to its own default environment variable:

  • openai: OPENAI_API_KEY
  • google: GOOGLE_API_KEY
  • anthropic: ANTHROPIC_API_KEY
  • deepseek: DEEPSEEK_API_KEY
  • mistralai: MISTRAL_API_KEY
  • xai: XAI_API_KEY
  • alibaba: DASHSCOPE_API_KEY
  • moonshotai: MOONSHOT_API_KEY
  • openrouter: OPENROUTER_API_KEY

This fallback also applies to judge provider configurations under judges[].provider.client-config.

[!NOTE] Anthropic and DeepSeek providers support configurable request timeout in the client-config section:

  • request-timeout: Sets the timeout duration for API requests (i.e. thinking).

Alibaba and Moonshot AI providers support endpoint configuration in the client-config section:

  • endpoint: Specifies the network endpoint URL for the API. If not specified, defaults are:
    • Alibaba: Singapore endpoint (https://dashscope-intl.aliyuncs.com/compatible-mode/v1) for better international access. For China mainland, use https://dashscope.aliyuncs.com/compatible-mode/v1.
    • Moonshot AI: Public API endpoint (https://api.moonshot.ai/v1).

[!NOTE] Some models support additional model-specific runtime configuration parameters. These can be provided in the model-parameters section of the run configuration.

Currently supported parameters for OpenAI models include:

  • text-response-format: If true, use plain-text response format (less reliable) for compatibility with models that do not support JSON.
  • reasoning-effort: Controls effort on reasoning for reasoning models. (values: none, minimal, low, medium, high, xhigh, max). The max level is supported on GPT-5.6 and later reasoning models. Legacy models may not support all values.
  • reasoning-context: Controls which prior reasoning items are reused across conversation turns (persisted reasoning). (values: auto, current_turn, all_turns). Supported on GPT-5.6 and later reasoning models. auto uses the model's default; current_turn keeps reasoning from the active turn only, without rendering earlier turns' reasoning into the next call; all_turns renders compatible reasoning items from earlier turns into the next call as well, which only has an effect when prior response items are available (e.g. via previous_response_id, which MindTrial already relies on for its multi-turn tool-calling loop).
  • reasoning-mode: Selects an alternate reasoning execution mode. (values: pro). Pro mode performs additional model work before returning a single final answer, increasing latency and token usage; use selectively for demanding tasks. Supported on GPT-5.6 and later models.
  • verbosity: Controls how many output tokens are generated. (values: low, medium, high). May not be supported by legacy models.
  • temperature: Controls randomness/creativity of responses (range: 0.0 to 2.0, default: 1.0). Lower values produce more focused and deterministic outputs.
  • top-p: Controls diversity via nucleus sampling (range: 0.0 to 1.0, default: 1.0). Lower values produce more focused outputs.
  • presence-penalty: Penalizes new tokens based on their presence in text so far (range: -2.0 to 2.0, default: 0.0). Positive values encourage model to use new tokens.
  • frequency-penalty: Penalizes new tokens based on their frequency in text so far (range: -2.0 to 2.0, default: 0.0). Positive values encourage model to use less frequent tokens.
  • max-completion-tokens: Controls the maximum number of tokens available to the model for generating a response.

Currently supported parameters for OpenRouter models include:

  • response-format: Controls response format (values: json-schema, json-object, text, default: json-schema). MindTrial adjusts its parsing logic based on this value, so always use this typed parameter rather than passing response_format directly (see note below about extra parameters).
  • temperature: Controls randomness/creativity of responses (range: 0.0 to 2.0, default: 1.0). Lower values produce more focused and deterministic outputs.
  • top-p: Controls diversity via nucleus sampling (range: 0.0 to 1.0, default: 1.0). Lower values produce more focused outputs.
  • top-k: Limits token selection to top K candidates (range: 0 or above, default: 0). Value 0 disables this setting.
  • min-p: Filters tokens below minimum probability relative to most likely token (range: 0.0 to 1.0, default: 0.0).
  • top-a: Considers tokens with sufficiently high probability relative to most likely token (range: 0.0 to 1.0, default: 0.0).
  • presence-penalty: Penalizes token repetition based on presence in input (range: -2.0 to 2.0, default: 0.0).
  • frequency-penalty: Penalizes token repetition based on frequency in input (range: -2.0 to 2.0, default: 0.0).
  • repetition-penalty: Reduces repetition of tokens from input (range: 0.0 to 2.0, default: 1.0).
  • max-tokens: Sets upper limit on generated tokens (range: 1 or above, limited by model context).
  • max-completion-tokens: Sets the modern upper limit on generated tokens. It cannot be combined with max-tokens.
  • reasoning-effort: Controls reasoning depth (values: none, minimal, low, medium, high, xhigh, max). It maps to OpenRouter's native top-level reasoning_effort field.
  • seed: Enables deterministic sampling when supported.
  • parallel-tool-calls: Enables parallel function calling during tool use (default: true).
  • verbosity: Controls response verbosity (values: low, medium, high, default: medium).
  • server-tools: Injects provider-managed server-side tools into every request for this run. Each entry requires a type (the tool identifier, e.g. openrouter:fusion) and an optional parameters map whose fields are tool-specific. Server tools are appended to the request's tools array alongside any local Docker-based tools. Use an extra parameter (e.g. tool_choice: required) to control invocation behavior.

Any additional parameters not listed above can be specified directly in model-parameters and will be forwarded to the OpenRouter API. Use for provider-specific or OpenRouter-specific parameters. Prefer typed parameters where they exist. Note: if both a typed parameter and an equivalent extra parameter are specified (e.g., max-tokens: 100 and max_tokens: 500), the extra parameter takes precedence and the API will receive the extra parameter's value.

Currently supported parameters for Anthropic models include:

  • max-tokens: Controls the maximum number of tokens available to the model for generating a response.
  • thinking-budget-tokens: Enables extended thinking with a fixed token budget, giving the model more reasoning capacity on complex tasks. Must be at least 1024 and less than max-tokens. Ignored when effort is also set. Without either setting, thinking follows the model default. Deprecated: Claude Opus 4.7+ removed fixed thinking budgets; setting this returns a 400 error. Use effort instead.
  • effort: Enables adaptive extended thinking and guides how deeply the model reasons before responding, from quick answers (low) to thorough multi-step reasoning (max) (values: low, medium, high, xhigh, max). Without effort or thinking-budget-tokens, thinking follows the model default. When set, thinking-budget-tokens is ignored. Use max-tokens to cap total output (thinking + response text). The xhigh level is recommended for coding and agentic use cases on Claude Opus 4.7+.
  • thinking: Explicitly overrides the thinking mode (values: disabled, between_tools). disabled explicitly disables thinking and cannot be combined with thinking-budget-tokens or effort above high. between_tools, available on Claude Sonnet 5.5, disables up-front thinking while allowing progress thinking between tool calls; it supports low, medium, and high effort and cannot be combined with thinking-budget-tokens. When omitted, the model's default thinking behaviour is used unless effort or thinking-budget-tokens requests an explicit mode.
  • temperature: Controls randomness/creativity of responses (range: 0.0 to 1.0, default: 1.0). Lower values produce more focused and deterministic outputs. Deprecated: Claude Opus 4.7+ rejects non-default values with a 400 error.
  • top-p: Controls diversity via nucleus sampling (range: 0.0 to 1.0). Lower values produce more focused outputs. Deprecated: Claude Opus 4.7+ rejects non-default values with a 400 error.
  • top-k: Limits tokens considered for each position to top K options. Higher values allow more diverse outputs. Deprecated: Claude Opus 4.7+ rejects any value with a 400 error.
  • stream: If true, enables streaming mode for the API response. Streaming is recommended for requests with large max-tokens values, especially when extended thinking is enabled, to prevent HTTP timeouts on long-running requests. Responses are streamed incrementally and buffered internally before processing.
  • legacy-structured-output: If true, uses tool-based structured output instead of native JSON schema output. This is a workaround for models that have difficulty producing valid responses with native output_config.format constrained decoding when extended thinking is enabled. When set, the provider registers a submit_response tool and instructs the model to use it to submit its response.
  • prompt-cache-ttl: Enables Anthropic prompt caching and selects the cache lifetime. Supported values are 5m and 1h. When set, MindTrial enables top-level automatic caching and places one explicit cache breakpoint on the most reusable request prefix: the final configured local tool, the final system block, or the final cacheable initial user content block. When omitted, no prompt cache controls are added.

Currently supported parameters for Google models include:

  • text-response-format: If true, use plain-text response format (less reliable) for compatibility with models that do not support JSON. This setting applies to all tasks, including those with and without tools enabled.
  • text-response-format-with-tools: If true, forces plain-text response format when tools are enabled (required for pre-Gemini 3 models). If false or unset, uses JSON schema mode with tools (Gemini 3+ default behavior). This setting only applies to tasks with tools enabled.
  • thinking-level: Controls the maximum depth of the model's internal reasoning process (values: minimal, low, medium, high). The default is model-dependent — for example, Gemini 3 Pro defaults to high, while Gemini 3.5 Flash defaults to medium. minimal minimizes reasoning for lowest latency (does not guarantee thinking is disabled), low minimizes latency and cost for simple tasks, medium balances reasoning depth and latency, while high maximizes reasoning depth for complex tasks (the model may take longer but output is more carefully reasoned).
  • media-resolution: Controls the maximum number of tokens allocated per input image (values: low, medium, high). Higher resolutions improve fine text reading and small detail identification but increase token usage and latency. low uses 280 tokens; medium uses 560 tokens; high uses 1120 tokens. If unspecified, the model uses optimal defaults.
  • temperature: Controls randomness/creativity of responses (range: 0.0 to 2.0, default: 1.0). Lower values produce more focused and deterministic outputs. For Gemini 3, it's recommended to keep temperature at default 1.0 for optimal reasoning performance.
  • top-p: Controls diversity via nucleus sampling (range: 0.0 to 1.0). Lower values produce more focused outputs.
  • top-k: Limits tokens considered for each position to top K options. Higher values allow more diverse outputs.
  • presence-penalty: Penalizes new tokens based on whether they appear in the text so far. Positive values discourage reuse of tokens, increasing vocabulary. Negative values encourage token reuse.
  • frequency-penalty: Penalizes new tokens based on their frequency in the text so far. Positive values discourage frequent tokens proportionally. Negative values encourage token repetition.
  • seed: Seed used for deterministic generation. When set, the model attempts to provide consistent responses for identical inputs.

Currently supported parameters for DeepSeek models include:

  • temperature: Controls randomness/creativity of responses (range: 0.0 to 2.0, default: 1.0). Lower values produce more focused and deterministic outputs.
  • top-p: Controls diversity via nucleus sampling (range: 0.0 to 1.0). Lower values produce more focused outputs.
  • presence-penalty: Penalizes new tokens based on their presence in text so far (range: -2.0 to 2.0, default: 0.0). Positive values encourage model to use new tokens.
  • frequency-penalty: Penalizes new tokens based on their frequency in text so far (range: -2.0 to 2.0, default: 0.0). Positive values encourage model to use less frequent tokens.
  • thinking: Toggles reasoning (thinking) mode for V4 and newer thinking-capable models. When enabled, the model produces chain-of-thought reasoning before the final answer (values: enabled, disabled; default: enabled for V4 models). Note: temperature, top-p, presence-penalty, and frequency-penalty are silently ignored in thinking mode.
  • reasoning-effort: Controls how deeply the model reasons in thinking mode (values: low, medium, high, xhigh, max). Default is high. Currently only high and max are distinct effective values — low and medium are mapped to high, and xhigh is mapped to max.
  • max-tokens: Controls the maximum number of tokens generated by the model.

Currently supported parameters for Mistral AI models include:

  • temperature: Controls randomness/creativity of responses (range: 0.0 to 1.5). Lower values produce more focused and deterministic outputs.
  • top-p: Controls diversity via nucleus sampling (range: 0.0 to 1.0). Lower values produce more focused outputs.
  • max-tokens: Controls the maximum number of tokens available to the model for generating a response.
  • presence-penalty: Penalizes new tokens based on their presence in text so far (range: -2.0 to 2.0, default: 0.0). Positive values encourage model to use new tokens.
  • frequency-penalty: Penalizes new tokens based on their frequency in text so far (range: -2.0 to 2.0, default: 0.0). Positive values encourage model to use less frequent tokens.
  • random-seed: Provides the seed to use for random sampling. If set, requests will generate deterministic results.
  • prompt-mode: When set to "reasoning", instructs the model to reason if supported.
  • reasoning-effort: Controls reasoning depth for current models (values: none, minimal, low, medium, high, xhigh). This is independent of the legacy prompt-mode; high can increase output-token cost.
  • safe-prompt: Enables content filtering to ensure outputs comply with usage policies.

Currently supported parameters for xAI models include:

  • temperature: Controls randomness/creativity of responses (range: 0.0 to 2.0, default: 1.0). Lower values produce more focused and deterministic outputs.
  • top-p: Controls diversity via nucleus sampling (range: 0.0 to 1.0, default: 1.0). Lower values produce more focused outputs.
  • max-completion-tokens: Controls the maximum number of tokens available to the model for generating a response.
  • presence-penalty: Penalizes new tokens based on their presence in text so far (range: -2.0 to 2.0, default: 0.0). Positive values encourage model to use new tokens.
  • frequency-penalty: Penalizes new tokens based on their frequency in text so far (range: -2.0 to 2.0, default: 0.0). Positive values encourage model to use less frequent tokens.
  • reasoning-effort: Controls effort on reasoning for supported reasoning-capable models (values: low, medium, high, xhigh). Not all xAI reasoning models (i.e. Grok 4) accept this parameter; Grok 4.5 defaults to high when unset. The xhigh level is supported on Grok 4.6 and later.
  • seed: Integer seed to request deterministic sampling when possible. Determinism is best-effort. xAI makes a best-effort to return repeatable outputs for identical inputs when seed and other parameters are the same.

Currently supported parameters for Alibaba models include:

  • response-format: Selects json-schema, json-object, or text. MindTrial applies its legacy schema-instruction behavior when this is omitted and structured output is not disabled. Cannot be combined with the deprecated text-response-format or disable-legacy-json-mode properties.
  • stream: If true, enables streaming mode for the API response. Some models (e.g. QwQ, QVQ, and Qwen-Omni) require streaming to be enabled. Responses are streamed incrementally and buffered internally before processing.
  • enable-thinking: Enables hybrid thinking on supported Qwen models.
  • preserve-thinking: Preserves reasoning_content across tool-call turns. Preserved reasoning is included in later input-token counts and billing.
  • thinking-budget: Optional positive token budget for thinking. This is distinct from max-tokens/max-completion-tokens, which limit the complete generated response. Mutually exclusive with reasoning-effort.
  • reasoning-effort: Controls reasoning depth for supported Qwen models (e.g. Qwen 3.8 Max) (values: none, minimal, low, medium, high, xhigh, max). Mutually exclusive with thinking-budget.
  • temperature: Controls randomness/creativity of responses (range: 0.0 to 2.0, default: 1.0). Lower values produce more focused and deterministic outputs.
  • top-p: Controls diversity via nucleus sampling (range: 0.0 to 1.0). Lower values produce more focused outputs.
  • max-tokens: Controls the maximum number of tokens available to the model for generating a response. Deprecated: use max-completion-tokens instead. Mutually exclusive with max-completion-tokens.
  • max-completion-tokens: Controls the maximum number of tokens available to the model for generating a response, including reasoning tokens for thinking models. Mutually exclusive with max-tokens.
  • presence-penalty: Penalizes new tokens based on whether they appear in the text so far (range: -2.0 to 2.0, default: 0.0). Positive values encourage introducing new topics.
  • frequency-penalty: Penalizes new tokens based on their frequency in text so far (range: -2.0 to 2.0, default: 0.0). Positive values encourage model to use less frequent tokens.
  • seed: Makes text generation more deterministic by using the same seed value. When using the same seed and keeping other parameters unchanged, the model makes best-effort to return consistent outputs for identical inputs.
  • text-response-format: If true, use plain-text response format (less reliable) for compatibility with models that do not support JSON (for example, when thinking is enabled on certain Qwen models). Deprecated: use response-format: text instead. Cannot be combined with response-format.
  • disable-legacy-json-mode: Compatibility toggle that controls legacy prompt injection for JSON formatting. Default: false (legacy mode on), which adds an explicit JSON formatting instruction to the prompt for improved compatibility with most Qwen models. Setting this to true disables the legacy prompt injection. For best compatibility and reliable JSON responses, keep this set to false unless you are certain the target model works correctly without legacy prompt injection. Deprecated: use response-format: json-schema instead. Cannot be combined with response-format.

Currently supported parameters for Moonshot AI models include:

  • temperature: Controls randomness/creativity of responses (range: 0.0 to 1.0, default: 0.0). Higher values make output more random, while lower values make it more focused and deterministic. Moonshot AI recommends 0.6 for kimi-k2 models and 1.0 for kimi-k2-thinking models.
  • top-p: Controls diversity via nucleus sampling (range: 0.0 to 1.0, default: 1.0). Lower values produce more focused outputs. Generally, change either this or temperature, but not both at the same time.
  • max-tokens: Controls the maximum number of tokens to generate for the chat completion.
  • max-completion-tokens: Controls the modern maximum number of generated tokens. It cannot be combined with deprecated max-tokens.
  • reasoning-effort: Controls Kimi K3 reasoning depth (values: low, high, max).
  • response-format: Selects json-schema, json-object, or text. MindTrial applies json-object when this is omitted (unless structured output is disabled). Kimi K3 can use native json-schema.
  • stream: Enables streaming and usage accumulation for long-running responses.
  • presence-penalty: Penalizes new tokens based on whether they appear in the text (range: -2.0 to 2.0, default: 0.0). Positive values increase the likelihood of the model discussing new topics.
  • frequency-penalty: Penalizes new tokens based on their existing frequency in the text (range: -2.0 to 2.0, default: 0.0). Positive values reduce the likelihood of the model repeating the same phrases verbatim.
  • thinking: Toggles the reasoning (thinking) capability for thinking-capable models such as kimi-k2.6. Accepted values are enabled (default for kimi-k2.6) and disabled. Older Kimi models that do not support this parameter should omit it.
  • preserve-thinking: Enables Moonshot's Preserved Thinking feature for kimi-k2.6, which preserves the model's chain-of-thought across model calls that share the same conversation context (e.g. successive calls in a tool-using task), so the model can build on its earlier reasoning. Accepted value: all; when omitted, prior reasoning is dropped between calls — reducing token cost at the expense of chain-of-thought continuity. Older Kimi models do not support this parameter and should omit it.

For kimi-k2.5 and kimi-k2.6, Moonshot AI fixes temperature, top-p, presence-penalty, and frequency-penalty to model-specific defaults — supplying any of these parameters will cause the API to reject the request.

[!NOTE] The results will be saved to <output-dir>/<output-basename>.<format>. If the result output file already exists, it will be replaced. If the log file already exists, it will be appended to.

[!TIP] The following placeholders are available for output paths and names:

  • {{.Year}}: Current year
  • {{.Month}}: Current month
  • {{.Day}}: Current day
  • {{.Hour}}: Current hour
  • {{.Minute}}: Current minute
  • {{.Second}}: Current second

[!TIP] If log-file and/or output-basename is blank, the log and/or output will be written to the stdout.

[!NOTE] MindTrial processes tasks across different AI providers simultaneously (in parallel). However, when running multiple configurations from the same provider (e.g. different OpenAI models), these are processed one after another (sequentially) by default.

[!TIP] To run multiple configurations from the same provider in parallel, set max-parallel-requests-per-minute on the provider. This enables parallel execution of all runs within that provider, while limiting the aggregate number of API requests per minute across all runs to the specified value. When set to 0 (or omitted), runs execute sequentially (the default behavior).

[!TIP] Models can use the max-requests-per-minute property in their run configurations to limit the number of requests made per minute.

[!TIP] To automatically retry failed requests due to rate limiting or other transient errors, set retry-policy at the provider level to apply to all runs. An individual run configuration can override this by setting its own retry-policy:

  • max-retry-attempts: Maximum number of retry attempts (default: 0 means no retry).
  • initial-delay-seconds: Initial delay before the first retry in seconds.

Retries use exponential backoff starting with the initial delay.

[!TIP] To estimate what a trial run costs, set pricing at the application, provider, or run level. Rates are per million tokens:

  • currency: ISO 4217 code the rates are expressed in (default: USD).
  • input-per-million: Price per million uncached input tokens.
  • output-per-million: Price per million generated output tokens.
  • cache-read-per-million: Price per million input tokens read from a prompt cache (default: the input price).
  • cache-write-per-million: Price per million input tokens written to a prompt cache (default: the input price).
  • reasoning-per-million: Price per million reasoning tokens (default: the output price).

pricing is inherited as a whole, not field-by-field: the nearest level (run, then provider, then application) that sets any field is used in its entirety, replacing rather than merging with whatever a less specific level configured. To override a single rate, restate the full price list at that level:

config:
  pricing:
    currency: USD
    input-per-million: 1.25
    output-per-million: 10.00
  providers:
    - name: openai
      runs:
        - name: "GPT-5.2"
          model: "gpt-5.2"
          # Overriding output-per-million requires restating the rest of the list.
          pricing:
            currency: USD
            input-per-million: 1.25
            output-per-million: 12.00

Estimated costs are reported by the stats command and are always estimates derived from these configured rates, never billed amounts. A rate left unset is treated as unknown rather than free, so any estimate depending on it is omitted instead of being understated. The effective prices are also recorded in JSON results, so historical estimates stay reproducible.

[!TIP] To disable all run configurations for a given provider, set disabled: true on that provider. An individual run configuration can override this by setting disabled: false (e.g. to enable just that one configuration).

Example snippet from config.yaml:

# config.yaml
config:
  log-file: ""
  output-dir: "./results/{{.Year}}-{{.Month}}-{{.Day}}/"
  output-basename: "{{.Hour}}-{{.Minute}}-{{.Second}}"
  task-source: "./tasks.yaml"
  providers:
    - name: openai
      disabled: true
      client-config:
        # Resolved from the OPENAI_API_KEY environment variable at load time.
        api-key: "{{.Env.OPENAI_API_KEY}}"
      retry-policy:
        max-retry-attempts: 5
        initial-delay-seconds: 30
      runs:
        - name: "4o-mini - latest"
          disabled: false
          model: "gpt-4o-mini"
          max-requests-per-minute: 3
        - name: "o1-mini - latest"
          model: "o1-mini"
          max-requests-per-minute: 3
          model-parameters:
            text-response-format: true
        - name: "o3-mini - latest (high reasoning)"
          model: "o3-mini"
          max-requests-per-minute: 3
          model-parameters:
            reasoning-effort: "high"
    - name: openrouter
      retry-policy:
        max-retry-attempts: 5
        initial-delay-seconds: 30
      max-parallel-requests-per-minute: 30
      client-config:
        api-key: "<your-api-key>"
      runs:
        - name: "OpenAI GPT-5.2 (xhigh reasoning)"
          model: "openai/gpt-5.2"
          max-requests-per-minute: 20
          model-parameters:
            verbosity: "medium"
            # Pass-through parameters use OpenAI API naming (underscores).
            reasoning_effort: "xhigh"
        - name: "GPT via Fusion (xhigh panel + judge, high outer)"
          # Outer model.
          model: "~openai/gpt-latest"
          max-requests-per-minute: 2
          model-parameters:
            # Outer model reasoning effort.
            reasoning:
              effort: "high"
            server-tools:
              - type: openrouter:fusion
                parameters:
                  # Inner panel models.
                  analysis_models:
                    - "~anthropic/claude-opus-latest"
                    - "~openai/gpt-latest"
                    - "~google/gemini-pro-latest"
                  # Judge model.
                  model: "~anthropic/claude-opus-latest"
                  # Inner panel + judge parameters.
                  reasoning:
                    effort: "xhigh"
                  max_completion_tokens: 65536
                  max_tool_calls: 16
        - name: "Google Gemma 3 27B IT (free)"
          model: "google/gemma-3-27b-it:free"
          max-requests-per-minute: 3
          model-parameters:
            response-format: "text"
    - name: google
      client-config:
        api-key: "<your-api-key>"
      runs:
        - name: "Gemini 2.5 Pro - latest"
          model: "gemini-2.5-pro"
          max-requests-per-minute: 3
          model-parameters:
            text-response-format-with-tools: true
        - name: "Gemini 3 Pro - latest"
          model: "gemini-3-pro-preview"
          max-requests-per-minute: 3
          model-parameters:
            thinking-level: "high"
            media-resolution: "high"
    - name: anthropic
      client-config:
        api-key: "<your-api-key>"
      runs:
        - name: "Claude 3.7 Sonnet - latest"
          model: "claude-3-7-sonnet-latest"
          max-requests-per-minute: 5
          model-parameters:
            max-tokens: 4096
        - name: "Claude 3.7 Sonnet - latest (extended thinking)"
          model: "claude-3-7-sonnet-latest"
          max-requests-per-minute: 5
          model-parameters:
            max-tokens: 8192
            thinking-budget-tokens: 2048
            stream: true
        - name: "Claude 4.6 Opus - latest (max adaptive thinking)"
          model: "claude-opus-4-6"
          max-requests-per-minute: 5
          model-parameters:
            max-tokens: 65536
            effort: max
            stream: true
        - name: "Claude Opus 4.7 (xhigh adaptive thinking)"
          model: "claude-opus-4-7"
          max-requests-per-minute: 5
          model-parameters:
            max-tokens: 65536
            effort: xhigh
            stream: true
        - name: "Claude Opus 5 (xhigh adaptive thinking with prompt caching)"
          model: "claude-opus-5"
          max-requests-per-minute: 5
          model-parameters:
            max-tokens: 65536
            effort: xhigh
            stream: true
            prompt-cache-ttl: 5m
    - name: deepseek
      client-config:
        api-key: "<your-api-key>"
        request-timeout: 10m
      runs:
        - name: "DeepSeek-V3.1 - latest (thinking mode)"
          model: "deepseek-reasoner"
          max-requests-per-minute: 15
    - name: mistralai
      client-config:
        api-key: "<your-api-key>"
      runs:
        - name: "Mistral Large - latest"
          model: "mistral-large-latest"
          max-requests-per-minute: 5
          retry-policy:
            max-retry-attempts: 5
            initial-delay-seconds: 30
        - name: "Mistral Medium 3.5 - latest (high reasoning)"
          model: "mistral-medium-3-5"
          max-requests-per-minute: 5
          model-parameters:
            max-tokens: 65536
            reasoning-effort: "high"
    - name: alibaba
      client-config:
        api-key: "<your-api-key>"
        endpoint: "https://dashscope-intl.aliyuncs.com/compatible-mode/v1"  # Singapore region
      retry-policy:
        max-retry-attempts: 5
        initial-delay-seconds: 30
      runs:
        - name: "Qwen3-Max-Preview"
          model: "qwen3-max-preview"
          max-requests-per-minute: 30
        - name: "Qwen3-Max-Preview - unstructured"
          model: "qwen3-max-preview"
          disable-structured-output: true
          max-requests-per-minute: 30
        - name: "Qwen-VL-Max-Latest"
          model: "qwen-vl-max-latest"
          max-requests-per-minute: 30
          model-parameters:
            disable-legacy-json-mode: true
        - name: "Qwen3-Next-80B-A3B-Thinking"
          model: "qwen3-next-80b-a3b-thinking"
          max-requests-per-minute: 30
          model-parameters:
            text-response-format: true
        - name: "QVQ-Max (vision reasoning)"
          model: "qvq-max"
          max-requests-per-minute: 30
          model-parameters:
            stream: true  # Required for QvQ models
        - name: "Qwen3.7 Plus - latest (thinking)"
          model: "qwen3.7-plus"
          max-requests-per-minute: 30
          model-parameters:
            enable-thinking: true
            preserve-thinking: true
            max-tokens: 65536
            stream: true
        - name: "Qwen3.8 Max - latest (xhigh reasoning)"
          model: "qwen3.8-max"
          max-requests-per-minute: 30
          model-parameters:
            reasoning-effort: "xhigh"
            preserve-thinking: true
            max-completion-tokens: 65536
            response-format: "json-object"
            stream: true
    - name: moonshotai
      client-config:
        api-key: "<your-api-key>"
      runs:
        - name: "Kimi K2 - latest (thinking)"
          model: "kimi-k2-thinking"
          text-only: true  # Skip tasks that require native file input
          max-requests-per-minute: 3
          model-parameters:
            temperature: 1.0
            max-tokens: 16000
        - name: "Kimi K2.6 (thinking)"
          model: "kimi-k2.6"
          max-requests-per-minute: 3
          model-parameters:
            max-tokens: 32000
            thinking: enabled     # default for kimi-k2.6; "disabled" turns off reasoning
            preserve-thinking: all  # preserve chain-of-thought across model calls (Preserved Thinking)
        - name: "Kimi K3 - latest (max reasoning)"
          model: "kimi-k3"
          max-requests-per-minute: 3
          model-parameters:
            reasoning-effort: "max"
            max-completion-tokens: 65536
            response-format: "json-schema"
            stream: true

tasks.yaml

This file defines the tasks to be executed on all enabled run configurations. Each task defines the following properties; name, prompt, and response-result-format are always required, and expected-result is required unless the task uses a custom validator:

  • name: A unique display-friendly name to be shown in the results.
  • prompt: The prompt (i.e. task) that will be sent to the AI model.
  • response-result-format: Defines how the AI should format the final answer to the prompt. This can be either:
    • Plain text format: A string instruction describing the expected answer format (e.g., "single number", "list of words separated by commas").
    • Structured schema format: A JSON schema object defining the structure of the expected response for complex data (e.g., objects with specific fields and types).
  • expected-result: Defines the accepted valid answer(s) to the prompt. The format depends on the response-result-format type:
    • For plain text format: A string value or list of string values that follow the format instruction precisely.
    • For structured schema format: An object value or list of object values that conform to the JSON schema definition. Only one expected result needs to match for the response to be considered correct. With a custom validator, expected-result is optional trusted reference data that is passed to the validator unchanged and does not need to match response-result-format.

Optionally, a task can include a list of files to be sent along with the prompt:

  • files: A list of files to attach to the prompt. Each file entry defines the following properties:

    • name: A unique name for the file. Every attached file is announced to the model in the prompt as [file: <name>], regardless of its access setting, and local tools mount the file under this exact name.
    • uri: The path or URI to the file. Local file paths and remote HTTP/HTTPS URLs are supported. The content is loaded on demand: it is sent with the request for files with native access, and copied to the tool auxiliary directory for files with local access.
    • type: The MIME type of the file (e.g., image/png, image/jpeg). If omitted, the tool will attempt to infer the type based on the file extension or content.
    • options: Optional per-file processing options that override the file-options defaults from the task-config section.
      • image-detail: Controls the fidelity level at which the model processes input images (values: auto, low, medium, high, original). If the provider does not natively support the requested level, the next higher level or the highest available level is selected. If not set or unknown, the provider uses its own default behavior. Currently only the OpenAI provider honors this setting. It also controls how PDF pages are rendered as images, but only on the OpenAI Responses API; the Chat Completions API does not accept a detail level for file inputs. The original level has no PDF equivalent and is treated as high.
      • access: Controls how the file is exposed. Supported values are native (provider native file/multimodal input) and local (local Docker tools). When omitted, the stable default is [native, local]. A per-file access replaces the inherited file-options value entirely (no union). An explicit empty list [] is invalid. The setting applies to every file type including images, so an image with access: [local] is never sent to the model and can only be inspected through a tool. Tasks with files that are only available to local tools (access: [local]) do not require provider file support and are not skipped by text-only.
  • file-options: Default file processing options for all task files in task-config. Individual files can override via files[].options.

    • image-detail: Default image fidelity (same values as above).
    • access: Default access list (same semantics as above). Example: file-options: { access: [native, local] }.

[!NOTE] If a task requires native file input (any file with native access), it will be skipped for provider configurations that do not support file uploads or the specific file type. Tasks with only local access never require native file support.

[!IMPORTANT] A file with local access is only readable if the task also enables at least one tool that defines auxiliary-dir. Otherwise the model sees the [file: <name>] reference but has no way to read the contents.

[!NOTE] Currently supported image types include: image/jpeg, image/jpg, image/png, image/gif, image/webp. Support may vary by provider. Native non-image (document) input is additionally supported by:

  • OpenAI: the complete accepted file types list — PDF; Word, Excel, PowerPoint, Pages, Keynote, Google Docs/Sheets/Slides, RTF and OpenDocument text; CSV, TSV and IIF; and a broad set of text and code formats including plain text, Markdown, HTML, XML, CSS, JSON, YAML, TOML, calendar, vCard, subtitles, email and most programming languages
  • Google: application/pdf, application/json, text/plain, text/html, text/css, text/xml, text/csv, text/rtf and text/javascript, per the supported content types. Only PDF is read with document vision; the other types are extracted as plain text, so charts and formatting are lost.
  • Anthropic: application/pdf and text/plain, per the citations documentation. A plain text document is sent verbatim rather than base64-encoded; other text formats such as CSV or Markdown must declare type: "text/plain" explicitly to use this path.
  • OpenRouter, Mistral AI: application/pdf
  • Alibaba, Moonshot AI, xAI: images only
  • DeepSeek: images only on vision-capable models

A file with native access whose type is not supported by the selected provider makes the task unsupported; it is never silently downgraded to local access.

Example file access configuration:

task-config:
  file-options:
    access: [native, local]  # default for all files
  tasks:
    - name: "vision task"
      prompt: "Describe the image."
      response-result-format: "single sentence"
      expected-result: "A cat."
      files:
        - name: "picture"
          uri: "./taskdata/cat.png"
          type: "image/png"
          # inherits access: [native, local]
    - name: "local-only task"
      prompt: "Use the data file to answer."
      response-result-format: "single number"
      expected-result: "42"
      files:
        - name: "data"
          uri: "./taskdata/data.csv"
          type: "text/csv"
          options:
            access: [local]  # mounted to tool auxiliary-dir only, never sent natively

[!TIP] To disable all tasks by default, set disabled: true in the task-config section. An individual task can override this by setting disabled: false (e.g. to enable just that one task).

Optionally, a task can also carry descriptive metadata used for filtering and grouping:

  • suite: A grouping label for organizing related tasks (e.g. a benchmark suite name).
  • category: A classification label for the task (e.g. "math", "coding").
  • difficulty: A free-form difficulty label for the task (e.g. "easy", "hard").
  • tags: A list of free-form labels for filtering and grouping tasks.

These fields are optional and have no effect on task execution or validation. When present, they are included in the JSON results (TaskMetadata), can be filtered on in the HTML report alongside the existing status/task filters, and are used for the suite/category/difficulty cycling hotkeys (s/c/d) and tag search (/) in the interactive task picker's checklist.

- name: "math problem"
  suite: "arithmetic-basics"
  category: "math"
  difficulty: "easy"
  tags: ["smoke", "regression"]
  prompt: "What is 2 + 2?"
  response-result-format: "single number"
  expected-result: "4"

Structured Response Formats

MindTrial supports two types of response formats for tasks:

Plain Text Format

For tasks where the final answer can be represented as a text value:

- name: "math problem"
  prompt: "What is 2 + 2?"
  response-result-format: "single number"
  expected-result: "4"
Structured Schema Format

For tasks requiring complex structured answers, you can define a JSON schema that describes the expected response format:

- name: "perfect square check"
  prompt: "For each number in [4, 9, 10, 16], determine if it's a perfect square and if so, provide the square root."
  response-result-format:
    type: array
    items:
      type: object
      additionalProperties: false
      properties:
        number:
          type: integer
        is_perfect_square:
          type: boolean
        square_root:
          type: integer
      required: ["number", "is_perfect_square"]
  expected-result:
    - - number: 4
        is_perfect_square: true
        square_root: 2
      - number: 9
        is_perfect_square: true
        square_root: 3
      - number: 10
        is_perfect_square: false
      - number: 16
        is_perfect_square: true
        square_root: 4

[!IMPORTANT] Structured schema format caveats:

  • Semantic validation (LLM judges) cannot be used with structured schema-based response formats.
  • All expected results must be objects that conform to the same schema. For array schemas, the entire expected array must be wrapped in a single list item under expected-result to avoid treating each array element as a separate expected answer.
  • Models must support structured JSON response generation for reliable results.
  • The OpenAI provider requires JSON schemas to have additionalProperties: false and all fields must be required (no optional fields allowed). Other providers may be more flexible.

System Prompt

The system prompt controls how the response format instruction is presented to the AI model.

You can customize this template globally for all tasks in the task-config section, and override it for individual tasks if needed. The template uses Go's template syntax and can reference {{.ResponseResultFormat}} to include the task's response-result-format.

Default system prompt for all tasks is:

Provide the final answer in exactly this format: {{.ResponseResultFormat}}

  • system-prompt: A configuration section for the system prompt.
    • template: The template string for the system prompt instruction. If not specified, uses the default.
    • enable-for: Controls when system prompt should be sent to AI models. Options:
      • "all": Send system prompt for all tasks (both plain text and structured schema formats).
      • "text": Send system prompt only for tasks with plain text response format (default).
      • "none": Do not send system prompt.

[!NOTE] For structured schema response formats, the JSON schema is automatically passed to the AI model through the provider's structured response mechanism, making explicit format instructions in the system prompt optional.

Validation Rules

These rules control how the validator compares the model's answer to the expected results. By default, comparisons are case-insensitive and only trim leading and trailing whitespace.

You can set validation rules globally for all tasks in the task-config section, and override them for individual tasks if needed; any option not specified at the task level will inherit the global setting from task-config:

  • validation-rules: Controls how model responses are validated against expected results.
    • case-sensitive: If true, comparison is case-sensitive. If false (default), comparison ignores case.
    • ignore-whitespace: If true, all whitespace (spaces, tabs, newlines) is removed before comparison. If false (default), only leading/trailing whitespace is trimmed, and internal whitespace is preserved.
    • trim-lines: If true, trims leading and trailing whitespace from each line before comparison while preserving internal spaces within lines. CRLF line endings are normalized to LF. This option is ignored when ignore-whitespace is enabled. If false (default), lines are not individually trimmed.
    • schema-validation: If true, validates the raw candidate answer against the single expected-result JSON Schema. See Schema-Based Validation for details. Mutually exclusive with judge and custom-validator.
    • judge: Optional LLM-based semantic validation. See Judge-Based Validation for details. Mutually exclusive with schema-validation and custom-validator.
      • enabled: If true, uses an LLM judge to evaluate semantic equivalence. If false (default), uses exact value matching.
      • name: The name of the judge configuration defined in the config.yaml file.
      • variant: The specific run variant from the judge's provider to use.
    • custom-validator: Name of a trusted Docker-backed validator defined in the config.yaml file that decides whether the answer is correct. See Custom Validators for details. Mutually exclusive with schema-validation and judge. Set it to "" in a task to opt out of an inherited custom validator.

Judge-Based Validation

For complex or open-ended tasks where exact value matching is insufficient, you can configure LLM judges to evaluate responses semantically. This is particularly useful for creative writing, reasoning tasks, or when multiple valid answer formats exist.

How it works: Instead of comparing text exactly, an LLM judge evaluates whether the model's response semantically matches the expected result, considering meaning and intent rather than exact wording.

To use judge validation:

  1. Define and configure judge models in config.yaml:

    config:
      # ... existing configuration ...
      judges:
        - name: "mistral-judge"  # A unique name for the judge configuration.
          provider:
            name: "mistralai"
            client-config:
              api-key: "<your-api-key>"
            runs:
              - name: "fast"
                model: "mistral-medium-latest"
                max-requests-per-minute: 30
                model-parameters:
                  temperature: 0.20
                  random-seed: 847629
              - name: "reasoning"
                model: "magistral-medium-latest"
                max-requests-per-minute: 30
                model-parameters:
                  prompt-mode: "reasoning"
                  temperature: 0.20
                  random-seed: 847629
        - name: "deepseek-judge"
          provider:
            name: "deepseek"
            client-config:
              api-key: "<your-api-key>"
            runs:
              - name: "fast"
                model: "deepseek-chat"
                max-requests-per-minute: 30
                model-parameters:
                  temperature: 0.20
              - name: "reasoning"
                model: "deepseek-reasoner"
                max-requests-per-minute: 30
    
  2. Enable judge validation in tasks.yaml:

    # Enable globally for all tasks.
    task-config:
      validation-rules:
        judge:
          enabled: true
          name: "mistral-judge"
          variant: "fast"
      tasks:
        # ... tasks will use judge validation by default ...
    
    # Override per-task (inherit global settings and override specific options).
    task-config:
      validation-rules:
        judge:
          enabled: false  # Default: use exact value matching.
          name: "mistral-judge"
          variant: "fast"
      tasks:
        - name: "exact matching task"
          prompt: "What is 2+2?"
          response-result-format: "single number"
          expected-result: "4"
          # Inherits global validation-rules (exact value matching).
        
        - name: "creative writing task"
          prompt: "Write a short story about..."
          response-result-format: "short story narrative"
          expected-result: "A creative and engaging short story"
          validation-rules:
            judge:
              enabled: true  # Override: enable judge validation for this task.
              # Inherits name: "mistral-judge" and variant: "fast" from global config.
        
        - name: "complex reasoning task"
          prompt: "Analyze this philosophical argument..."
          response-result-format: "structured analysis with reasoning"
          expected-result: "A thoughtful analysis with logical reasoning"
          validation-rules:
            judge:
              enabled: true
              variant: "reasoning"  # Override: use reasoning run variant instead of fast.
              # Inherits name: "mistral-judge" from global config.
    

Schema-Based Validation

For tasks where exact value matching is too strict but an LLM judge is unnecessary, you can validate responses against a JSON Schema. This is particularly useful for numeric ranges, tolerances, bounding boxes, or any structured output where the set of acceptable answers is better expressed as constraints.

How it works: Instead of comparing the candidate answer to literal expected values, the raw candidate answer is validated directly against the single expected-result JSON Schema without canonicalization, normalization, or type coercion. The schema's $schema is optional and defaults to Draft 2020-12; normalization flags (case-sensitive, ignore-whitespace, trim-lines) are ignored and the mode is mutually exclusive with judge.

[!IMPORTANT] Schema validation performs no type coercion. A plain-text response-result-format always produces a string, so its expected-result schema must accept strings (e.g., type: string with pattern). A schema with type: number will not match the string "10" — use a structured response-result-format such as type: number when you need numeric validation.

To use schema validation, enable it in tasks.yaml:

# Simple range — any number between 9.9 and 10.1 is accepted.
task-config:
  tasks:
    - name: "numeric range"
      prompt: "Pick a number between 9.9 and 10.1"
      response-result-format:
        type: number
      validation-rules:
        schema-validation: true
      expected-result:
        $schema: "https://json-schema.org/draft/2020-12/schema"
        type: number
        minimum: 9.9
        maximum: 10.1

Plain-text responses can still use schema validation — the schema just needs to describe a string:

- name: "hex colour"
  prompt: "Return a six-digit hexadecimal colour starting with #"
  response-result-format: "a six-digit hexadecimal colour starting with #"
  validation-rules:
    schema-validation: true
  expected-result:
    $schema: "https://json-schema.org/draft/2020-12/schema"
    type: string
    pattern: "^#[0-9A-Fa-f]{6}$"

When the model is asked to produce structured JSON, response-result-format and the validator schema are distinct — the former is sent to the model, the latter is evaluator-only:

- name: locate-object
  prompt: >
    Locate the red object in the image and return its normalized
    bounding box.

  response-result-format:
    type: object
    properties:
      x:
        type: number
        minimum: 0
        maximum: 1
      y:
        type: number
        minimum: 0
        maximum: 1
      width:
        type: number
        minimum: 0
        maximum: 1
      height:
        type: number
        minimum: 0
        maximum: 1
    required: [x, y, width, height]
    additionalProperties: false

  expected-result:
    $schema: "https://json-schema.org/draft/2020-12/schema"
    type: object
    properties:
      x:
        type: number
        minimum: 0.31
        maximum: 0.35
      y:
        type: number
        minimum: 0.42
        maximum: 0.46
      width:
        type: number
        minimum: 0.19
        maximum: 0.23
      height:
        type: number
        minimum: 0.27
        maximum: 0.31
    required: [x, y, width, height]
    additionalProperties: false

  validation-rules:
    schema-validation: true

[!TIP] Use anyOf/oneOf/allOf/enum inside the single schema to express alternatives instead of multiple outer expected-result values. A literal object containing $schema without schema-validation: true is still compared via standard value matching — the canonical values are compared exactly (e.g., when the model was asked to generate a schema).

Judge Prompt Customization

MindTrial automatically applies a built-in semantic evaluation template that compares candidate responses against expected answers. For advanced use cases, you can customize the judge prompt template, response format, and acceptance criteria.

Judge prompts can be customized in the validation-rules.judge.prompt section of your tasks.yaml file, either globally in task-config or individually per task.

Customization Fields:

  • template: Custom prompt template for the judge (supports template variables listed below).
  • verdict-format: Expected response format from the judge (plain text instruction or JSON schema).
  • passing-verdicts: Set of verdict values that indicate a passing evaluation, or a single explicit JSON Schema object (identified by a $schema field) for threshold-style criteria (e.g. a minimum score).

[!IMPORTANT]

  • template is independently optional: you can customize it without also overriding verdict-format/passing-verdicts, and vice versa.
  • verdict-format and passing-verdicts must either both be specified together, or both left unset to fall back to the built-in defaults; specifying only one leaves the other ambiguous and is rejected.
  • When passing-verdicts is a set of literal value(s) (the common case), each value must conform to the verdict-format structure.
  • When passing-verdicts is an explicit JSON Schema, it is validated as a standalone schema and matched directly against the judge's raw verdict, without any case/whitespace normalization.

[!TIP] The following template variables are available for judge prompts:

  • {{.OriginalTask.Prompt}}: The original task prompt
  • {{.OriginalTask.ResponseResultFormat}}: Format instruction from the task
  • {{.OriginalTask.ExpectedResults}}: Array of expected answers
  • {{.Candidate.Response}}: The model's response being evaluated
  • {{.Rules.CaseSensitive}}: Boolean case-sensitive validation flag
  • {{.Rules.IgnoreWhitespace}}: Boolean ignore whitespace flag
  • {{.Rules.TrimLines}}: Boolean trim lines flag
  • {{.Verdict.Format}}: The resolved verdict format, rendered as plain text or pretty-printed JSON schema

A sample task from tasks.yaml:

# tasks.yaml
task-config:
  disabled: true
  file-options:
    image-detail: high
  system-prompt:
    enable-for: "text"
    template: |
      Provide the final answer in exactly this format: {{.ResponseResultFormat}}
      Treat every substring enclosed in `<` and `>` as a variable placeholder.
      Substitute only the raw value in place of `<variable name>`, removing the `<` and `>` characters.
      Do not add any extra words, punctuation, quotes, or whitespace beyond what the format string shows.
  validation-rules:
    case-sensitive: false
    ignore-whitespace: false
  tasks:
    - name: "riddle - split words - v1"
      disabled: false
      prompt: |-
        There are four 8-letter words (animals) that have been split into 2-letter pieces.
        Find these four words by putting appropriate pieces back together:

        RR TE KA DG EH AN SQ EL UI OO HE LO AR PE NG OG
      response-result-format: |-
        list of words in alphabetical order separated by ", "
      system-prompt:
        template: "Provide the final answer in exactly this format: {{.ResponseResultFormat}}"
      expected-result: |-
        ANTELOPE, HEDGEHOG, KANGAROO, SQUIRREL
    - name: "visual - shapes - v1"
      prompt: |-
        The attached picture contains various shapes marked by letters.
        It also contains a set of same shapes that have been rotated marked by numbers.
        Your task is to find all matching pairs.
      response-result-format: |-
        <shape number>: <shape letter> pairs separated by ", " and ordered by shape number
      expected-result: |-
        1: G, 2: F, 3: B, 4: A, 5: C, 6: D, 7: E
      validation-rules:
        ignore-whitespace: true
      files:
        - name: "picture"
          uri: "./taskdata/visual-shapes-v1.png"
          type: "image/png"
          options:
            image-detail: original
    - name: "riddle - anagram - v3"
      prompt: |-
        Two words (each individual word is a fruit) have been combined and their letters arranged in alphabetical order forming a single group.
        Find the original words for each of these 2 groups:

        1. AACEEGHPPR
        2. ACEILMNOOPRT
      response-result-format: |-
        1. <word>, <word>
        2. <word>, <word>
        (words in each group must be alphabetically ordered)
      expected-result:
        - |
          1. GRAPE, PEACH
          2. APRICOT, MELON
        - |
          1. GRAPE, PEACH
          2. APRICOT, LEMON
    - name: "chemistry - observable phenomena - v1"
      disabled: true
      prompt: |-
        What are the primary observable results of mixing household vinegar (an aqueous solution of acetic acid, $CH_3COOH$) with baking soda (sodium bicarbonate, $NaHCO_3$)?
      response-result-format: |-
        Provide a bulleted list of the main, directly observable phenomena. Focus on what one would see and hear. Do not include the chemical equation.
      expected-result: |-
        The response must correctly identify the two main observable results of the chemical reaction.
        Crucially, it must mention the production of a gas, described as fizzing, bubbling, or effervescence.
        It should also note that the solid baking soda dissolves or disappears as it reacts with the vinegar.
      validation-rules:
        judge:
          enabled: true
          name: "mistral-judge"
          variant: "reasoning"
          # Uses default judge prompt configuration for semantic evaluation.
    - name: "code quality - custom judge"
      prompt: |-
        Write a Python function that finds the maximum value in a list.
      response-result-format: |-
        complete Python function with proper naming and structure
      expected-result: |-
        A well-written Python function that correctly finds the maximum value with good practices
      validation-rules:
        judge:
          enabled: true
          name: "mistral-judge"
          variant: "reasoning"
          prompt:
            template: |-
              Evaluate this Python code for both correctness and quality:
              {{.Candidate.Response}}

              Criteria: 1) Correctly finds max value, 2) Proper function name/structure, 3) Handles edge cases, 4) Good Python style
            verdict-format:
              type: object
              properties:
                quality_score:
                  type: string
                  enum: ["excellent", "good", "poor"]
              required: ["quality_score"]
              additionalProperties: false
            passing-verdicts:
              - quality_score: "excellent"
              - quality_score: "good"
    - name: "essay quality - score threshold"
      prompt: |-
        Write a short essay explaining photosynthesis for a middle school audience.
      response-result-format: |-
        a clear, well-organized short essay
      expected-result: |-
        A clear and accurate explanation of photosynthesis appropriate for the target audience
      validation-rules:
        judge:
          enabled: true
          name: "mistral-judge"
          variant: "reasoning"
          prompt:
            template: |-
              Score this essay from 0 to 100 for clarity, accuracy, and audience appropriateness:
              {{.Candidate.Response}}
            verdict-format:
              type: object
              properties:
                score:
                  type: integer
                  minimum: 0
                  maximum: 100
              required: ["score"]
              additionalProperties: false
            passing-verdicts:
              # An explicit JSON Schema (identified by "$schema") is matched directly against
              # the judge's raw verdict, enabling threshold-style criteria.
              $schema: "https://json-schema.org/draft/2020-12/schema"
              type: object
              properties:
                score:
                  type: integer
                  exclusiveMinimum: 80
              required: ["score"]
    - name: "structured response - log parsing"
      prompt: |-
        Parse the following log lines and extract the timestamp, log level, and message for each. If a user ID is present, extract that as well.
        Log lines:
        [2025-09-14 10:30:00] INFO: User 'admin' logged in successfully.
        [2025-09-14 10:31:15] WARN: System memory usage is high.
      response-result-format:
        type: array
        items:
          type: object
          additionalProperties: false
          properties:
            timestamp:
              type: string
              format: "date-time"
            level:
              type: string
              enum: ["INFO", "WARN", "ERROR"]
            message:
              type: string
            user_id:
              type: string
          required: ["timestamp", "level", "message", "user_id"]
      expected-result:
        - - timestamp: "2025-09-14T10:30:00Z"
            level: "INFO"
            message: "User 'admin' logged in successfully."
            user_id: "admin"
          - timestamp: "2025-09-14T10:31:15Z"
            level: "WARN"
            message: "System memory usage is high."
            user_id: ""

Tools

MindTrial supports tool use for tasks, allowing AI models to execute external tools during task solving. Tools are executed in sandboxed Docker containers with resource limits and network isolation.

Tool Definitions

Tools must be defined in config.yaml under the tools section. Each tool defines how to execute a specific capability:

  • name: A unique name for the tool.
  • image: Docker image to use for the tool execution.
  • description: A detailed description of what the tool does and how to use it. This description is provided to the LLM to help it understand when and how to use the tool. Be specific and avoid ambiguity to help the LLM choose the correct tool and provide appropriate parameters.
  • parameters: JSON schema defining the tool's input parameters. The LLM will generate the actual parameter values based on this schema. Provide comprehensive descriptions that explain parameter purpose and format.
  • parameter-files: Mapping of parameter names to container file paths where argument values should be written. Argument values are converted to strings, non-string values are marshaled to JSON. The tool's command should read these files as needed.
  • auxiliary-dir: Directory path inside the container where task files with local access will be automatically mounted. If specified, files with local access attached to the task will be mounted to this directory using each file's unique reference name exactly as provided. Files in this directory are reset between tool calls.
  • shared-dir: Directory path inside the container that persists across all tool calls within a single task. If specified, files created in this directory will be available for any subsequent tool calls but will be removed when the task completes.
  • command: Command to run inside the container. The standard output of the command execution is captured and passed back to the LLM as is.
  • env: Environment variables to set in the container.
  • dependencies: Task services the tool can access, each with a service name and optional env templates.

[!IMPORTANT] Tool use requires Docker to be installed and running on the system. Tools are executed in isolated containers with no network access, unless they depend on task services.

Example tool definition in config.yaml:

config:
  tools:
    - name: python-code-executor
      image: python:latest
      description: |
        Executes Python 3 code in a secure, sandboxed environment to perform calculations, data manipulation, or algorithmic tasks.
        IMPORTANT:
        - Only the Python standard library is available. No third-party packages (like pandas or numpy) can be imported.
        - The environment has no network access.
        - Task files made available to this tool are mounted under /app/data/ using their [file: filename] names.
        - Use standard file operations like open('/app/data/filename', 'r') to read attached files, where 'filename' matches the name shown in [file: filename] references.
        - A persistent shared directory is available at /app/shared/ that persists across ALL tool calls within the same task (regardless of which tool is being called). Files created in this directory will be available in any subsequent tool call.
        - Any files or changes outside of /app/shared/ are ephemeral and will be reset between tool calls.
        - The code must print its final result to standard output to be returned.
      parameters:
        type: object
        properties:
          code:
            type: string
            description: "A string containing a self-contained Python 3 script. The script must use the `print()` function to return a final result. Example: `print(sum([i for i in range(101) if i % 2 == 0]))`. To read attached files, use open('/app/data/filename', 'r') where 'filename' matches what appears in [file: filename] references."
        required:
          - code
        additionalProperties: false
      parameter-files:
        code: /app/main.py
      auxiliary-dir: /app/data
      shared-dir: /app/shared
      command:
        - python
        - /app/main.py
      env:
        PYTHONIOENCODING: "UTF-8"
        PYTHONUNBUFFERED: "1"
        PYTHONHASHSEED: "847629"
Tool Selection

You can configure tool selection globally for all tasks in the task-config section, and override it for individual tasks if needed. Tools must be defined in config.yaml first.

  • tool-selector: Configuration for tool availability during task execution.
    • disabled: If true, no tools are available for tasks (default: false).
    • tools: List of tools to make available, with per-tool limits.
      • name: Name of the tool as defined in config.yaml.
      • disabled: If true, this tool is not available (default: false).
      • max-calls: Maximum number of times this tool can be called per task (optional).
      • timeout: Maximum execution time per tool call (e.g., 60s, optional).
      • max-memory-mb: Maximum memory usage in MB per tool call (optional).
      • cpu-percent: Maximum CPU usage as percentage per tool call (optional).
    • service-inputs: Inputs for the task services that the task uses, keyed by service name and then by input name (optional). Inputs set in task-config apply to every task, and a task can override single inputs. Inputs only take effect for tasks that use the service.

Example tool configuration in tasks.yaml:

task-config:
  tool-selector:
    disabled: false
    tools:
      - name: python-code-executor
        disabled: false
        max-calls: 10
        timeout: 60s
        max-memory-mb: 512
        cpu-percent: 25
  tasks:
    - name: "math calculation"
      prompt: "Calculate the sum of even numbers from 1 to 100."
      response-result-format: "single number"
      expected-result: "2550"
      # Inherits global tool-selector configuration.
    - name: "simple math"
      prompt: "What is 2 + 2?"
      response-result-format: "single number"
      expected-result: "4"
      tool-selector:
        tools:
          - name: python-code-executor
            disabled: true  # Selectively disable tool for this simple task.
Conversation Turn Limit

You can set a maximum number of conversation turns per task to act as a safety net against infinite conversation loops (e.g., when a model repeatedly requests exhausted tools). The limit can be configured globally in the task-config section, and overridden for individual tasks if needed. A value of 0 means unlimited.

  • max-turns: Maximum number of conversation turns allowed per task (default: 0, unlimited).

Example configuration in tasks.yaml:

task-config:
  max-turns: 100  # Default limit for all tasks.
  tasks:
    - name: "trivia - geography - Asia"
      prompt: "What is the capital of Japan?"
      response-result-format: "city name"
      expected-result: "Tokyo"
      max-turns: 200  # Override: allow more turns for this task.
    - name: "trivia - geography"
      prompt: "What is the capital of Australia?"
      response-result-format: "city name"
      expected-result: "Canberra"
      # Inherits the global limit of 100 turns.

Task Services

A task service is a Docker container that keeps state while the model works on a task, such as a simulated environment, a database, or a web shop. Services are used together with:

  • tools, which let the model read and change the service state, and
  • custom validators, which can check the final service state after the model answers.

Services and custom validators are independent of each other. Tools that use services work with any validation method (exact match, schema, judge, or custom validator), and a custom validator does not need any services.

[!IMPORTANT] Task services require Docker Engine 25.0 or newer (Engine API 1.44+). Before an evaluation that uses services starts, MindTrial checks the Docker daemon and stops with an error if it is too old.

Services are defined in config.yaml under the services section:

  • name: A unique name for the service, used in dependencies and service-inputs.
  • image: Docker image used to run the service.
  • command: Command overriding the image's default command (optional).
  • env: Environment variables set for every instance of the service (optional).
  • endpoint: Where tools and validators connect to the service.
    • port: Container port on which the service accepts connections.
    • scheme: URL scheme, http (default) or https.
  • input-env: Inputs that tasks can set, mapping each input name to the environment variable that receives its value in the service container (optional). An input must not set a different value for a variable that env already defines.
  • healthcheck: Command that checks whether the service is ready, run inside the service container without a shell (optional). If set, the service is ready once the command succeeds. If not set, the service is ready as soon as its container is running.
  • startup-timeout: Maximum time a service with a healthcheck may take to become ready (optional, default 15s). The check is repeated until it succeeds or this time runs out, and a single check may also take up to this long.
  • max-memory-mb: Maximum memory available to the service container in MB (optional).
  • cpu-percent: Maximum CPU usage as a percentage of total host CPU (optional).

A tool or custom validator gets access to a service by listing it in dependencies:

  • dependencies: List of services the tool or validator can access.
    • service: Name of the service.
    • env: Environment variables set in the tool or validator container (optional). Each value is a template filled in from the running service, for example CART_URL: "{{ .Endpoint }}".

A task sets service inputs with service-inputs in its tool selector:

tool-selector:
  service-inputs:
    cart-service:        # service name
      customer_id: "42"  # input name declared in the service's input-env
  • Input values must be strings, numbers, or booleans.
  • String values are templates, so a task can derive an input from the evaluation seed, e.g. "{{ hash .Evaluation.Seed .Task.Name }}". MindTrial generates a new evaluation seed for each evaluation, writes it to the log, and records it in every result (Evaluation.Seed in the JSON output); pass the same seed with --evaluation-seed to get the same inputs again.
  • service-inputs set in task-config apply to every task, and a task can override single inputs. Inputs only take effect for tasks that use the service.
  • When an input is not set, the service uses its own default.

How services run:

  • Which services start: The services that the task's enabled tools and its custom validator depend on. Tasks that need no services start none.
  • When services start: Before the model receives the prompt. Every service must be ready within its startup-timeout (15 seconds by default); otherwise the attempt fails and its services are removed.
  • One set of services per attempt: An attempt is one try of one task by one run configuration. Each attempt gets its own service instances, so runs never share state. A retry is a new attempt with fresh services started from the same inputs. If a service does not derive its initial state from its inputs (for example, from a seed input), a retry may start from a different state.
  • Validation: After a successful attempt, a custom validator that depends on a service connects to the same instance that the model's tools used. The services are removed after validation.
  • Network isolation: Each service has its own internal Docker network without internet access. A tool or validator joins only the networks of the services in its dependencies, and one without dependencies has no network access. Services cannot reach each other.
  • Service failure: A service that stops during an attempt is not restarted, because that would silently reset its state. The next tool call or validation that needs it fails, and the task result is an error.

Custom Validators

A custom validator is a Docker container that you provide to decide whether the model's answer is correct. Use it when an answer can only be checked by running code, for example to run generated code against hidden tests or to check the final state of a simulated environment.

A custom validator does not need task services. Without dependencies, it runs in an isolated container with no network access. With dependencies, it can inspect the same service instances that the model's tools used.

How it works: After a successful model attempt, MindTrial runs the validator container once. It passes the task and the model's answer to the validator through templated command arguments, environment variables, and files. The validator prints its verdict as JSON on standard output.

[!IMPORTANT] Custom validators are trusted: they are part of your evaluation setup, not tools for the model, and they decide task outcomes. Only use validator images you control.

Validators are defined in config.yaml under the validators section:

  • name: A unique name for the validator, used by custom-validator in task validation rules.
  • image: Docker image used to run the validator.
  • command: Command overriding the image's default command (optional). Each argument is a template and is passed directly to Docker, without a shell.
  • env: Environment variables set in the validator container (optional). Each value is a template.
  • template-files: Files created for each validation and mounted read-only into the validator container (optional).
    • path: Absolute path of the file inside the container. Each path must be unique.
    • template: Template that produces the file content.
  • dependencies: Task services the validator can access (optional).
  • timeout: Maximum duration of one validator run (e.g., 60s). If not set, there is no timeout.
  • max-memory-mb: Maximum memory available to the validator container in MB (optional).
  • cpu-percent: Maximum CPU usage as a percentage of total host CPU (optional).

[!TIP] Pass the model's answer through template-files. Use command and env only for short values you control, such as names, seeds, and flags. The model's answer can be of any size and may contain characters that command arguments and environment variables cannot carry; if the validator cannot start, the task result is an error rather than a failed answer.

A task selects a validator with the custom-validator validation rule:

  • response-result-format is still required: it tells the model how to write its final answer.
  • expected-result is optional. When set, it is passed to the validator unchanged as reference data and does not need to match response-result-format.
  • case-sensitive, ignore-whitespace, and trim-lines are passed to the validator, which decides whether to use them. MindTrial does not change the answer.

Example of a validator that runs hidden tests without any services:

# config.yaml
config:
  validators:
    - name: hidden-tests
      image: example/code-checker:latest
      command: [check, --solution, /input/solution.py]
      template-files:
        - path: /input/solution.py
          template: "{{ .Candidate.Response }}"
      timeout: 60s
# tasks.yaml
task-config:
  tasks:
    - name: fizzbuzz
      prompt: "Write a Python function fizzbuzz(n) that returns the FizzBuzz sequence from 1 to n."
      response-result-format: "complete Python source code"
      validation-rules:
        custom-validator: hidden-tests

The validator must exit with code 0 and print exactly one JSON object to standard output, for example:

{"correct": false, "title": "2 of 5 tests failed", "explanation": "fizzbuzz(15) returned '15' instead of 'FizzBuzz'."}

The object must match this schema:

{
  "type": "object",
  "properties": {
    "correct": {"type": "boolean"},
    "title": {"type": "string", "pattern": "\\S"},
    "explanation": {"type": "string", "pattern": "\\S"}
  },
  "required": ["correct", "title", "explanation"],
  "additionalProperties": false
}
  • correct: true marks the answer as correct, and correct: false marks it as failed.
  • title and explanation must contain non-whitespace text.
  • Anything else makes the task result an error, not a failed answer: a non-zero exit code, a timeout, a Docker error, missing or unknown fields, or any output after the JSON object.

Results checked by a custom validator record custom as their validation method. Failed answers are shown exactly as the model returned them, without a comparison against expected-result. Validator runs are not counted as model tool calls.

Templates

Service inputs, dependency environment variables, and custom validator settings use Go template syntax. MindTrial checks template syntax before any task runs, and using a field that does not exist is an error.

Each kind of template has its own fields:

TemplateAvailable fields
service-inputs string values.Evaluation.Seed, .Task.Name, .Provider.Name, .Run.Name
dependencies[].env values.Name, .Host, .Port, .Endpoint
Validator command, env, and template-files.OriginalTask.*, .Candidate.Response, .Rules.*, .Evaluation.Seed, .Task.Name, .Provider.Name, .Run.Name

Fields:

  • {{ .Evaluation.Seed }}: Evaluation seed shared by all providers, runs, tasks, and attempts of one evaluation.
  • {{ .Task.Name }}, {{ .Provider.Name }}, {{ .Run.Name }}: Names of the task, the provider, and the run configuration.
  • {{ .Name }}: Name of the service.
  • {{ .Host }}: Network host name of the service. MindTrial generates host names, so use .Host or .Endpoint instead of the service name.
  • {{ .Port }}: Endpoint port of the service.
  • {{ .Endpoint }}: Endpoint URL of the service, e.g. http://<host>:8080.
  • {{ .OriginalTask.Prompt }}: The task prompt.
  • {{ .OriginalTask.ResponseResultFormat }}: The task's response-result-format.
  • {{ .OriginalTask.ExpectedResults }}: List of the task's expected results. It is an empty list (not null) when the task has no expected-result.
  • {{ .Candidate.Response }}: The model's final answer. For structured response formats, this is a structured value.
  • {{ .Rules.CaseSensitive }}, {{ .Rules.IgnoreWhitespace }}, {{ .Rules.TrimLines }}: The task's validation flags.

The validator fields use the same names as judge prompt templates.

Helper functions:

  • hash: Returns a stable unsigned 64-bit number derived from all of its arguments, e.g. {{ hash .Evaluation.Seed .Task.Name }}.
  • json: Encodes its argument as JSON, e.g. {{ json .OriginalTask.ExpectedResults }}.

[!IMPORTANT] Go prints structured values in its own format, which is not JSON. Use json to pass structured data, e.g. {{ json .Candidate.Response }}.

Example: Stateful Shopping Cart

In this example, the model uses the cart tool to change a shopping cart held by the cart-service service. After the model answers, the cart-state validator checks the same cart:

# config.yaml
config:
  services:
    - name: cart-service
      image: example/cart-service:latest
      command: [cart-server]
      env:
        CART_CURRENCY: CAD
      endpoint:
        port: 8080
      input-env:
        customer_id: CART_CUSTOMER_ID
      healthcheck: [cart-server, --check]

  tools:
    - name: cart
      image: example/cart-client:latest
      command: [cart-client, --request, /input/request.json]
      description: >
        Modify the current customer's shopping cart by adding or removing items.
      parameters:
        type: object
        properties:
          request:
            type: object
            properties:
              action:
                type: string
                enum: [ADD, REMOVE]
                description: Operation to perform on the cart.
              item:
                type: string
                enum: [apple, orange, bread]
                description: Item to add or remove.
              quantity:
                type: integer
                minimum: 1
                description: Number of items to add or remove.
            required: [action, item, quantity]
            additionalProperties: false
        required: [request]
        additionalProperties: false
      parameter-files:
        request: /input/request.json
      dependencies:
        - service: cart-service
          env:
            CART_URL: "{{ .Endpoint }}"

  validators:
    - name: cart-state
      image: example/cart-validator:latest
      command: [cart-validator, --expected, /input/expected.json]
      template-files:
        - path: /input/expected.json
          template: "{{ json .OriginalTask.ExpectedResults }}"
      timeout: 60s
      dependencies:
        - service: cart-service
          env:
            CART_URL: "{{ .Endpoint }}"

CART_CURRENCY is a fixed service setting. customer_id is an input that each task can set; the service receives it as CART_CUSTOMER_ID. The task below derives the customer ID from the evaluation seed, so all providers and runs in one evaluation start with the same customer. Passing the same --evaluation-seed again gives the same customer:

# tasks.yaml
task-config:
  tasks:
    - name: update-shopping-cart
      prompt: >
        Use the cart tool to add two apples and one loaf of bread.
      response-result-format: "short summary of the final cart contents"
      expected-result:
        items:
          apple: 2
          bread: 1
      tool-selector:
        tools:
          - name: cart
        service-inputs:
          cart-service:
            customer_id: "{{ hash .Evaluation.Seed .Task.Name }}"
      validation-rules:
        custom-validator: cart-state

The expected-result object is reference data for the validator; it does not have to match the plain-text response-result-format. The validator reads it from /input/expected.json as [{"items":{"apple":2,"bread":1}}] and compares it with the cart state it gets from CART_URL.

Command Reference

mindtrial [options] [command]

Commands:
  run                       Start the trials
  merge-results             Merge results from multiple runs
  stats                     Compute derived statistics from result files
  help                      Show help
  version                   Show version

Options:
  --config string           Configuration file path (default: config.yaml)
  --tasks string            Task definitions file path
  --output-dir string       Results output directory
  --output-basename string  Base filename for results; replace if exists; blank = stdout
  --html                    Generate HTML output (default: true)
  --csv                     Generate CSV output (default: false)
  --json                    Generate JSON output (default: false)
  --input string            Input result file path for merge-results/stats; can be specified multiple times
  --log string              Log file path; append if exists; blank = stdout
  --verbose                 Enable detailed logging
  --debug                   Enable low-level debug logging (implies --verbose)
  --interactive             Enable interactive interface for run configuration, and real-time progress monitoring (default: false)
  --evaluation-seed string  Seed for reproducible evaluation behavior (e.g., derived task service inputs); generated for each evaluation when omitted
  --group-by string         Comma-separated stats grouping dimensions: provider, run, model, suite, category, difficulty, tag (default: provider,run)
  --stats-format string     Stats output format: text, csv, json, or jsonl (default: text)
  --provider string         Filter stats to this provider; can be specified multiple times
  --run string              Filter stats to this run configuration; can be specified multiple times
  --model string            Filter stats to this model; can be specified multiple times
  --suite string            Filter stats to this task suite; can be specified multiple times
  --category string         Filter stats to this task category; can be specified multiple times
  --difficulty string       Filter stats to this task difficulty; can be specified multiple times
  --status string           Filter stats to this result status (passed, failed, error, skipped); can be specified multiple times
  --tag string              Filter stats to results tagged with this value; can be specified multiple times
  --tag-mode string         How multiple --tag filters combine for stats: all or any (default: all)

Contributing

Contributions are welcome! Please review our CONTRIBUTING.md guidelines for more details.

Getting the Source Code

Clone the repository and install dependencies:

git clone https://github.com/petmal/mindtrial.git
cd mindtrial
go mod download

Running Tests

Execute the unit tests with:

go test -tags=test -race -v ./...

Project Details

/
├── cmd/
│   └── mindtrial/       # Command-line interface and main entry point
│       └── tui/         # Terminal-based UI and interactive mode functionality
├── config/              # Data models and management for configuration and task definitions
├── formatters/          # Output formatting for results
├── pkg/                 # Shared packages and utilities
├── providers/           # AI model service provider connectors
│   ├── execution/       # Provider run execution utilities and coordination
│   └── tools/           # Execution engine for external tools used by models
├── runners/             # Task execution and result aggregation
├── stats/               # Derived statistics (filtering, grouping, aggregation) for the stats command
├── taskdata/            # Auxiliary files referenced by tasks in tasks.yaml
├── validators/          # Result validation logic
└── version/             # Application metadata

License

This project is licensed under the Mozilla Public License 2.0 - see the LICENSE file for details.

ai-benchmark
ai-evaluation-tools
ai-model-comparison
ai-tool
anthropic
artificial-intelligence-projects
deepseek
google-gemini-ai
grok-ai
language-models-ai
llm-benchmarking
llm-comparison
llm-evaluation-framework
mistral-ai
moonshot-ai
openai
openrouter
opensource
qwen
xai

Significant stargazers

mitchell

88 followers · starred May 2026

Languages

Go

93.1%

Go Template

6.9%