Does your OpenAI-compatible LLM endpoint actually hoour the parameters you send? Reasoning switches, effrt, caching, model substitution.
TypeScript
0
1 commits
updated Oct 6, 2026
Does this OpenAI-compatible endpoint actually honour what I send? Send controlled request variants and see what really changed.
With LLM providers and gateways, almost every misconfiguration returns HTTP 200. The reply looks normal, the parameter did nothing, and the only symptom shows up days later as changed behaviour or a larger bill. The classic silent failures:
low, medium and high produce the same output length, or one spelling multiplies output tokens several times over.extra_body sent as a literal field. It is a Python SDK argument that the client merges into the body root before sending. On the wire it is just an unknown key.llm-wire-check uses Node's native fetch with no SDK, so every parameter reaches the wire exactly as written, and measures the response rather than trusting any flag the provider reports about itself.
Requires Node 20 or newer. No runtime dependencies.
git clone https://github.com/danidante/llm-wire-check.git
cd llm-wire-check
npm install
npx tsx src/cli.ts reasoning --base-url https://api.example.com/v1 --model example-model
The examples below use npx llm-wire-check; from a clone, run npx tsx src/cli.ts instead (or npm run build && npm link once).
The API key is read from an environment variable (default OPENAI_API_KEY) and never printed:
export MY_PROVIDER_KEY=...
npx llm-wire-check reasoning \
--base-url https://gateway.example.net/v1 \
--model example-model \
--api-key-env MY_PROVIDER_KEY \
--served-model-header x-served-model \
--effort-sweep
# only two arms, streamed, five repetitions, JSON for CI
npx llm-wire-check reasoning --base-url https://api.example.com/v1 --model example-model \
--arms none,template_thinking --stream --reps 5 --json > reasoning.json
# does the provider cache a long prefix, and does prompt_cache_key change anything?
npx llm-wire-check cache --base-url https://api.example.com/v1 --model example-model --compare-key
Common options:
| Option | Default | Meaning |
|---|---|---|
--base-url URL | required | Endpoint base; /chat/completions is appended |
--model ID | required | Model id to request |
--api-key-env NAME | OPENAI_API_KEY | Env var holding the key (sent as Authorization: Bearer) |
--header "K: V" | Extra header, repeatable; an explicit Authorization header wins | |
--reps N | 3 | Repetitions per arm |
--prompt-tokens N | 8000 | Synthetic prompt size |
--max-tokens N | 3000 | Output budget per call; warns below 2500 |
--max-tokens-field NAME | max_tokens | e.g. max_completion_tokens for endpoints that require it |
--timeout-ms N | 180000 | Per-call timeout |
--served-model-header NAME | Response header naming the served model, repeatable | |
--stream | off | SSE streaming (reasoning command) |
--json | off | One JSON object on stdout; progress stays on stderr |
Exit codes: 0 no findings, 1 findings, 2 usage error or nothing reached the endpoint.
Each arm is one request body variant. Every arm is sent --reps times, and the arm order rotates on each repetition so no arm always runs first (provider load drifts over minutes).
| Arm | Adds to the body | What it tests |
|---|---|---|
none | nothing | Baseline. Always included. |
effort | reasoning_effort: "high" | Top-level effort field |
template_thinking | chat_template_kwargs: { thinking: true } | Chat-template switch, thinking spelling |
template_enable_thinking | chat_template_kwargs: { enable_thinking: true } | Chat-template switch, enable_thinking spelling |
template_effort | chat_template_kwargs: { thinking: true, reasoning_effort: "high" } | Effort nested inside the template kwargs |
extra_body_literal | extra_body: { chat_template_kwargs: { enable_thinking: true } } | The classic mistake: a literal extra_body key on the wire |
effort_low, effort_medium, effort_high | reasoning_effort: "low" / "medium" / "high" | Added by --effort-sweep: does effort scale anything? |
--arms a,b,c selects a subset; none is added automatically. Note that effort and effort_high send the same body.
Reasoning text is detected in message.reasoning_content, message.reasoning (string), message.reasoning_details[].text, and inline <think>...</think> in the content (reported as inline_think, and stripped before counting content). A stray </think> without an opener (the template opened it in the prompt) and an unclosed <think> (cut off mid-thought) count as inline too. In --stream mode the same fields are read from delta, with stream_options: { include_usage: true } sent so usage arrives.
| Code | Exit code | Meaning |
|---|---|---|
SWITCH_WORKS:<arm> | info | The arm reasons and the none baseline does not |
SWITCH_IGNORED:<arm> | finding | The arm sends a switch, the call succeeds, and reasoning stays at zero like the baseline |
REASONS_UNCONDITIONALLY | info | The baseline already reasons, so switch presence cannot be attributed (switch verdicts are skipped) |
ARM_REJECTED:<arm> | info | Every call of the arm got a non-2xx. A 400 is information: the endpoint validates the field |
BASELINE_FAILED | finding | The none arm never succeeded; switch verdicts unavailable |
EFFORT_IGNORED | finding | Sweep medians for low / medium / high are within 15% of each other (reasoning chars, or completion tokens when no reasoning text is exposed) |
OUTPUT_INFLATION:<arm> | finding | Median completion tokens are at least 2x the lightest arm that also reasons |
READING_UNRELIABLE:<arm> | finding | At least one call hit the output budget; its reading is an artifact |
REASONING_TOKENS_UNREPORTED | finding | Reasoning text came back but usage.completion_tokens_details.reasoning_tokens was missing or 0 |
MODEL_SUBSTITUTED | finding | A --served-model-header named a different model (a vendor/ prefix alone does not count) |
CACHED | info | Call 2 or 3 read more than 50% of the prompt from cache |
PARTIAL_CACHE | info | Between 5% and 50% read from cache |
NOT_CACHED | finding | A cached-token field is present and reports about 0 on repeat calls |
UNVERIFIABLE | finding | No cached-token field at all. This means "cannot verify", not "no cache" |
COST_IDENTICAL | finding | usage.cost is identical on all 3 calls of a large prompt; a working cache normally shows in the price |
WARM_BEFORE_TEST | info | Call 1 was already cached |
CACHE_KEY_NO_EFFECT | info | With --compare-key: the result is the same with and without the key |
CACHE_CALLS_FAILED | finding | Repeat calls failed, nothing to compare |
With --compare-key, cache codes carry a :with_key or :no_key suffix. Cached tokens are read from usage.prompt_tokens_details.cached_tokens, then usage.cached_tokens, then any numeric usage key matching /cach/i (cache reads preferred over cache writes and misses).
--prompt-tokens to match your production prompts.max_tokens budget, leaving empty content that looks like a failure. Any call with completion_tokens >= max_tokens or finish_reason == "length" is flagged TRUNC and its reading is treated as an artifact. That equality is the tell. Keep --max-tokens at 2500 or more.reasoning_tokens. Many providers omit it, or report 0 while returning thousands of characters of reasoning. The tool counts reasoning characters and shows chars / 3.9 labelled as an estimate.--compare-key series do not share a prefix.Example only: synthetic numbers against an imaginary endpoint, trimmed to three arms.
llm-wire-check reasoning model=example-model reps=3 prompt~8000 tok max_tokens=3000
ARM none
rep http ms r_chars r_source est_rtok rep_rtok compl prompt content finish flags
--- ---- ---- ------- -------- -------- -------- ----- ------ ------- ------ -----
1 200 2140 0 none 0 - 38 8041 151 stop -
2 200 1985 0 none 0 - 41 8041 163 stop -
3 200 2310 0 none 0 - 36 8041 144 stop -
summary: ok 3/3 r_chars median 0 (min 0, max 0) compl median 38 latency median 2140ms truncated 0 empty 0
ARM template_thinking
rep http ms r_chars r_source est_rtok rep_rtok compl prompt content finish flags
--- ---- ---- ------- ----------------- -------- -------- ----- ------ ------- ------ -----
1 200 6420 3074 reasoning_content 788 0 826 8041 148 stop -
2 200 7015 3219 reasoning_content 825 0 867 8041 159 stop -
3 200 5890 2494 reasoning_content 639 0 679 8041 152 stop -
summary: ok 3/3 r_chars median 3074 (min 2494, max 3219) compl median 826 latency median 6420ms truncated 0 empty 0
ARM extra_body_literal
rep http ms r_chars r_source est_rtok rep_rtok compl prompt content finish flags
--- ---- ---- ------- -------- -------- -------- ----- ------ ------- ------ -----
1 400 212 0 none 0 - - - 0 - -
2 400 198 0 none 0 - - - 0 - -
3 400 205 0 none 0 - - - 0 - -
rep 1 body: {"error":{"message":"invalid request: unknown field \"extra_body\""}}
...
summary: ok 0/3 r_chars median 0 (min 0, max 0) compl median - latency median - truncated 0 empty 0
FINDINGS
[info] ARM_REJECTED:extra_body_literal
every call returned HTTP 400; see the body snippet
[info] SWITCH_WORKS:template_thinking
median 3074 reasoning chars vs 0 with no params
[FIND] REASONING_TOKENS_UNREPORTED
3 call(s) returned reasoning text but usage reasoning_tokens was missing or 0 (arms: template_thinking); do not use that field as proof
Each run ends with a short READING THIS legend explaining every column.
Each call sends the full synthetic prompt. The reasoning command makes arms x reps calls (18 with the defaults, about 144,000 input tokens at 8,000 tokens each, plus up to 3,000 output tokens per call). The cache command makes 3 calls, or 6 with --compare-key, at max_tokens 16. The tool prints this estimate before it starts. Some gateways reserve input plus max_tokens per call up front, so a low balance can fail early even when actual spend is small. Use --arms, --reps and --prompt-tokens to keep runs cheap.
npm test # node:test via tsx, includes an end-to-end run against a local mock endpoint
npm run typecheck
npm run build # emits dist/, used by the bin entry
MIT, Copyright (c) 2026 Daniel Valle. See LICENSE.
Does your OpenAI-compatible LLM endpoint actually hoour the parameters you send? Reasoning switches, effrt, caching, model substitution.
TypeScript
0
1 commits
updated Oct 6, 2026
Does this OpenAI-compatible endpoint actually honour what I send? Send controlled request variants and see what really changed.
With LLM providers and gateways, almost every misconfiguration returns HTTP 200. The reply looks normal, the parameter did nothing, and the only symptom shows up days later as changed behaviour or a larger bill. The classic silent failures:
low, medium and high produce the same output length, or one spelling multiplies output tokens several times over.extra_body sent as a literal field. It is a Python SDK argument that the client merges into the body root before sending. On the wire it is just an unknown key.llm-wire-check uses Node's native fetch with no SDK, so every parameter reaches the wire exactly as written, and measures the response rather than trusting any flag the provider reports about itself.
Requires Node 20 or newer. No runtime dependencies.
git clone https://github.com/danidante/llm-wire-check.git
cd llm-wire-check
npm install
npx tsx src/cli.ts reasoning --base-url https://api.example.com/v1 --model example-model
The examples below use npx llm-wire-check; from a clone, run npx tsx src/cli.ts instead (or npm run build && npm link once).
The API key is read from an environment variable (default OPENAI_API_KEY) and never printed:
export MY_PROVIDER_KEY=...
npx llm-wire-check reasoning \
--base-url https://gateway.example.net/v1 \
--model example-model \
--api-key-env MY_PROVIDER_KEY \
--served-model-header x-served-model \
--effort-sweep
# only two arms, streamed, five repetitions, JSON for CI
npx llm-wire-check reasoning --base-url https://api.example.com/v1 --model example-model \
--arms none,template_thinking --stream --reps 5 --json > reasoning.json
# does the provider cache a long prefix, and does prompt_cache_key change anything?
npx llm-wire-check cache --base-url https://api.example.com/v1 --model example-model --compare-key
Common options:
| Option | Default | Meaning |
|---|---|---|
--base-url URL | required | Endpoint base; /chat/completions is appended |
--model ID | required | Model id to request |
--api-key-env NAME | OPENAI_API_KEY | Env var holding the key (sent as Authorization: Bearer) |
--header "K: V" | Extra header, repeatable; an explicit Authorization header wins | |
--reps N | 3 | Repetitions per arm |
--prompt-tokens N | 8000 | Synthetic prompt size |
--max-tokens N | 3000 | Output budget per call; warns below 2500 |
--max-tokens-field NAME | max_tokens | e.g. max_completion_tokens for endpoints that require it |
--timeout-ms N | 180000 | Per-call timeout |
--served-model-header NAME | Response header naming the served model, repeatable | |
--stream | off | SSE streaming (reasoning command) |
--json | off | One JSON object on stdout; progress stays on stderr |
Exit codes: 0 no findings, 1 findings, 2 usage error or nothing reached the endpoint.
Each arm is one request body variant. Every arm is sent --reps times, and the arm order rotates on each repetition so no arm always runs first (provider load drifts over minutes).
| Arm | Adds to the body | What it tests |
|---|---|---|
none | nothing | Baseline. Always included. |
effort | reasoning_effort: "high" | Top-level effort field |
template_thinking | chat_template_kwargs: { thinking: true } | Chat-template switch, thinking spelling |
template_enable_thinking | chat_template_kwargs: { enable_thinking: true } | Chat-template switch, enable_thinking spelling |
template_effort | chat_template_kwargs: { thinking: true, reasoning_effort: "high" } | Effort nested inside the template kwargs |
extra_body_literal | extra_body: { chat_template_kwargs: { enable_thinking: true } } | The classic mistake: a literal extra_body key on the wire |
effort_low, effort_medium, effort_high | reasoning_effort: "low" / "medium" / "high" | Added by --effort-sweep: does effort scale anything? |
--arms a,b,c selects a subset; none is added automatically. Note that effort and effort_high send the same body.
Reasoning text is detected in message.reasoning_content, message.reasoning (string), message.reasoning_details[].text, and inline <think>...</think> in the content (reported as inline_think, and stripped before counting content). A stray </think> without an opener (the template opened it in the prompt) and an unclosed <think> (cut off mid-thought) count as inline too. In --stream mode the same fields are read from delta, with stream_options: { include_usage: true } sent so usage arrives.
| Code | Exit code | Meaning |
|---|---|---|
SWITCH_WORKS:<arm> | info | The arm reasons and the none baseline does not |
SWITCH_IGNORED:<arm> | finding | The arm sends a switch, the call succeeds, and reasoning stays at zero like the baseline |
REASONS_UNCONDITIONALLY | info | The baseline already reasons, so switch presence cannot be attributed (switch verdicts are skipped) |
ARM_REJECTED:<arm> | info | Every call of the arm got a non-2xx. A 400 is information: the endpoint validates the field |
BASELINE_FAILED | finding | The none arm never succeeded; switch verdicts unavailable |
EFFORT_IGNORED | finding | Sweep medians for low / medium / high are within 15% of each other (reasoning chars, or completion tokens when no reasoning text is exposed) |
OUTPUT_INFLATION:<arm> | finding | Median completion tokens are at least 2x the lightest arm that also reasons |
READING_UNRELIABLE:<arm> | finding | At least one call hit the output budget; its reading is an artifact |
REASONING_TOKENS_UNREPORTED | finding | Reasoning text came back but usage.completion_tokens_details.reasoning_tokens was missing or 0 |
MODEL_SUBSTITUTED | finding | A --served-model-header named a different model (a vendor/ prefix alone does not count) |
CACHED | info | Call 2 or 3 read more than 50% of the prompt from cache |
PARTIAL_CACHE | info | Between 5% and 50% read from cache |
NOT_CACHED | finding | A cached-token field is present and reports about 0 on repeat calls |
UNVERIFIABLE | finding | No cached-token field at all. This means "cannot verify", not "no cache" |
COST_IDENTICAL | finding | usage.cost is identical on all 3 calls of a large prompt; a working cache normally shows in the price |
WARM_BEFORE_TEST | info | Call 1 was already cached |
CACHE_KEY_NO_EFFECT | info | With --compare-key: the result is the same with and without the key |
CACHE_CALLS_FAILED | finding | Repeat calls failed, nothing to compare |
With --compare-key, cache codes carry a :with_key or :no_key suffix. Cached tokens are read from usage.prompt_tokens_details.cached_tokens, then usage.cached_tokens, then any numeric usage key matching /cach/i (cache reads preferred over cache writes and misses).
--prompt-tokens to match your production prompts.max_tokens budget, leaving empty content that looks like a failure. Any call with completion_tokens >= max_tokens or finish_reason == "length" is flagged TRUNC and its reading is treated as an artifact. That equality is the tell. Keep --max-tokens at 2500 or more.reasoning_tokens. Many providers omit it, or report 0 while returning thousands of characters of reasoning. The tool counts reasoning characters and shows chars / 3.9 labelled as an estimate.--compare-key series do not share a prefix.Example only: synthetic numbers against an imaginary endpoint, trimmed to three arms.
llm-wire-check reasoning model=example-model reps=3 prompt~8000 tok max_tokens=3000
ARM none
rep http ms r_chars r_source est_rtok rep_rtok compl prompt content finish flags
--- ---- ---- ------- -------- -------- -------- ----- ------ ------- ------ -----
1 200 2140 0 none 0 - 38 8041 151 stop -
2 200 1985 0 none 0 - 41 8041 163 stop -
3 200 2310 0 none 0 - 36 8041 144 stop -
summary: ok 3/3 r_chars median 0 (min 0, max 0) compl median 38 latency median 2140ms truncated 0 empty 0
ARM template_thinking
rep http ms r_chars r_source est_rtok rep_rtok compl prompt content finish flags
--- ---- ---- ------- ----------------- -------- -------- ----- ------ ------- ------ -----
1 200 6420 3074 reasoning_content 788 0 826 8041 148 stop -
2 200 7015 3219 reasoning_content 825 0 867 8041 159 stop -
3 200 5890 2494 reasoning_content 639 0 679 8041 152 stop -
summary: ok 3/3 r_chars median 3074 (min 2494, max 3219) compl median 826 latency median 6420ms truncated 0 empty 0
ARM extra_body_literal
rep http ms r_chars r_source est_rtok rep_rtok compl prompt content finish flags
--- ---- ---- ------- -------- -------- -------- ----- ------ ------- ------ -----
1 400 212 0 none 0 - - - 0 - -
2 400 198 0 none 0 - - - 0 - -
3 400 205 0 none 0 - - - 0 - -
rep 1 body: {"error":{"message":"invalid request: unknown field \"extra_body\""}}
...
summary: ok 0/3 r_chars median 0 (min 0, max 0) compl median - latency median - truncated 0 empty 0
FINDINGS
[info] ARM_REJECTED:extra_body_literal
every call returned HTTP 400; see the body snippet
[info] SWITCH_WORKS:template_thinking
median 3074 reasoning chars vs 0 with no params
[FIND] REASONING_TOKENS_UNREPORTED
3 call(s) returned reasoning text but usage reasoning_tokens was missing or 0 (arms: template_thinking); do not use that field as proof
Each run ends with a short READING THIS legend explaining every column.
Each call sends the full synthetic prompt. The reasoning command makes arms x reps calls (18 with the defaults, about 144,000 input tokens at 8,000 tokens each, plus up to 3,000 output tokens per call). The cache command makes 3 calls, or 6 with --compare-key, at max_tokens 16. The tool prints this estimate before it starts. Some gateways reserve input plus max_tokens per call up front, so a low balance can fail early even when actual spend is small. Use --arms, --reps and --prompt-tokens to keep runs cheap.
npm test # node:test via tsx, includes an end-to-end run against a local mock endpoint
npm run typecheck
npm run build # emits dist/, used by the bin entry
MIT, Copyright (c) 2026 Daniel Valle. See LICENSE.