danidante/llm-wire-check

Does your OpenAI-compatible LLM endpoint actually hoour the parameters you send? Reasoning switches, effrt, caching, model substitution.

TypeScript

0

1 commits

updated Oct 6, 2026

See the code

README

llm-wire-check

Does this OpenAI-compatible endpoint actually honour what I send? Send controlled request variants and see what really changed.

Why

With LLM providers and gateways, almost every misconfiguration returns HTTP 200. The reply looks normal, the parameter did nothing, and the only symptom shows up days later as changed behaviour or a larger bill. The classic silent failures:

  • A reasoning switch is silently dropped. The provider expects a different key (or a different nesting) than the one you sent, accepts the request anyway, and runs with reasoning off.
  • A reasoning switch is translated, not forwarded. A gateway schema-validates the body and maps a field onto its own idea of the upstream's switch, which can be a no-op or a much heavier setting.
  • An effort level is ignored, or inflated. low, medium and high produce the same output length, or one spelling multiplies output tokens several times over.
  • A cache key never reaches the upstream. The field is accepted and inert; every call pays full input price.
  • extra_body sent as a literal field. It is a Python SDK argument that the client merges into the body root before sending. On the wire it is just an unknown key.
  • A different model is served. A gateway falls back or substitutes a model and says so only in a response header, if at all.

llm-wire-check uses Node's native fetch with no SDK, so every parameter reaches the wire exactly as written, and measures the response rather than trusting any flag the provider reports about itself.

Install and usage

Requires Node 20 or newer. No runtime dependencies.

git clone https://github.com/danidante/llm-wire-check.git
cd llm-wire-check
npm install
npx tsx src/cli.ts reasoning --base-url https://api.example.com/v1 --model example-model

The examples below use npx llm-wire-check; from a clone, run npx tsx src/cli.ts instead (or npm run build && npm link once).

The API key is read from an environment variable (default OPENAI_API_KEY) and never printed:

export MY_PROVIDER_KEY=...
npx llm-wire-check reasoning \
  --base-url https://gateway.example.net/v1 \
  --model example-model \
  --api-key-env MY_PROVIDER_KEY \
  --served-model-header x-served-model \
  --effort-sweep

# only two arms, streamed, five repetitions, JSON for CI
npx llm-wire-check reasoning --base-url https://api.example.com/v1 --model example-model \
  --arms none,template_thinking --stream --reps 5 --json > reasoning.json

# does the provider cache a long prefix, and does prompt_cache_key change anything?
npx llm-wire-check cache --base-url https://api.example.com/v1 --model example-model --compare-key

Common options:

OptionDefaultMeaning
--base-url URLrequiredEndpoint base; /chat/completions is appended
--model IDrequiredModel id to request
--api-key-env NAMEOPENAI_API_KEYEnv var holding the key (sent as Authorization: Bearer)
--header "K: V"Extra header, repeatable; an explicit Authorization header wins
--reps N3Repetitions per arm
--prompt-tokens N8000Synthetic prompt size
--max-tokens N3000Output budget per call; warns below 2500
--max-tokens-field NAMEmax_tokense.g. max_completion_tokens for endpoints that require it
--timeout-ms N180000Per-call timeout
--served-model-header NAMEResponse header naming the served model, repeatable
--streamoffSSE streaming (reasoning command)
--jsonoffOne JSON object on stdout; progress stays on stderr

Exit codes: 0 no findings, 1 findings, 2 usage error or nothing reached the endpoint.

Arms (reasoning command)

Each arm is one request body variant. Every arm is sent --reps times, and the arm order rotates on each repetition so no arm always runs first (provider load drifts over minutes).

ArmAdds to the bodyWhat it tests
nonenothingBaseline. Always included.
effortreasoning_effort: "high"Top-level effort field
template_thinkingchat_template_kwargs: { thinking: true }Chat-template switch, thinking spelling
template_enable_thinkingchat_template_kwargs: { enable_thinking: true }Chat-template switch, enable_thinking spelling
template_effortchat_template_kwargs: { thinking: true, reasoning_effort: "high" }Effort nested inside the template kwargs
extra_body_literalextra_body: { chat_template_kwargs: { enable_thinking: true } }The classic mistake: a literal extra_body key on the wire
effort_low, effort_medium, effort_highreasoning_effort: "low" / "medium" / "high"Added by --effort-sweep: does effort scale anything?

--arms a,b,c selects a subset; none is added automatically. Note that effort and effort_high send the same body.

Reasoning text is detected in message.reasoning_content, message.reasoning (string), message.reasoning_details[].text, and inline <think>...</think> in the content (reported as inline_think, and stripped before counting content). A stray </think> without an opener (the template opened it in the prompt) and an unclosed <think> (cut off mid-thought) count as inline too. In --stream mode the same fields are read from delta, with stream_options: { include_usage: true } sent so usage arrives.

Findings reference

CodeExit codeMeaning
SWITCH_WORKS:<arm>infoThe arm reasons and the none baseline does not
SWITCH_IGNORED:<arm>findingThe arm sends a switch, the call succeeds, and reasoning stays at zero like the baseline
REASONS_UNCONDITIONALLYinfoThe baseline already reasons, so switch presence cannot be attributed (switch verdicts are skipped)
ARM_REJECTED:<arm>infoEvery call of the arm got a non-2xx. A 400 is information: the endpoint validates the field
BASELINE_FAILEDfindingThe none arm never succeeded; switch verdicts unavailable
EFFORT_IGNOREDfindingSweep medians for low / medium / high are within 15% of each other (reasoning chars, or completion tokens when no reasoning text is exposed)
OUTPUT_INFLATION:<arm>findingMedian completion tokens are at least 2x the lightest arm that also reasons
READING_UNRELIABLE:<arm>findingAt least one call hit the output budget; its reading is an artifact
REASONING_TOKENS_UNREPORTEDfindingReasoning text came back but usage.completion_tokens_details.reasoning_tokens was missing or 0
MODEL_SUBSTITUTEDfindingA --served-model-header named a different model (a vendor/ prefix alone does not count)
CACHEDinfoCall 2 or 3 read more than 50% of the prompt from cache
PARTIAL_CACHEinfoBetween 5% and 50% read from cache
NOT_CACHEDfindingA cached-token field is present and reports about 0 on repeat calls
UNVERIFIABLEfindingNo cached-token field at all. This means "cannot verify", not "no cache"
COST_IDENTICALfindingusage.cost is identical on all 3 calls of a large prompt; a working cache normally shows in the price
WARM_BEFORE_TESTinfoCall 1 was already cached
CACHE_KEY_NO_EFFECTinfoWith --compare-key: the result is the same with and without the key
CACHE_CALLS_FAILEDfindingRepeat calls failed, nothing to compare

With --compare-key, cache codes carry a :with_key or :no_key suffix. Cached tokens are read from usage.prompt_tokens_details.cached_tokens, then usage.cached_tokens, then any numeric usage key matching /cach/i (cache reads preferred over cache writes and misses).

Measurement traps built in

  • Prompt size. The default prompt is about 8,000 tokens of varied, numbered, distinct filler sentences. A 30-token probe can report a provider healthy that hangs, slows to a crawl, or behaves differently on real prompt sizes. Use --prompt-tokens to match your production prompts.
  • The conversation ends on a user turn. Some models degenerate when the last message is a system message, which would look like a provider defect.
  • The output budget. Reasoning can consume the whole max_tokens budget, leaving empty content that looks like a failure. Any call with completion_tokens >= max_tokens or finish_reason == "length" is flagged TRUNC and its reading is treated as an artifact. That equality is the tell. Keep --max-tokens at 2500 or more.
  • Never trust reported reasoning_tokens. Many providers omit it, or report 0 while returning thousands of characters of reasoning. The tool counts reasoning characters and shows chars / 3.9 labelled as an estimate.
  • A catalog flag is not evidence. A model list saying "supports reasoning" says nothing about whether your request shape turns it on. Only a returned reasoning trace counts.
  • 400s are information. A rejected arm shows that the endpoint schema-validates its input; the first 300 characters of the body are kept so you can see which field it refused.
  • Interleave arms. Arm order rotates every repetition so slow minutes are spread across arms instead of landing on one.
  • Medians plus tails. Per arm you get median, min and max. A clean median can hide a tail that breaks real traffic; re-run at a different time of day before concluding.
  • Caching needs a fresh prefix. Each cache series uses a newly randomised prefix, so a previous run cannot pre-warm it, and --compare-key series do not share a prefix.

Sample output

Example only: synthetic numbers against an imaginary endpoint, trimmed to three arms.

llm-wire-check reasoning  model=example-model  reps=3  prompt~8000 tok  max_tokens=3000

ARM none
rep  http  ms    r_chars  r_source  est_rtok  rep_rtok  compl  prompt  content  finish  flags
---  ----  ----  -------  --------  --------  --------  -----  ------  -------  ------  -----
1    200   2140  0        none      0         -         38     8041    151      stop    -
2    200   1985  0        none      0         -         41     8041    163      stop    -
3    200   2310  0        none      0         -         36     8041    144      stop    -
  summary: ok 3/3  r_chars median 0 (min 0, max 0)  compl median 38  latency median 2140ms  truncated 0  empty 0

ARM template_thinking
rep  http  ms    r_chars  r_source           est_rtok  rep_rtok  compl  prompt  content  finish  flags
---  ----  ----  -------  -----------------  --------  --------  -----  ------  -------  ------  -----
1    200   6420  3074     reasoning_content  788       0         826    8041    148      stop    -
2    200   7015  3219     reasoning_content  825       0         867    8041    159      stop    -
3    200   5890  2494     reasoning_content  639       0         679    8041    152      stop    -
  summary: ok 3/3  r_chars median 3074 (min 2494, max 3219)  compl median 826  latency median 6420ms  truncated 0  empty 0

ARM extra_body_literal
rep  http  ms    r_chars  r_source  est_rtok  rep_rtok  compl  prompt  content  finish  flags
---  ----  ----  -------  --------  --------  --------  -----  ------  -------  ------  -----
1    400   212   0        none      0         -         -      -       0        -       -
2    400   198   0        none      0         -         -      -       0        -       -
3    400   205   0        none      0         -         -      -       0        -       -
  rep 1 body: {"error":{"message":"invalid request: unknown field \"extra_body\""}}
  ...
  summary: ok 0/3  r_chars median 0 (min 0, max 0)  compl median -  latency median -  truncated 0  empty 0

FINDINGS
  [info] ARM_REJECTED:extra_body_literal
         every call returned HTTP 400; see the body snippet
  [info] SWITCH_WORKS:template_thinking
         median 3074 reasoning chars vs 0 with no params
  [FIND] REASONING_TOKENS_UNREPORTED
         3 call(s) returned reasoning text but usage reasoning_tokens was missing or 0 (arms: template_thinking); do not use that field as proof

Each run ends with a short READING THIS legend explaining every column.

Cost note

Each call sends the full synthetic prompt. The reasoning command makes arms x reps calls (18 with the defaults, about 144,000 input tokens at 8,000 tokens each, plus up to 3,000 output tokens per call). The cache command makes 3 calls, or 6 with --compare-key, at max_tokens 16. The tool prints this estimate before it starts. Some gateways reserve input plus max_tokens per call up front, so a low balance can fail early even when actual spend is small. Use --arms, --reps and --prompt-tokens to keep runs cheap.

Development

npm test           # node:test via tsx, includes an end-to-end run against a local mock endpoint
npm run typecheck
npm run build      # emits dist/, used by the bin entry

License

MIT, Copyright (c) 2026 Daniel Valle. See LICENSE.

danidante/llm-wire-check

Does your OpenAI-compatible LLM endpoint actually hoour the parameters you send? Reasoning switches, effrt, caching, model substitution.

TypeScript

0

1 commits

updated Oct 6, 2026

See the code

README

llm-wire-check

Does this OpenAI-compatible endpoint actually honour what I send? Send controlled request variants and see what really changed.

Why

With LLM providers and gateways, almost every misconfiguration returns HTTP 200. The reply looks normal, the parameter did nothing, and the only symptom shows up days later as changed behaviour or a larger bill. The classic silent failures:

  • A reasoning switch is silently dropped. The provider expects a different key (or a different nesting) than the one you sent, accepts the request anyway, and runs with reasoning off.
  • A reasoning switch is translated, not forwarded. A gateway schema-validates the body and maps a field onto its own idea of the upstream's switch, which can be a no-op or a much heavier setting.
  • An effort level is ignored, or inflated. low, medium and high produce the same output length, or one spelling multiplies output tokens several times over.
  • A cache key never reaches the upstream. The field is accepted and inert; every call pays full input price.
  • extra_body sent as a literal field. It is a Python SDK argument that the client merges into the body root before sending. On the wire it is just an unknown key.
  • A different model is served. A gateway falls back or substitutes a model and says so only in a response header, if at all.

llm-wire-check uses Node's native fetch with no SDK, so every parameter reaches the wire exactly as written, and measures the response rather than trusting any flag the provider reports about itself.

Install and usage

Requires Node 20 or newer. No runtime dependencies.

git clone https://github.com/danidante/llm-wire-check.git
cd llm-wire-check
npm install
npx tsx src/cli.ts reasoning --base-url https://api.example.com/v1 --model example-model

The examples below use npx llm-wire-check; from a clone, run npx tsx src/cli.ts instead (or npm run build && npm link once).

The API key is read from an environment variable (default OPENAI_API_KEY) and never printed:

export MY_PROVIDER_KEY=...
npx llm-wire-check reasoning \
  --base-url https://gateway.example.net/v1 \
  --model example-model \
  --api-key-env MY_PROVIDER_KEY \
  --served-model-header x-served-model \
  --effort-sweep

# only two arms, streamed, five repetitions, JSON for CI
npx llm-wire-check reasoning --base-url https://api.example.com/v1 --model example-model \
  --arms none,template_thinking --stream --reps 5 --json > reasoning.json

# does the provider cache a long prefix, and does prompt_cache_key change anything?
npx llm-wire-check cache --base-url https://api.example.com/v1 --model example-model --compare-key

Common options:

OptionDefaultMeaning
--base-url URLrequiredEndpoint base; /chat/completions is appended
--model IDrequiredModel id to request
--api-key-env NAMEOPENAI_API_KEYEnv var holding the key (sent as Authorization: Bearer)
--header "K: V"Extra header, repeatable; an explicit Authorization header wins
--reps N3Repetitions per arm
--prompt-tokens N8000Synthetic prompt size
--max-tokens N3000Output budget per call; warns below 2500
--max-tokens-field NAMEmax_tokense.g. max_completion_tokens for endpoints that require it
--timeout-ms N180000Per-call timeout
--served-model-header NAMEResponse header naming the served model, repeatable
--streamoffSSE streaming (reasoning command)
--jsonoffOne JSON object on stdout; progress stays on stderr

Exit codes: 0 no findings, 1 findings, 2 usage error or nothing reached the endpoint.

Arms (reasoning command)

Each arm is one request body variant. Every arm is sent --reps times, and the arm order rotates on each repetition so no arm always runs first (provider load drifts over minutes).

ArmAdds to the bodyWhat it tests
nonenothingBaseline. Always included.
effortreasoning_effort: "high"Top-level effort field
template_thinkingchat_template_kwargs: { thinking: true }Chat-template switch, thinking spelling
template_enable_thinkingchat_template_kwargs: { enable_thinking: true }Chat-template switch, enable_thinking spelling
template_effortchat_template_kwargs: { thinking: true, reasoning_effort: "high" }Effort nested inside the template kwargs
extra_body_literalextra_body: { chat_template_kwargs: { enable_thinking: true } }The classic mistake: a literal extra_body key on the wire
effort_low, effort_medium, effort_highreasoning_effort: "low" / "medium" / "high"Added by --effort-sweep: does effort scale anything?

--arms a,b,c selects a subset; none is added automatically. Note that effort and effort_high send the same body.

Reasoning text is detected in message.reasoning_content, message.reasoning (string), message.reasoning_details[].text, and inline <think>...</think> in the content (reported as inline_think, and stripped before counting content). A stray </think> without an opener (the template opened it in the prompt) and an unclosed <think> (cut off mid-thought) count as inline too. In --stream mode the same fields are read from delta, with stream_options: { include_usage: true } sent so usage arrives.

Findings reference

CodeExit codeMeaning
SWITCH_WORKS:<arm>infoThe arm reasons and the none baseline does not
SWITCH_IGNORED:<arm>findingThe arm sends a switch, the call succeeds, and reasoning stays at zero like the baseline
REASONS_UNCONDITIONALLYinfoThe baseline already reasons, so switch presence cannot be attributed (switch verdicts are skipped)
ARM_REJECTED:<arm>infoEvery call of the arm got a non-2xx. A 400 is information: the endpoint validates the field
BASELINE_FAILEDfindingThe none arm never succeeded; switch verdicts unavailable
EFFORT_IGNOREDfindingSweep medians for low / medium / high are within 15% of each other (reasoning chars, or completion tokens when no reasoning text is exposed)
OUTPUT_INFLATION:<arm>findingMedian completion tokens are at least 2x the lightest arm that also reasons
READING_UNRELIABLE:<arm>findingAt least one call hit the output budget; its reading is an artifact
REASONING_TOKENS_UNREPORTEDfindingReasoning text came back but usage.completion_tokens_details.reasoning_tokens was missing or 0
MODEL_SUBSTITUTEDfindingA --served-model-header named a different model (a vendor/ prefix alone does not count)
CACHEDinfoCall 2 or 3 read more than 50% of the prompt from cache
PARTIAL_CACHEinfoBetween 5% and 50% read from cache
NOT_CACHEDfindingA cached-token field is present and reports about 0 on repeat calls
UNVERIFIABLEfindingNo cached-token field at all. This means "cannot verify", not "no cache"
COST_IDENTICALfindingusage.cost is identical on all 3 calls of a large prompt; a working cache normally shows in the price
WARM_BEFORE_TESTinfoCall 1 was already cached
CACHE_KEY_NO_EFFECTinfoWith --compare-key: the result is the same with and without the key
CACHE_CALLS_FAILEDfindingRepeat calls failed, nothing to compare

With --compare-key, cache codes carry a :with_key or :no_key suffix. Cached tokens are read from usage.prompt_tokens_details.cached_tokens, then usage.cached_tokens, then any numeric usage key matching /cach/i (cache reads preferred over cache writes and misses).

Measurement traps built in

  • Prompt size. The default prompt is about 8,000 tokens of varied, numbered, distinct filler sentences. A 30-token probe can report a provider healthy that hangs, slows to a crawl, or behaves differently on real prompt sizes. Use --prompt-tokens to match your production prompts.
  • The conversation ends on a user turn. Some models degenerate when the last message is a system message, which would look like a provider defect.
  • The output budget. Reasoning can consume the whole max_tokens budget, leaving empty content that looks like a failure. Any call with completion_tokens >= max_tokens or finish_reason == "length" is flagged TRUNC and its reading is treated as an artifact. That equality is the tell. Keep --max-tokens at 2500 or more.
  • Never trust reported reasoning_tokens. Many providers omit it, or report 0 while returning thousands of characters of reasoning. The tool counts reasoning characters and shows chars / 3.9 labelled as an estimate.
  • A catalog flag is not evidence. A model list saying "supports reasoning" says nothing about whether your request shape turns it on. Only a returned reasoning trace counts.
  • 400s are information. A rejected arm shows that the endpoint schema-validates its input; the first 300 characters of the body are kept so you can see which field it refused.
  • Interleave arms. Arm order rotates every repetition so slow minutes are spread across arms instead of landing on one.
  • Medians plus tails. Per arm you get median, min and max. A clean median can hide a tail that breaks real traffic; re-run at a different time of day before concluding.
  • Caching needs a fresh prefix. Each cache series uses a newly randomised prefix, so a previous run cannot pre-warm it, and --compare-key series do not share a prefix.

Sample output

Example only: synthetic numbers against an imaginary endpoint, trimmed to three arms.

llm-wire-check reasoning  model=example-model  reps=3  prompt~8000 tok  max_tokens=3000

ARM none
rep  http  ms    r_chars  r_source  est_rtok  rep_rtok  compl  prompt  content  finish  flags
---  ----  ----  -------  --------  --------  --------  -----  ------  -------  ------  -----
1    200   2140  0        none      0         -         38     8041    151      stop    -
2    200   1985  0        none      0         -         41     8041    163      stop    -
3    200   2310  0        none      0         -         36     8041    144      stop    -
  summary: ok 3/3  r_chars median 0 (min 0, max 0)  compl median 38  latency median 2140ms  truncated 0  empty 0

ARM template_thinking
rep  http  ms    r_chars  r_source           est_rtok  rep_rtok  compl  prompt  content  finish  flags
---  ----  ----  -------  -----------------  --------  --------  -----  ------  -------  ------  -----
1    200   6420  3074     reasoning_content  788       0         826    8041    148      stop    -
2    200   7015  3219     reasoning_content  825       0         867    8041    159      stop    -
3    200   5890  2494     reasoning_content  639       0         679    8041    152      stop    -
  summary: ok 3/3  r_chars median 3074 (min 2494, max 3219)  compl median 826  latency median 6420ms  truncated 0  empty 0

ARM extra_body_literal
rep  http  ms    r_chars  r_source  est_rtok  rep_rtok  compl  prompt  content  finish  flags
---  ----  ----  -------  --------  --------  --------  -----  ------  -------  ------  -----
1    400   212   0        none      0         -         -      -       0        -       -
2    400   198   0        none      0         -         -      -       0        -       -
3    400   205   0        none      0         -         -      -       0        -       -
  rep 1 body: {"error":{"message":"invalid request: unknown field \"extra_body\""}}
  ...
  summary: ok 0/3  r_chars median 0 (min 0, max 0)  compl median -  latency median -  truncated 0  empty 0

FINDINGS
  [info] ARM_REJECTED:extra_body_literal
         every call returned HTTP 400; see the body snippet
  [info] SWITCH_WORKS:template_thinking
         median 3074 reasoning chars vs 0 with no params
  [FIND] REASONING_TOKENS_UNREPORTED
         3 call(s) returned reasoning text but usage reasoning_tokens was missing or 0 (arms: template_thinking); do not use that field as proof

Each run ends with a short READING THIS legend explaining every column.

Cost note

Each call sends the full synthetic prompt. The reasoning command makes arms x reps calls (18 with the defaults, about 144,000 input tokens at 8,000 tokens each, plus up to 3,000 output tokens per call). The cache command makes 3 calls, or 6 with --compare-key, at max_tokens 16. The tool prints this estimate before it starts. Some gateways reserve input plus max_tokens per call up front, so a low balance can fail early even when actual spend is small. Use --arms, --reps and --prompt-tokens to keep runs cheap.

Development

npm test           # node:test via tsx, includes an end-to-end run against a local mock endpoint
npm run typecheck
npm run build      # emits dist/, used by the bin entry

License

MIT, Copyright (c) 2026 Daniel Valle. See LICENSE.