Arthur031221/toolcall-check

Check nested tool arguments, streamed calls and the result handoff on a chat endpoint.

Python

0

3 commits

updated Oct 3, 2026

See the code

See what people are saying

SourceMessageScoreDate

A small CLI for checking nested tool calls, streaming, and the next turn (r/LocalLLaMA)

I'm the author of toolcall-check, a small Python CLI for checking chat completions compatible endpoints. It exercises two forced function calls, two streamed calls, and one two turn round trip that returns a local result and checks the exact normal answer. Nested argument values retain JSON types,…

2

Oct 3, 2026

README


toolcall-check

Check exact tool arguments, streamed calls, and the return trip through a local tool result.

GitHub stars Tests MIT license

Quickstart | How it works | Examples | FAQ

[!TIP] Try the local synthetic fixture without configuring an endpoint:

uvx --from git+https://github.com/Arthur031221/toolcall-check toolcall-check --demo

The real CLI checks a synthetic local endpoint, showing two nested argument failures followed by five passing checks.

Why toolcall-check

A successful HTTP response does not show whether nested arguments kept their JSON types, whether a streamed call survived fragment boundaries, or whether the assistant can use a tool result on its next turn. When one of those steps goes wrong, the response body alone can be difficult to inspect.

This command sends five small probes and saves a private HTML report with sanitized request and response traces. It does not run generated tools. The local demo uses a synthetic fixture and says so in its output. A failed forced-call probe describes that case, and does not establish whether automatic tool selection works.

Features

  • Checks nested values: compares exact keys, arrays, strings, numbers, and booleans.
  • Rebuilds streamed calls: joins fragmented UTF-8 SSE events by choice and call index.
  • Checks the next turn: returns a fixed local result and requires an exact normal assistant answer.
  • Keeps failure evidence: each probe records expected values, received values, and its bounded trace.
  • Limits requests: rejects redirects and bounds response bytes and overall request time.
  • Writes private reports: creates a new directory with restrictive permissions and escapes HTML evidence.

Quickstart

The hero command requires uv. For a persistent command, install from GitHub:

uv tool install git+https://github.com/Arthur031221/toolcall-check

Try the synthetic fixture in one command:

uvx \
  --from git+https://github.com/Arthur031221/toolcall-check \
  toolcall-check \
  --demo \
  --output fixture-report

For a local endpoint, set a key only when the endpoint requires one:

export TOOLCALL_CHECK_API_KEY="your-local-endpoint-key"
toolcall-check --base-url http://127.0.0.1:1234/v1 --model local-model --output endpoint-report

The command exits 0 when all checks pass, 1 when one or more checks fail, and 2 for invalid configuration or report-write errors. It refuses to reuse an output path.

Examples

A successful synthetic run prints a report path followed by the check count:

Checks: 5/5 passed
PASS Forced echo call
PASS Forced nested JSON call
PASS Streaming echo
PASS Streaming nested JSON
PASS Two-turn echo round trip

Show two intentional argument mismatches and inspect their evidence:

toolcall-check --demo-broken --output broken-report

This command exits 1. Its report is labeled synthetic and is not evidence of a remote model's compatibility.

How it works

The checks force one named function at a time, then reconstruct streamed deltas before comparing arguments. The final check replays the validated assistant call with a locally generated result and asks for a normal assistant response containing that result exactly. A strict completion marker policy requires finish_reason: tool_calls and [DONE] for streams. The parser rejects duplicate JSON keys and non-finite numbers. Booleans, integers, and floats remain distinct during exact comparison.

Each single-call probe runs in a disposable Python process. The round-trip probe shares one process across its two exchanges. The parent enforces an overall wall-time limit and reaps a timed-out worker. HTTP redirects are rejected. Responses are capped at 256 KiB and requests at 64 KiB. Single-call probes have a 20 second wall limit, while the complete round trip has a 40 second limit. Socket operations use a 15 second timeout. The tool never invokes a generated function or command.

ToolWhat it checksHow this project is scoped
CompatCanaryA broader chat API compatibility scan that documents forced calls, streaming, and structured output.This project concentrates on nested argument integrity, reconstructing streamed tool fragments, a fixed second-turn result, and inspectable traces.
toolcall-checkFive forced, streaming, and round-trip probes with exact argument comparison.The report keeps one result and trace for each named check.
View a generated HTML report The synthetic endpoint report shows five passing checks with expected and received arguments.

Details

Options and saved files
toolcall-check --base-url URL [--model NAME] [--api-key-env VARIABLE] [--output DIRECTORY]
toolcall-check --demo [--output DIRECTORY]
toolcall-check --demo-broken [--output DIRECTORY]

The default key variable is TOOLCALL_CHECK_API_KEY. Use --api-key-env to select another variable. The key is sent as an authorization header. Saved strings and reconstructed stream fields mask the supplied key. Authorization headers are omitted from traces. Other credential fields and common token patterns are redacted recursively. Redaction is best effort, so inspect a report before sharing it. When joined stream fragments contain a detected secret, their trace fragments are masked and the sanitized assembled field is saved.

A new output directory contains report.html, results.json, and trace.json. The directory uses mode 0700 and each file uses mode 0600 on Unix-like systems. The HTML has no scripts, remote assets, or external requests.

Continuous integration

The workflow runs the standard-library test suite on Python 3.10 through 3.14. Install the project with python -m pip install --editable . and run python -m unittest discover -s tests -v.

FAQ

Does the demo test a remote model? No. It starts a local synthetic HTTP fixture and exercises the same HTTP client and report path.

What does an empty or malformed call mean? The relevant probe fails and keeps bounded response evidence in the report. The remaining checks still run.

Can the report be shared as-is? Review the HTML and both JSON files first. Secret redaction is best effort.

Contributing

Report a problem or send a focused pull request. See CONTRIBUTING.md for the test command and report-sharing guidance.

License

MIT. See LICENSE.

cli
llm
ollama
python
streaming
testing
tool-calling

Arthur031221/toolcall-check

Check nested tool arguments, streamed calls and the result handoff on a chat endpoint.

Python

0

3 commits

updated Oct 3, 2026

See the code

See what people are saying

SourceMessageScoreDate

A small CLI for checking nested tool calls, streaming, and the next turn (r/LocalLLaMA)

I'm the author of toolcall-check, a small Python CLI for checking chat completions compatible endpoints. It exercises two forced function calls, two streamed calls, and one two turn round trip that returns a local result and checks the exact normal answer. Nested argument values retain JSON types,…

2

Oct 3, 2026

README


toolcall-check

Check exact tool arguments, streamed calls, and the return trip through a local tool result.

GitHub stars Tests MIT license

Quickstart | How it works | Examples | FAQ

[!TIP] Try the local synthetic fixture without configuring an endpoint:

uvx --from git+https://github.com/Arthur031221/toolcall-check toolcall-check --demo

The real CLI checks a synthetic local endpoint, showing two nested argument failures followed by five passing checks.

Why toolcall-check

A successful HTTP response does not show whether nested arguments kept their JSON types, whether a streamed call survived fragment boundaries, or whether the assistant can use a tool result on its next turn. When one of those steps goes wrong, the response body alone can be difficult to inspect.

This command sends five small probes and saves a private HTML report with sanitized request and response traces. It does not run generated tools. The local demo uses a synthetic fixture and says so in its output. A failed forced-call probe describes that case, and does not establish whether automatic tool selection works.

Features

  • Checks nested values: compares exact keys, arrays, strings, numbers, and booleans.
  • Rebuilds streamed calls: joins fragmented UTF-8 SSE events by choice and call index.
  • Checks the next turn: returns a fixed local result and requires an exact normal assistant answer.
  • Keeps failure evidence: each probe records expected values, received values, and its bounded trace.
  • Limits requests: rejects redirects and bounds response bytes and overall request time.
  • Writes private reports: creates a new directory with restrictive permissions and escapes HTML evidence.

Quickstart

The hero command requires uv. For a persistent command, install from GitHub:

uv tool install git+https://github.com/Arthur031221/toolcall-check

Try the synthetic fixture in one command:

uvx \
  --from git+https://github.com/Arthur031221/toolcall-check \
  toolcall-check \
  --demo \
  --output fixture-report

For a local endpoint, set a key only when the endpoint requires one:

export TOOLCALL_CHECK_API_KEY="your-local-endpoint-key"
toolcall-check --base-url http://127.0.0.1:1234/v1 --model local-model --output endpoint-report

The command exits 0 when all checks pass, 1 when one or more checks fail, and 2 for invalid configuration or report-write errors. It refuses to reuse an output path.

Examples

A successful synthetic run prints a report path followed by the check count:

Checks: 5/5 passed
PASS Forced echo call
PASS Forced nested JSON call
PASS Streaming echo
PASS Streaming nested JSON
PASS Two-turn echo round trip

Show two intentional argument mismatches and inspect their evidence:

toolcall-check --demo-broken --output broken-report

This command exits 1. Its report is labeled synthetic and is not evidence of a remote model's compatibility.

How it works

The checks force one named function at a time, then reconstruct streamed deltas before comparing arguments. The final check replays the validated assistant call with a locally generated result and asks for a normal assistant response containing that result exactly. A strict completion marker policy requires finish_reason: tool_calls and [DONE] for streams. The parser rejects duplicate JSON keys and non-finite numbers. Booleans, integers, and floats remain distinct during exact comparison.

Each single-call probe runs in a disposable Python process. The round-trip probe shares one process across its two exchanges. The parent enforces an overall wall-time limit and reaps a timed-out worker. HTTP redirects are rejected. Responses are capped at 256 KiB and requests at 64 KiB. Single-call probes have a 20 second wall limit, while the complete round trip has a 40 second limit. Socket operations use a 15 second timeout. The tool never invokes a generated function or command.

ToolWhat it checksHow this project is scoped
CompatCanaryA broader chat API compatibility scan that documents forced calls, streaming, and structured output.This project concentrates on nested argument integrity, reconstructing streamed tool fragments, a fixed second-turn result, and inspectable traces.
toolcall-checkFive forced, streaming, and round-trip probes with exact argument comparison.The report keeps one result and trace for each named check.
View a generated HTML report The synthetic endpoint report shows five passing checks with expected and received arguments.

Details

Options and saved files
toolcall-check --base-url URL [--model NAME] [--api-key-env VARIABLE] [--output DIRECTORY]
toolcall-check --demo [--output DIRECTORY]
toolcall-check --demo-broken [--output DIRECTORY]

The default key variable is TOOLCALL_CHECK_API_KEY. Use --api-key-env to select another variable. The key is sent as an authorization header. Saved strings and reconstructed stream fields mask the supplied key. Authorization headers are omitted from traces. Other credential fields and common token patterns are redacted recursively. Redaction is best effort, so inspect a report before sharing it. When joined stream fragments contain a detected secret, their trace fragments are masked and the sanitized assembled field is saved.

A new output directory contains report.html, results.json, and trace.json. The directory uses mode 0700 and each file uses mode 0600 on Unix-like systems. The HTML has no scripts, remote assets, or external requests.

Continuous integration

The workflow runs the standard-library test suite on Python 3.10 through 3.14. Install the project with python -m pip install --editable . and run python -m unittest discover -s tests -v.

FAQ

Does the demo test a remote model? No. It starts a local synthetic HTTP fixture and exercises the same HTTP client and report path.

What does an empty or malformed call mean? The relevant probe fails and keeps bounded response evidence in the report. The remaining checks still run.

Can the report be shared as-is? Review the HTML and both JSON files first. Secret redaction is best effort.

Contributing

Report a problem or send a focused pull request. See CONTRIBUTING.md for the test command and report-sharing guidance.

License

MIT. See LICENSE.

cli
llm
ollama
python
streaming
testing
tool-calling

Languages

Python

95.2%

HTML

4.4%