Find the LLM calls your AI agent may never have needed.
Python
2
2 commits
updated Oct 7, 2026
Find the LLM calls your AI agent may never have needed.
AUDR can tell you what an agent run did and what it cost. KORA Doctor asks the next question:
Did all of that inference need to happen?
KORA Doctor is a small open-source CLI that analyzes AUDR JSON/JSONL and flags calls that may be worth removing, caching, replacing with deterministic logic, or moving to a cheaper model.
No dashboard. No account. No hosted service.
Install directly from GitHub:
pipx install git+https://github.com/Krako-Labs/kora-doctor.git
Then audit an AUDR file:
kora-doctor audit audr.jsonl
From a clone:
python3 -m kora_doctor audit samples/inefficient_agent.jsonl
Machine-readable output:
kora-doctor audit audr.jsonl --json
The included inefficient trace is synthetic and deliberately pathological. It is a demo of the reporting surface, not a benchmark or a claim that typical agents waste 89%.
KORA Doctor
Find the LLM calls your AI agent may never have needed.
Observed: 11 records · 2 runs · 11 model calls · 0 tool calls
Observed cost: $0.1050
Potentially avoidable: $0.0936 (89%)
Estimated optimized cost: $0.0114
Candidates
-----------------------------------------------
Duplicate/repeated calls 5
Cache/reuse candidates 5
Deterministic candidates 11
Smaller-model candidates 11
Orchestration overhead 5
Run the included intentionally inefficient trace to see the full current output:
kora-doctor audit samples/inefficient_agent.jsonl
KORA Doctor v0 intentionally starts with simple heuristics:
The goal is not to prove that an inference call was unnecessary. The goal is to narrow a long trace down to the calls a developer should inspect first.
AUDR v1.0.0 deliberately records usage/cost telemetry without prompt content or secrets. That is good for security, but it also means an AUDR record alone usually cannot prove that two model calls were semantically identical or that a task could definitely have been deterministic.
KORA Doctor therefore uses:
Savings are only calculated when cost.total_cost is present. Multiple currencies are never silently converted.
The dollar estimate is a scenario estimate attached to each candidate, not a measured future bill:
If one call matches several rules, KORA Doctor uses only the largest ratio for that call; it never stacks savings estimates. These defaults are intentionally easy to inspect and change as real traces arrive.
KORA Doctor currently targets AUDR v1.0.0 and accepts:
It checks the core fields needed for analysis, but it is not a replacement for the official AUDR JSON Schema conformance validator.
AUDR upstream currently provides adapters for LiteLLM, Vercel AI SDK, Mastra, NVIDIA NeMo Relay, and Merge Gateway, plus sinks including Chargebee.
python3 -m kora_doctor audit samples/simple.jsonl
python3 -m kora_doctor audit samples/multi_step.jsonl
python3 -m kora_doctor audit samples/inefficient_agent.jsonl
The samples are synthetic AUDR-compatible traces created for KORA Doctor. The inefficient trace is intentionally constructed to trigger multiple heuristics.
KORA Doctor has no runtime dependencies.
python3 -m venv .venv
source .venv/bin/activate
python3 -m pip install -e .
python3 -m unittest discover -s tests -v
./scripts/demo.sh
These are deliberate v0 constraints. If real traces show that an extra signal is necessary, we will add the smallest useful one.
KORA Doctor is standalone. It does not require the full KORA runtime.
Today:
AUDR trace -> KORA Doctor -> diagnose
If users pull for it later:
AUDR trace -> diagnose -> recommend -> optimize automatically with KORA
Real AUDR traces, false positives, and missed optimization opportunities are the highest-value feedback.
See CONTRIBUTING.md or open an issue.
Apache-2.0.
Find the LLM calls your AI agent may never have needed.
Python
2
2 commits
updated Oct 7, 2026
Find the LLM calls your AI agent may never have needed.
AUDR can tell you what an agent run did and what it cost. KORA Doctor asks the next question:
Did all of that inference need to happen?
KORA Doctor is a small open-source CLI that analyzes AUDR JSON/JSONL and flags calls that may be worth removing, caching, replacing with deterministic logic, or moving to a cheaper model.
No dashboard. No account. No hosted service.
Install directly from GitHub:
pipx install git+https://github.com/Krako-Labs/kora-doctor.git
Then audit an AUDR file:
kora-doctor audit audr.jsonl
From a clone:
python3 -m kora_doctor audit samples/inefficient_agent.jsonl
Machine-readable output:
kora-doctor audit audr.jsonl --json
The included inefficient trace is synthetic and deliberately pathological. It is a demo of the reporting surface, not a benchmark or a claim that typical agents waste 89%.
KORA Doctor
Find the LLM calls your AI agent may never have needed.
Observed: 11 records · 2 runs · 11 model calls · 0 tool calls
Observed cost: $0.1050
Potentially avoidable: $0.0936 (89%)
Estimated optimized cost: $0.0114
Candidates
-----------------------------------------------
Duplicate/repeated calls 5
Cache/reuse candidates 5
Deterministic candidates 11
Smaller-model candidates 11
Orchestration overhead 5
Run the included intentionally inefficient trace to see the full current output:
kora-doctor audit samples/inefficient_agent.jsonl
KORA Doctor v0 intentionally starts with simple heuristics:
The goal is not to prove that an inference call was unnecessary. The goal is to narrow a long trace down to the calls a developer should inspect first.
AUDR v1.0.0 deliberately records usage/cost telemetry without prompt content or secrets. That is good for security, but it also means an AUDR record alone usually cannot prove that two model calls were semantically identical or that a task could definitely have been deterministic.
KORA Doctor therefore uses:
Savings are only calculated when cost.total_cost is present. Multiple currencies are never silently converted.
The dollar estimate is a scenario estimate attached to each candidate, not a measured future bill:
If one call matches several rules, KORA Doctor uses only the largest ratio for that call; it never stacks savings estimates. These defaults are intentionally easy to inspect and change as real traces arrive.
KORA Doctor currently targets AUDR v1.0.0 and accepts:
It checks the core fields needed for analysis, but it is not a replacement for the official AUDR JSON Schema conformance validator.
AUDR upstream currently provides adapters for LiteLLM, Vercel AI SDK, Mastra, NVIDIA NeMo Relay, and Merge Gateway, plus sinks including Chargebee.
python3 -m kora_doctor audit samples/simple.jsonl
python3 -m kora_doctor audit samples/multi_step.jsonl
python3 -m kora_doctor audit samples/inefficient_agent.jsonl
The samples are synthetic AUDR-compatible traces created for KORA Doctor. The inefficient trace is intentionally constructed to trigger multiple heuristics.
KORA Doctor has no runtime dependencies.
python3 -m venv .venv
source .venv/bin/activate
python3 -m pip install -e .
python3 -m unittest discover -s tests -v
./scripts/demo.sh
These are deliberate v0 constraints. If real traces show that an extra signal is necessary, we will add the smallest useful one.
KORA Doctor is standalone. It does not require the full KORA runtime.
Today:
AUDR trace -> KORA Doctor -> diagnose
If users pull for it later:
AUDR trace -> diagnose -> recommend -> optimize automatically with KORA
Real AUDR traces, false positives, and missed optimization opportunities are the highest-value feedback.
See CONTRIBUTING.md or open an issue.
Apache-2.0.