Deterministic PDF coordinate grounding for LLM extractions. Pin every extracted value to its exact source bbox - no ML, no API, no hallucination. 72.6% Word F1.
Python
1
80 commits
updated Oct 6, 2026
Deterministic evidence grounding for LLM document extraction.
TonerHound resolves extracted values to exact physical PDF coordinates — no coordinate hallucination, no neural networks, no cloud APIs. Given any extracted JSON and its source PDF, it returns the precise bounding box for every value, or tells you the value cannot be grounded.
LLMs extract values from documents. They rarely say where those values came from. TonerHound closes that gap.
Extracted: {"total_revenue": "$8,420.50"}
↓ TonerHound
Grounded: {"total_revenue": {"value": "$8,420.50",
"page": 47,
"bbox": [0.412, 0.318, 0.089, 0.011],
"status": "VERIFIED"}}
Every bounding box is deterministic: same input, same output, every time. No sampling, no temperature, no API calls. Runs on a laptop, offline.
| Property | TonerHound | Typical VLM grounding |
|---|---|---|
| Reproducible | ✅ Same output every run | ❌ Varies with sampling |
| Auditable | ✅ Traceable to character offsets | ❌ Black-box neural net |
| Offline | ✅ No network required | ❌ Cloud API |
| Cost | ✅ $0.00 per page | ⚠️ $0.01–$0.40 per page |
| Accuracy (ExtractBench) | 72.62% Word F1 | 82.2% Word F1 (frontier) |
TonerHound trades ~10 points of grounding accuracy for reproducibility, auditability, and zero cost. That trade is worth it when you need to prove where a number came from.
git clone https://github.com/vanrajsinh650/TonerHound
cd TonerHound
uv sync
Requires Python 3.12+ and Tesseract OCR. Install Tesseract via brew install tesseract (macOS) or apt install tesseract-ocr (Linux).
import tonerhound
results = tonerhound.resolve(
document="invoice.pdf",
extraction=[
{"field": "invoice_number", "value": "INV-2024-0847"},
{"field": "total", "value": "$8,420.50"},
{"field": "due_date", "value": "2024-03-15"},
],
)
for res in results:
if res.is_grounded:
print(f"{res.field}: {res.status.value} → page {res.page}, bbox {res.bbox}")
else:
print(f"{res.field}: UNGROUNDED ({res.status.value})")
Output:
invoice_number: exact → page 1, bbox [0.12, 0.08, 0.21, 0.02]
total: normalized → page 3, bbox [0.41, 0.32, 0.09, 0.01]
due_date: exact → page 1, bbox [0.68, 0.08, 0.14, 0.02]
If a value cannot be grounded, res.is_grounded is False with status not_found or ambiguous, and res.bbox is None.
python -m tonerhound.benchmark --data-dir research/data/full --exp-id VERIFY
# Expected: Word Grounding F1 = 72.6179%
Evaluated on the official ExtractBench 370-document corpus (498,140 fields). Full results in docs/benchmark-results.md.
| Metric | Score |
|---|---|
| Word Grounding F1 | 72.6179% |
| Page Grounding F1 | 83.8490% |
| Word Precision | 77.7935% |
| Word Recall | 68.9729% |
| Passing Fields | 327,671 / 498,140 |
| Regressions (across 7 experiments) | 0 |
What drove the score (three techniques did ~75% of the work):
What didn't work (and why it matters): semantic normalization (2/37,850), column rail constraints (0), hyphen joiner v2 (0), Needleman-Wunsch alignment (13). These negative results define the deterministic ceiling. See docs/how-tonerhound-reached-72.md for the full research narrative.
PDF ──► pypdfium2 ──► character stream
│
├──► inverted token index
├──► numeric index
├──► date index
└──► OCR noise 3-gram index
│
Extracted JSON ──► Evidence Resolver
│
├── Hungarian table assignment
├── Multi-token sequence matching
├── Visual pixel-statistics fallback
├── Date literal variants
└── Multi-region assembly
│
▼
Grounded JSON + bounding boxes
Stack: pypdfium2 · pytesseract · numpy · scipy · OpenCV · PyMuPDF · rapidfuzz · python-dateutil
No PyTorch. No TensorFlow. No models. No training. No GPU.
TonerHound/
├── src/tonerhound/ # Production code
│ ├── document/ # PDF parsing, indexing, OCR routing
│ ├── resolution/ # Evidence resolver, Hungarian assignment
│ ├── matching/ # Token matching, date variants
│ ├── geometry/ # Bounding box assembly, multi-region
│ └── vision/ # Deterministic checkbox detection
├── tests/ # 261 unit tests (~20s)
├── benchmarks/ # Official ExtractBench runners
├── research/
│ ├── experiments/ # EXP-001 through EXP-043
│ └── observer/ # Failure Microscope V4
├── docs/
│ ├── benchmark-results.md
│ └── how-tonerhound-reached-72.md
├── pyproject.toml
├── uv.lock
└── LICENSE
Contributions welcome. Before opening a PR:
pytest tests/ -v # 261 tests must pass
python -m tonerhound.benchmark --exp-id YOUR-EXP # Score must not regress below 72.6179%
research/experiments/.MIT License. See LICENSE for details.
You can use TonerHound commercially, modify it, redistribute it, or embed it in proprietary software. Attribution is appreciated but not required.
If you use TonerHound in research, cite:
@software{tonerhound2026,
title = {TonerHound: Deterministic Evidence Grounding for LLM Document Extraction},
author = {Vanrajsinh},
year = {2026},
url = {https://github.com/vanrajsinh650/TonerHound}
}
ExtractBench (LlamaIndex) for the benchmark corpus. pypdfium2 for fast PDF character extraction. Tesseract for the OCR fallback. The scipy team for linear_sum_assignment.
Built without neural networks, by choice.
Deterministic PDF coordinate grounding for LLM extractions. Pin every extracted value to its exact source bbox - no ML, no API, no hallucination. 72.6% Word F1.
Python
1
80 commits
updated Oct 6, 2026
Deterministic evidence grounding for LLM document extraction.
TonerHound resolves extracted values to exact physical PDF coordinates — no coordinate hallucination, no neural networks, no cloud APIs. Given any extracted JSON and its source PDF, it returns the precise bounding box for every value, or tells you the value cannot be grounded.
LLMs extract values from documents. They rarely say where those values came from. TonerHound closes that gap.
Extracted: {"total_revenue": "$8,420.50"}
↓ TonerHound
Grounded: {"total_revenue": {"value": "$8,420.50",
"page": 47,
"bbox": [0.412, 0.318, 0.089, 0.011],
"status": "VERIFIED"}}
Every bounding box is deterministic: same input, same output, every time. No sampling, no temperature, no API calls. Runs on a laptop, offline.
| Property | TonerHound | Typical VLM grounding |
|---|---|---|
| Reproducible | ✅ Same output every run | ❌ Varies with sampling |
| Auditable | ✅ Traceable to character offsets | ❌ Black-box neural net |
| Offline | ✅ No network required | ❌ Cloud API |
| Cost | ✅ $0.00 per page | ⚠️ $0.01–$0.40 per page |
| Accuracy (ExtractBench) | 72.62% Word F1 | 82.2% Word F1 (frontier) |
TonerHound trades ~10 points of grounding accuracy for reproducibility, auditability, and zero cost. That trade is worth it when you need to prove where a number came from.
git clone https://github.com/vanrajsinh650/TonerHound
cd TonerHound
uv sync
Requires Python 3.12+ and Tesseract OCR. Install Tesseract via brew install tesseract (macOS) or apt install tesseract-ocr (Linux).
import tonerhound
results = tonerhound.resolve(
document="invoice.pdf",
extraction=[
{"field": "invoice_number", "value": "INV-2024-0847"},
{"field": "total", "value": "$8,420.50"},
{"field": "due_date", "value": "2024-03-15"},
],
)
for res in results:
if res.is_grounded:
print(f"{res.field}: {res.status.value} → page {res.page}, bbox {res.bbox}")
else:
print(f"{res.field}: UNGROUNDED ({res.status.value})")
Output:
invoice_number: exact → page 1, bbox [0.12, 0.08, 0.21, 0.02]
total: normalized → page 3, bbox [0.41, 0.32, 0.09, 0.01]
due_date: exact → page 1, bbox [0.68, 0.08, 0.14, 0.02]
If a value cannot be grounded, res.is_grounded is False with status not_found or ambiguous, and res.bbox is None.
python -m tonerhound.benchmark --data-dir research/data/full --exp-id VERIFY
# Expected: Word Grounding F1 = 72.6179%
Evaluated on the official ExtractBench 370-document corpus (498,140 fields). Full results in docs/benchmark-results.md.
| Metric | Score |
|---|---|
| Word Grounding F1 | 72.6179% |
| Page Grounding F1 | 83.8490% |
| Word Precision | 77.7935% |
| Word Recall | 68.9729% |
| Passing Fields | 327,671 / 498,140 |
| Regressions (across 7 experiments) | 0 |
What drove the score (three techniques did ~75% of the work):
What didn't work (and why it matters): semantic normalization (2/37,850), column rail constraints (0), hyphen joiner v2 (0), Needleman-Wunsch alignment (13). These negative results define the deterministic ceiling. See docs/how-tonerhound-reached-72.md for the full research narrative.
PDF ──► pypdfium2 ──► character stream
│
├──► inverted token index
├──► numeric index
├──► date index
└──► OCR noise 3-gram index
│
Extracted JSON ──► Evidence Resolver
│
├── Hungarian table assignment
├── Multi-token sequence matching
├── Visual pixel-statistics fallback
├── Date literal variants
└── Multi-region assembly
│
▼
Grounded JSON + bounding boxes
Stack: pypdfium2 · pytesseract · numpy · scipy · OpenCV · PyMuPDF · rapidfuzz · python-dateutil
No PyTorch. No TensorFlow. No models. No training. No GPU.
TonerHound/
├── src/tonerhound/ # Production code
│ ├── document/ # PDF parsing, indexing, OCR routing
│ ├── resolution/ # Evidence resolver, Hungarian assignment
│ ├── matching/ # Token matching, date variants
│ ├── geometry/ # Bounding box assembly, multi-region
│ └── vision/ # Deterministic checkbox detection
├── tests/ # 261 unit tests (~20s)
├── benchmarks/ # Official ExtractBench runners
├── research/
│ ├── experiments/ # EXP-001 through EXP-043
│ └── observer/ # Failure Microscope V4
├── docs/
│ ├── benchmark-results.md
│ └── how-tonerhound-reached-72.md
├── pyproject.toml
├── uv.lock
└── LICENSE
Contributions welcome. Before opening a PR:
pytest tests/ -v # 261 tests must pass
python -m tonerhound.benchmark --exp-id YOUR-EXP # Score must not regress below 72.6179%
research/experiments/.MIT License. See LICENSE for details.
You can use TonerHound commercially, modify it, redistribute it, or embed it in proprietary software. Attribution is appreciated but not required.
If you use TonerHound in research, cite:
@software{tonerhound2026,
title = {TonerHound: Deterministic Evidence Grounding for LLM Document Extraction},
author = {Vanrajsinh},
year = {2026},
url = {https://github.com/vanrajsinh650/TonerHound}
}
ExtractBench (LlamaIndex) for the benchmark corpus. pypdfium2 for fast PDF character extraction. Tesseract for the OCR fallback. The scipy team for linear_sum_assignment.
Built without neural networks, by choice.