vanrajsinh650/TonerHound

Deterministic PDF coordinate grounding for LLM extractions. Pin every extracted value to its exact source bbox - no ML, no API, no hallucination. 72.6% Word F1.

Python

1

80 commits

updated Oct 6, 2026

See the code

See what people are saying

SourceMessageScoreDate

Document grounding using PDF geometry, OCR and visual evidence — 72.6% Word F1 (r/computervision)

I've been working on document grounding where the goal is to map extracted values back to their exact locations in a PDF. The interesting constraint is that the grounding layer works from the document itself: * PDF character coordinates * OCR * spatial geometry * string/format matching * table…

1

Oct 6, 2026

README

TonerHound

Deterministic evidence grounding for LLM document extraction.

TonerHound resolves extracted values to exact physical PDF coordinates — no coordinate hallucination, no neural networks, no cloud APIs. Given any extracted JSON and its source PDF, it returns the precise bounding box for every value, or tells you the value cannot be grounded.

License: MIT Python 3.12+ Tests Word F1


What It Does

LLMs extract values from documents. They rarely say where those values came from. TonerHound closes that gap.

Extracted: {"total_revenue": "$8,420.50"}
     ↓ TonerHound
Grounded:  {"total_revenue": {"value": "$8,420.50",
                              "page": 47,
                              "bbox": [0.412, 0.318, 0.089, 0.011],
                              "status": "VERIFIED"}}

Every bounding box is deterministic: same input, same output, every time. No sampling, no temperature, no API calls. Runs on a laptop, offline.


Why Deterministic Grounding

PropertyTonerHoundTypical VLM grounding
Reproducible✅ Same output every run❌ Varies with sampling
Auditable✅ Traceable to character offsets❌ Black-box neural net
Offline✅ No network required❌ Cloud API
Cost✅ $0.00 per page⚠️ $0.01–$0.40 per page
Accuracy (ExtractBench)72.62% Word F182.2% Word F1 (frontier)

TonerHound trades ~10 points of grounding accuracy for reproducibility, auditability, and zero cost. That trade is worth it when you need to prove where a number came from.


Quick Start

Install

git clone https://github.com/vanrajsinh650/TonerHound
cd TonerHound
uv sync

Requires Python 3.12+ and Tesseract OCR. Install Tesseract via brew install tesseract (macOS) or apt install tesseract-ocr (Linux).

Resolve Evidence for Extracted Fields

import tonerhound

results = tonerhound.resolve(
    document="invoice.pdf",
    extraction=[
        {"field": "invoice_number", "value": "INV-2024-0847"},
        {"field": "total", "value": "$8,420.50"},
        {"field": "due_date", "value": "2024-03-15"},
    ],
)

for res in results:
    if res.is_grounded:
        print(f"{res.field}: {res.status.value} → page {res.page}, bbox {res.bbox}")
    else:
        print(f"{res.field}: UNGROUNDED ({res.status.value})")

Output:

invoice_number: exact → page 1, bbox [0.12, 0.08, 0.21, 0.02]
total:          normalized → page 3, bbox [0.41, 0.32, 0.09, 0.01]
due_date:       exact → page 1, bbox [0.68, 0.08, 0.14, 0.02]

If a value cannot be grounded, res.is_grounded is False with status not_found or ambiguous, and res.bbox is None.

Run the benchmark

python -m tonerhound.benchmark --data-dir research/data/full --exp-id VERIFY
# Expected: Word Grounding F1 = 72.6179%

Benchmark Summary

Evaluated on the official ExtractBench 370-document corpus (498,140 fields). Full results in docs/benchmark-results.md.

MetricScore
Word Grounding F172.6179%
Page Grounding F183.8490%
Word Precision77.7935%
Word Recall68.9729%
Passing Fields327,671 / 498,140
Regressions (across 7 experiments)0

What drove the score (three techniques did ~75% of the work):

  • Hungarian bipartite table assignment — globally optimal value-to-cell matching eliminates cascading row-swap errors. +5,845 fields.
  • Date literal variants — exact matching of 18 canonical date renderings. +3,690 fields.
  • Visual pixel-statistics fallback — deterministic checkbox detection via morphology and edge density. +1,882 fields.

What didn't work (and why it matters): semantic normalization (2/37,850), column rail constraints (0), hyphen joiner v2 (0), Needleman-Wunsch alignment (13). These negative results define the deterministic ceiling. See docs/how-tonerhound-reached-72.md for the full research narrative.


Architecture

PDF ──► pypdfium2 ──► character stream
          │
          ├──► inverted token index
          ├──► numeric index
          ├──► date index
          └──► OCR noise 3-gram index
          │
Extracted JSON ──► Evidence Resolver
                     │
                     ├── Hungarian table assignment
                     ├── Multi-token sequence matching
                     ├── Visual pixel-statistics fallback
                     ├── Date literal variants
                     └── Multi-region assembly
                     │
                     ▼
             Grounded JSON + bounding boxes

Stack: pypdfium2 · pytesseract · numpy · scipy · OpenCV · PyMuPDF · rapidfuzz · python-dateutil
No PyTorch. No TensorFlow. No models. No training. No GPU.


Repository Structure

TonerHound/
├── src/tonerhound/          # Production code
│   ├── document/            # PDF parsing, indexing, OCR routing
│   ├── resolution/          # Evidence resolver, Hungarian assignment
│   ├── matching/            # Token matching, date variants
│   ├── geometry/            # Bounding box assembly, multi-region
│   └── vision/              # Deterministic checkbox detection
├── tests/                   # 261 unit tests (~20s)
├── benchmarks/              # Official ExtractBench runners
├── research/
│   ├── experiments/         # EXP-001 through EXP-043
│   └── observer/            # Failure Microscope V4
├── docs/
│   ├── benchmark-results.md
│   └── how-tonerhound-reached-72.md
├── pyproject.toml
├── uv.lock
└── LICENSE

Contributing

Contributions welcome. Before opening a PR:

pytest tests/ -v  # 261 tests must pass
python -m tonerhound.benchmark --exp-id YOUR-EXP  # Score must not regress below 72.6179%
  • Keep it deterministic. No neural networks, no LLMs, no embeddings.
  • One experiment per PR. Document the delta in research/experiments/.
  • Negative results are welcome. They're the most valuable part of the research.

License

MIT License. See LICENSE for details.

You can use TonerHound commercially, modify it, redistribute it, or embed it in proprietary software. Attribution is appreciated but not required.


Citation

If you use TonerHound in research, cite:

@software{tonerhound2026,
  title = {TonerHound: Deterministic Evidence Grounding for LLM Document Extraction},
  author = {Vanrajsinh},
  year = {2026},
  url = {https://github.com/vanrajsinh650/TonerHound}
}

Acknowledgments

ExtractBench (LlamaIndex) for the benchmark corpus. pypdfium2 for fast PDF character extraction. Tesseract for the OCR fallback. The scipy team for linear_sum_assignment.

Built without neural networks, by choice.

vanrajsinh650/TonerHound

Deterministic PDF coordinate grounding for LLM extractions. Pin every extracted value to its exact source bbox - no ML, no API, no hallucination. 72.6% Word F1.

Python

1

80 commits

updated Oct 6, 2026

See the code

See what people are saying

SourceMessageScoreDate

Document grounding using PDF geometry, OCR and visual evidence — 72.6% Word F1 (r/computervision)

I've been working on document grounding where the goal is to map extracted values back to their exact locations in a PDF. The interesting constraint is that the grounding layer works from the document itself: * PDF character coordinates * OCR * spatial geometry * string/format matching * table…

1

Oct 6, 2026

README

TonerHound

Deterministic evidence grounding for LLM document extraction.

TonerHound resolves extracted values to exact physical PDF coordinates — no coordinate hallucination, no neural networks, no cloud APIs. Given any extracted JSON and its source PDF, it returns the precise bounding box for every value, or tells you the value cannot be grounded.

License: MIT Python 3.12+ Tests Word F1


What It Does

LLMs extract values from documents. They rarely say where those values came from. TonerHound closes that gap.

Extracted: {"total_revenue": "$8,420.50"}
     ↓ TonerHound
Grounded:  {"total_revenue": {"value": "$8,420.50",
                              "page": 47,
                              "bbox": [0.412, 0.318, 0.089, 0.011],
                              "status": "VERIFIED"}}

Every bounding box is deterministic: same input, same output, every time. No sampling, no temperature, no API calls. Runs on a laptop, offline.


Why Deterministic Grounding

PropertyTonerHoundTypical VLM grounding
Reproducible✅ Same output every run❌ Varies with sampling
Auditable✅ Traceable to character offsets❌ Black-box neural net
Offline✅ No network required❌ Cloud API
Cost✅ $0.00 per page⚠️ $0.01–$0.40 per page
Accuracy (ExtractBench)72.62% Word F182.2% Word F1 (frontier)

TonerHound trades ~10 points of grounding accuracy for reproducibility, auditability, and zero cost. That trade is worth it when you need to prove where a number came from.


Quick Start

Install

git clone https://github.com/vanrajsinh650/TonerHound
cd TonerHound
uv sync

Requires Python 3.12+ and Tesseract OCR. Install Tesseract via brew install tesseract (macOS) or apt install tesseract-ocr (Linux).

Resolve Evidence for Extracted Fields

import tonerhound

results = tonerhound.resolve(
    document="invoice.pdf",
    extraction=[
        {"field": "invoice_number", "value": "INV-2024-0847"},
        {"field": "total", "value": "$8,420.50"},
        {"field": "due_date", "value": "2024-03-15"},
    ],
)

for res in results:
    if res.is_grounded:
        print(f"{res.field}: {res.status.value} → page {res.page}, bbox {res.bbox}")
    else:
        print(f"{res.field}: UNGROUNDED ({res.status.value})")

Output:

invoice_number: exact → page 1, bbox [0.12, 0.08, 0.21, 0.02]
total:          normalized → page 3, bbox [0.41, 0.32, 0.09, 0.01]
due_date:       exact → page 1, bbox [0.68, 0.08, 0.14, 0.02]

If a value cannot be grounded, res.is_grounded is False with status not_found or ambiguous, and res.bbox is None.

Run the benchmark

python -m tonerhound.benchmark --data-dir research/data/full --exp-id VERIFY
# Expected: Word Grounding F1 = 72.6179%

Benchmark Summary

Evaluated on the official ExtractBench 370-document corpus (498,140 fields). Full results in docs/benchmark-results.md.

MetricScore
Word Grounding F172.6179%
Page Grounding F183.8490%
Word Precision77.7935%
Word Recall68.9729%
Passing Fields327,671 / 498,140
Regressions (across 7 experiments)0

What drove the score (three techniques did ~75% of the work):

  • Hungarian bipartite table assignment — globally optimal value-to-cell matching eliminates cascading row-swap errors. +5,845 fields.
  • Date literal variants — exact matching of 18 canonical date renderings. +3,690 fields.
  • Visual pixel-statistics fallback — deterministic checkbox detection via morphology and edge density. +1,882 fields.

What didn't work (and why it matters): semantic normalization (2/37,850), column rail constraints (0), hyphen joiner v2 (0), Needleman-Wunsch alignment (13). These negative results define the deterministic ceiling. See docs/how-tonerhound-reached-72.md for the full research narrative.


Architecture

PDF ──► pypdfium2 ──► character stream
          │
          ├──► inverted token index
          ├──► numeric index
          ├──► date index
          └──► OCR noise 3-gram index
          │
Extracted JSON ──► Evidence Resolver
                     │
                     ├── Hungarian table assignment
                     ├── Multi-token sequence matching
                     ├── Visual pixel-statistics fallback
                     ├── Date literal variants
                     └── Multi-region assembly
                     │
                     ▼
             Grounded JSON + bounding boxes

Stack: pypdfium2 · pytesseract · numpy · scipy · OpenCV · PyMuPDF · rapidfuzz · python-dateutil
No PyTorch. No TensorFlow. No models. No training. No GPU.


Repository Structure

TonerHound/
├── src/tonerhound/          # Production code
│   ├── document/            # PDF parsing, indexing, OCR routing
│   ├── resolution/          # Evidence resolver, Hungarian assignment
│   ├── matching/            # Token matching, date variants
│   ├── geometry/            # Bounding box assembly, multi-region
│   └── vision/              # Deterministic checkbox detection
├── tests/                   # 261 unit tests (~20s)
├── benchmarks/              # Official ExtractBench runners
├── research/
│   ├── experiments/         # EXP-001 through EXP-043
│   └── observer/            # Failure Microscope V4
├── docs/
│   ├── benchmark-results.md
│   └── how-tonerhound-reached-72.md
├── pyproject.toml
├── uv.lock
└── LICENSE

Contributing

Contributions welcome. Before opening a PR:

pytest tests/ -v  # 261 tests must pass
python -m tonerhound.benchmark --exp-id YOUR-EXP  # Score must not regress below 72.6179%
  • Keep it deterministic. No neural networks, no LLMs, no embeddings.
  • One experiment per PR. Document the delta in research/experiments/.
  • Negative results are welcome. They're the most valuable part of the research.

License

MIT License. See LICENSE for details.

You can use TonerHound commercially, modify it, redistribute it, or embed it in proprietary software. Attribution is appreciated but not required.


Citation

If you use TonerHound in research, cite:

@software{tonerhound2026,
  title = {TonerHound: Deterministic Evidence Grounding for LLM Document Extraction},
  author = {Vanrajsinh},
  year = {2026},
  url = {https://github.com/vanrajsinh650/TonerHound}
}

Acknowledgments

ExtractBench (LlamaIndex) for the benchmark corpus. pypdfium2 for fast PDF character extraction. Tesseract for the OCR fallback. The scipy team for linear_sum_assignment.

Built without neural networks, by choice.