BOOTH A lightweight checkpoint layer for LLM outputs. BOOTH sits between your application and an LLM call and decides whether an answer should pass through, be reconsidered, be flagged as resting on more than one valid interpretation, or be marked uncertain.
8
stars
18
commits
Python
primary language
Sep 8, 2026
updated
A lightweight checkpoint library for LLM outputs.
BOOTH sits between your application and an LLM call and provides structured checkpoints for deciding whether an output should pass through, be reconsidered, be flagged as ambiguous, be checked against a custom validation rule, or be checked against evidence supplied by your application.
BOOTH does not claim to know the truth. It checks whether an output meets a defined acceptance condition.
The name comes from the idea of a ticket booth, toll booth, or parking/payment booth: a booth doesn't need to know everything about what is happening beyond it. It checks whether the required condition has been met before allowing something to pass.
v0.4.3
BOOTH currently provides:
validator for custom pass/fail rules on check()/acheck(), with its own distinct retry promptresult.method to identify which mechanism produced a result, and result.parsed to expose the model's raw JSON responseVERIFIED, REPAIRED, AMBIGUOUS, UNCERTAIN, and BLOCKED statusesBOOTH is provider-agnostic. It does not require a particular LLM provider, retrieval system, vector database, or framework.
A normal LLM call might look like:
answer = call_llm(prompt)
BOOTH adds a checkpoint around the model call:
import booth
result = booth.check(
call_llm,
"What is the capital of France?"
)
if result.ok:
print(result.answer)
else:
print(f"BOOTH returned {result.status}")
BOOTH asks the model to provide structured information about its response, including:
{
"ambiguous": false,
"interpretations": [],
"chosen_interpretation": null,
"answer": "Paris",
"confidence": 0.95
}
BOOTH then applies its configured acceptance rules to that response, in this order:
AMBIGUOUS immediately.validator was supplied and the answer fails it, BOOTH asks the model to reconsider, showing it the specific validation failure.If the model reaches the threshold (and passes validation, if supplied) after reconsideration, the result is REPAIRED.
If BOOTH cannot obtain an acceptable result, it returns UNCERTAIN.
BOOTH also provides check_with_evidence() for applications that already have evidence from their own RAG, search, database, or tool pipeline.
BOOTH asks the model to identify whether the question has multiple valid interpretations before accepting the answer.
For example:
What is the capital of Georgia?
could refer to:
Georgia (the country) -> Tbilisi
Georgia (the US state) -> Atlanta
BOOTH can return:
AMBIGUOUS
with the detected interpretations available through:
result.interpretations
Ambiguity takes priority over everything else BOOTH checks. A highly confident, validator-passing answer can still be returned as AMBIGUOUS if the model identifies multiple valid readings — and a validator is never even invoked on an ambiguous attempt.
BOOTH uses the model's reported confidence as an acceptance signal.
The default threshold is:
0.7
You can configure it:
result = booth.check(
call_llm,
prompt,
threshold=0.8
)
The confidence value is self-reported by the model. BOOTH does not calibrate or independently validate that probability.
When an answer is not ambiguous but its confidence is below the configured threshold, BOOTH can ask the model to reconsider its previous answer.
For example:
Previous answer: "Lyon"
Previous confidence: 0.3
Reconsider carefully. If that answer is correct, restate it.
If it is wrong, give the corrected answer.
If the reconsidered answer reaches the threshold, BOOTH returns:
REPAIRED
You can control the number of retries:
result = booth.check(
call_llm,
prompt,
max_retries=2
)
max_retries=0 means only the initial model call is made.
LLM responses do not always follow the requested format.
BOOTH handles parse failures separately from low-confidence answers.
When a response cannot be parsed, the retry prompt tells the model that its previous response failed to meet the required format rather than simply repeating the original request.
You can determine whether an UNCERTAIN result occurred because no response could ever be parsed:
result.all_parse_failed
A value of True means that every attempt failed to produce a valid BOOTH response.
validatorBOOTH can run a caller-supplied validation rule against each attempt's answer, in addition to (and checked separately from) ambiguity and confidence.
def is_valid_order_id(answer: str) -> bool:
return answer.strip().upper().startswith("ORD-")
result = booth.check(
call_llm,
"What is the order ID for this request?",
validator=is_valid_order_id,
)
validator receives the attempt's answer and returns either:
True / False — a plain pass/fail. False produces a generic failure message.(bool, str) — pass/fail plus a specific reason, shown to the model verbatim on the retry prompt:def validate_amount(answer: str):
if not answer.replace(".", "", 1).isdigit():
return False, "The answer must be a plain numeric amount, e.g. 42.50"
return True, ""
result = booth.check(call_llm, prompt, validator=validate_amount)
Ordering: an attempt is only run through validator if it parsed successfully and was not flagged ambiguous — a question that's ambiguous as asked isn't something a validator should be judging, and there is nothing to validate if the response never parsed. A validation failure is checked before the confidence gate: an answer that fails your validator does not get a chance to pass purely on high self-reported confidence.
An exception raised inside validator, or a return value that isn't bool or (bool, str), is treated as a failed validation — it never propagates out of check()/acheck().
validator=None (the default) is a true no-op: every code path this parameter introduces is unreachable if you never pass it, so existing calls are unaffected.
validator must be synchronous, for both check() and acheck(). If your validation logic needs to await something (an API call, a DB lookup), resolve it yourself first and pass a plain sync closure in.
result.methodTells you which of BOOTH's mechanisms actually determined a result, derived entirely from existing fields:
result.method
# "ambiguity" — status is AMBIGUOUS
# "evidence" — result came from check_with_evidence()
# "parse_failure" — UNCERTAIN because every attempt failed to parse
# "validation" — UNCERTAIN because the last attempt parsed fine
# and was confident enough, but failed your validator
# "confidence" — the ordinary case: VERIFIED / REPAIRED, or
# UNCERTAIN from persistent low confidence on an
# attempt that did parse and did pass validation
This is most useful for UNCERTAIN results, where it distinguishes genuinely different problems that call for different fixes:
if result.status == booth.UNCERTAIN:
if result.method == "parse_failure":
print("Model never produced a parseable response — check call_fn / prompt formatting.")
elif result.method == "validation":
print("Model was confident, but never satisfied the custom validator.")
else:
print("Model tried, but confidence never reached the threshold.")
method reflects the last attempt's determining factor for a mixed history (e.g. a parse failure followed by a validation failure reports "validation"), the same rule all_parse_failed already follows — it is not a full history of every attempt's outcome.
result.parsedExposes the model's raw JSON response object, exactly as returned — before any of BOOTH's own coercion (str(answer), float(confidence), forcing interpretations into a list of strings, and so on).
result = booth.check(call_llm, "What is the order ID for this request?")
result.answer # BOOTH's normalized answer, e.g. "ORD-4471"
result.parsed # the raw dict the model returned, unmodified
This is a transparency layer, not a second validation layer — BOOTH already decided answer/confidence/ambiguous from this same object; parsed exists so you can see exactly what the model said, including fields BOOTH itself doesn't use:
def call_llm(prompt: str) -> str:
# imagine the model, prompted to also include a source field,
# returns:
# {"answer": "30 days", "confidence": "0.95", "ambiguous": false,
# "source_document": "policy.pdf"}
...
result = booth.check(call_llm, prompt)
result.confidence # 0.95 (float — BOOTH coerced it)
result.parsed["confidence"] # "0.95" (str — the raw value, untouched)
result.parsed["source_document"] # "policy.pdf" — a field BOOTH never asked for or uses
That type divergence between result.confidence and result.parsed["confidence"] is deliberate, not a bug — parsed is meant to show you precisely what came back, even where BOOTH normalized it for its own use.
parsed reflects the attempt that determined the result: the winning attempt for VERIFIED/REPAIRED/AMBIGUOUS, the last successfully parsed attempt for UNCERTAIN, or None if no attempt ever parsed. check_with_evidence() results always have parsed=None — there is no LLM JSON parse involved on that path at all.
BOOTH provides both:
booth.check()
and:
await booth.acheck()
The synchronous version accepts:
Callable[[str], str]
The asynchronous version accepts:
Callable[[str], Awaitable[str]]
Example:
import asyncio
import booth
async def call_llm(prompt: str) -> str:
response = await async_client(...)
return response
async def main():
result = await booth.acheck(
call_llm,
"What is the capital of France?"
)
if result.ok:
print(result.answer)
asyncio.run(main())
Both APIs use the same decision logic, including validator and parsed. The difference is how the supplied LLM function is called.
BOOTH also provides:
booth.check_with_evidence()
This checks whether an answer agrees with evidence that your application has already retrieved.
Example:
result = booth.check_with_evidence(
answer="Paris is the capital of France.",
evidence=[
"France's capital city is Paris."
],
compare_fn=compare_answer_to_evidence,
)
The comparison function belongs to the caller:
def compare_answer_to_evidence(answer, evidence):
...
BOOTH does not choose a retrieval system or comparison algorithm for you.
The comparison function can return either True/False for a simple pass/fail comparison, or a float between 0.0 and 1.0:
0.87
When a float is returned, BOOTH compares it with evidence_threshold:
result = booth.check_with_evidence(
answer=answer,
evidence=evidence,
compare_fn=compare_answer_to_evidence,
evidence_threshold=0.8,
)
A score of 0.87 passes. A score of 0.62 does not.
Boolean comparison results are treated as strict pass/fail values. evidence_threshold is not applied to boolean results.
check_with_evidence() has no validator concept, and its results always have parsed=None — it is a standalone, single-purpose comparison gate, untouched by either the validator or parsed additions.
check_with_evidence() checks agreement with the evidence supplied to it.
It does not establish that the evidence itself is true.
For example, if your application retrieves an incorrect document:
Digital downloads are never eligible for refunds.
and your comparison function determines that the answer agrees with that document, BOOTH can return:
VERIFIED
That means the answer passed the supplied evidence comparison. It does not mean BOOTH independently established that the evidence is correct.
The quality, relevance, completeness, freshness, and correctness of retrieved evidence remain the responsibility of the application. This applies with equal force when evidence is baked into a prompt as RAG context and then separately checked — the model can produce a highly confident, unambiguous, evidence-agreeing answer that is still simply wrong, if the retrieved evidence itself was wrong. Neither check()'s confidence check nor check_with_evidence()'s agreement check can catch that; only the quality of retrieval can.
booth.check()booth.check(
call_fn,
prompt,
threshold=0.7,
max_retries=1,
on_attempt=None,
*,
validator=None,
)
Checks an LLM response using ambiguity detection, confidence checking, reconsideration, and (if supplied) a custom validator.
call_fnA synchronous function, Callable[[str], str], that receives a prompt and returns the model's raw response.
promptThe original application or user prompt.
thresholdMinimum self-reported confidence required to accept an unambiguous, validator-passing answer. Default 0.7. Must be between 0.0 and 1.0.
max_retriesNumber of retries after the initial attempt. Default 1.
on_attemptOptional callback invoked after each attempt.
validator (keyword-only, 0.4.2+)Optional Callable[[str], bool | tuple[bool, str]]. Runs on an attempt's answer only if that attempt parsed successfully and was not ambiguous. See Custom validation with validator above for the full contract. Must be synchronous. Default None — a true no-op.
booth.acheck()await booth.acheck(
call_fn,
prompt,
threshold=0.7,
max_retries=1,
on_attempt=None,
*,
validator=None,
)
Asynchronous equivalent of check(), including full validator support (still required to be synchronous itself). The supplied call_fn must be asynchronous:
async def call_llm(prompt: str) -> str:
...
booth.check_with_evidence()booth.check_with_evidence(
answer,
evidence,
compare_fn,
evidence_threshold=0.7,
)
Checks an answer against caller-supplied evidence. It:
BoothResultvalidator parameter — it is a standalone comparison gateparsed=None — there is no LLM JSON response involvedcompare_fnanswerThe answer being checked.
evidenceA sequence of evidence strings already retrieved by the application.
compare_fnA caller-supplied comparison function, Callable[[str, Sequence[str]], bool | float]. Receives answer and evidence, returns either a boolean or a score from 0.0 to 1.0.
evidence_thresholdMinimum score required when compare_fn returns a float. Default 0.7. Separate from check()'s threshold because the two values represent different things.
BOOTH returns a BoothResult.
Important fields include:
result.answer
result.status
result.confidence
result.evidence_agreement
result.attempts
result.n_attempts
result.ok
result.ambiguous
result.interpretations
result.all_parse_failed
result.method
result.parsed
answerThe answer produced by the model or supplied to the evidence checker. May be None when no usable answer exists.
statusOne of VERIFIED, REPAIRED, AMBIGUOUS, UNCERTAIN, BLOCKED.
confidenceFor normal LLM checks, the model's self-reported confidence. For evidence checks, the comparison score when available.
evidence_agreementThe comparison score produced by check_with_evidence(). None for normal check() / acheck() results.
attemptsThe full history of LLM attempts made by check() or acheck(), each including per-attempt passed_validation / validation_error (always True / None if no validator was supplied) and parsed (the raw JSON object for that specific attempt, None if it failed to parse). Evidence checks do not make attempts, so their attempt list is empty.
n_attemptsNumber of recorded attempts.
okTrue only for VERIFIED / REPAIRED. False for AMBIGUOUS, UNCERTAIN, BLOCKED.
ambiguousWhether the model marked the question as ambiguous.
interpretationsThe interpretations reported when the model marks a question as ambiguous.
all_parse_failedTrue if every LLM attempt failed to produce a parseable BOOTH response. Useful for distinguishing a formatting/integration problem from persistent model uncertainty or validation failure.
method (0.4.2+)Which mechanism produced the result — "ambiguity", "evidence", "parse_failure", "validation", or "confidence". See result.method above.
parsed (0.4.3+)The model's raw, uncoerced JSON response object. See result.parsed above.
VERIFIEDThe result passed BOOTH's acceptance condition on the relevant check. For normal LLM checking, the answer was not ambiguous, passed validation (if supplied), and met the confidence threshold on the initial attempt. For evidence checking, the supplied comparison passed. VERIFIED does not mean independently proven true.
REPAIREDThe initial LLM answer did not meet the confidence or validation requirement, but a reconsideration attempt produced an acceptable result.
AMBIGUOUSThe model identified multiple valid interpretations of the question. BOOTH returns this immediately rather than using a confidence retry or a validator to resolve it.
UNCERTAINBOOTH could not obtain an acceptable result. This can occur because:
validator on every attemptcheck_with_evidence() was emptyCheck result.method to tell these apart.
BLOCKEDThe supplied evidence comparison did not pass — a float score below evidence_threshold, or a boolean False from compare_fn.
BOOTH currently does not:
validator is itself correct — a validator can pass a wrong answer or reject a correct one, same as any other application-supplied ruleresult.parsed — it is exposed as-is, entirely unvalidatedvalidator gives you a documented hook to plug your own logic into BOOTH's retry loop rather than reimplementing that loop yourself)BOOTH is a checkpoint library, not an LLM framework, search engine, RAG framework, or autonomous verification system.
A model can report {"confidence": 0.99} and still be wrong. BOOTH does not independently calibrate that number.
Ambiguity detection depends on the model recognizing the ambiguity. BOOTH can detect useful structural ambiguities, but it cannot guarantee every possible interpretation is identified. A model can also mistake its own uncertainty for ambiguity.
validator is exactly as reliable as the logic you give it. BOOTH enforces that a validator's decision is respected consistently in the retry loop — it does not, and cannot, check whether the validator's own logic is actually correct for your use case.
result.parsed is unvalidatedresult.parsed is the model's raw JSON object, exposed as-is. BOOTH does not validate its shape, enforce a schema on it, or guarantee any field beyond the five it extracts for itself (answer, confidence, ambiguous, interpretations, chosen_interpretation) is present or well-typed. Reading extra fields from parsed is entirely the caller's responsibility, including handling missing keys or unexpected types.
Evidence checking is only as useful as the evidence and comparison function supplied by the application. If the evidence is wrong, incomplete, outdated, or unrelated, BOOTH does not independently detect that. Likewise, a weak compare_fn can produce a misleading result. This includes the case where retrieved evidence is baked into the model's own prompt as RAG context — a wrong document can make the model's answer both more confident and more evidence-consistent, without becoming more correct.
check_with_evidence() deliberately does not retrieve documents. The application owns retrieval:
Application
↓
Retrieve evidence
↓
BOOTH.check_with_evidence()
↓
VERIFIED / BLOCKED / UNCERTAIN
This keeps BOOTH small and provider-agnostic.
check_with_evidence() is a standalone evidence checkpoint. It does not automatically consume or modify the result of check() or acheck(). If an application wants to use multiple BOOTH checks together — including building a reconsideration loop that runs check() again after a BLOCKED evidence result — the application decides how those results should be combined and how many extra attempts that composition is allowed to cost. BOOTH's own max_retries only bounds a single check()/acheck() call; it has no visibility into, or control over, retries you build on top across multiple calls.
For example:
b_result = booth.check(call_llm, prompt)
if b_result.ok:
a_result = booth.check_with_evidence(
b_result.answer,
evidence,
compare_fn,
)
if a_result.ok:
print(a_result.answer)
The composition logic remains under application control.
Future BOOTH development may explore:
result.parsed — deferred deliberately until there's a clear, minimal shape for it, rather than reaching for a general schema/validation dependencyThese are future directions, not capabilities currently guaranteed by the library.
pip install boothpy
BOOTH is also installable directly from GitHub:
pip install git+https://github.com/Vedantgitbot/booth.git
Clone the repository and install the development dependencies:
pip install -e ".[dev]"
Run the test suite:
pytest
The test suite covers the core checkpoint behavior, asynchronous API, parsing behavior, ambiguity handling, reconsideration, custom validation, evidence checking, and raw-response exposure via result.parsed. CI runs the full suite on push/PR across Python 3.9–3.12.
The parsed tests specifically verify: the raw object matches the model's actual JSON (including cases where a field's type diverges from BOOTH's own coerced field, and cases where the model includes extra fields BOOTH doesn't use), correct behavior across VERIFIED/REPAIRED/AMBIGUOUS/UNCERTAIN, parsed reflecting the last successfully parsed attempt in a mixed history, parsed=None on total parse failure, parsed=None for every check_with_evidence() outcome, and check()/acheck() parity.
result.parsed is a transparency layer over data BOOTH already has, not a second parser or validator — it shows the model's raw response rather than deciding what it should mean.This is the official BOOTH repository — Vedant Brahmbhatt
BOOTH is released under the MIT License.
See LICENSE for the full license text.
18 commits
Python
100.0%
BOOTH A lightweight checkpoint layer for LLM outputs. BOOTH sits between your application and an LLM call and decides whether an answer should pass through, be reconsidered, be flagged as resting on more than one valid interpretation, or be marked uncertain.
8
stars
18
commits
Python
primary language
Sep 8, 2026
updated
A lightweight checkpoint library for LLM outputs.
BOOTH sits between your application and an LLM call and provides structured checkpoints for deciding whether an output should pass through, be reconsidered, be flagged as ambiguous, be checked against a custom validation rule, or be checked against evidence supplied by your application.
BOOTH does not claim to know the truth. It checks whether an output meets a defined acceptance condition.
The name comes from the idea of a ticket booth, toll booth, or parking/payment booth: a booth doesn't need to know everything about what is happening beyond it. It checks whether the required condition has been met before allowing something to pass.
v0.4.3
BOOTH currently provides:
validator for custom pass/fail rules on check()/acheck(), with its own distinct retry promptresult.method to identify which mechanism produced a result, and result.parsed to expose the model's raw JSON responseVERIFIED, REPAIRED, AMBIGUOUS, UNCERTAIN, and BLOCKED statusesBOOTH is provider-agnostic. It does not require a particular LLM provider, retrieval system, vector database, or framework.
A normal LLM call might look like:
answer = call_llm(prompt)
BOOTH adds a checkpoint around the model call:
import booth
result = booth.check(
call_llm,
"What is the capital of France?"
)
if result.ok:
print(result.answer)
else:
print(f"BOOTH returned {result.status}")
BOOTH asks the model to provide structured information about its response, including:
{
"ambiguous": false,
"interpretations": [],
"chosen_interpretation": null,
"answer": "Paris",
"confidence": 0.95
}
BOOTH then applies its configured acceptance rules to that response, in this order:
AMBIGUOUS immediately.validator was supplied and the answer fails it, BOOTH asks the model to reconsider, showing it the specific validation failure.If the model reaches the threshold (and passes validation, if supplied) after reconsideration, the result is REPAIRED.
If BOOTH cannot obtain an acceptable result, it returns UNCERTAIN.
BOOTH also provides check_with_evidence() for applications that already have evidence from their own RAG, search, database, or tool pipeline.
BOOTH asks the model to identify whether the question has multiple valid interpretations before accepting the answer.
For example:
What is the capital of Georgia?
could refer to:
Georgia (the country) -> Tbilisi
Georgia (the US state) -> Atlanta
BOOTH can return:
AMBIGUOUS
with the detected interpretations available through:
result.interpretations
Ambiguity takes priority over everything else BOOTH checks. A highly confident, validator-passing answer can still be returned as AMBIGUOUS if the model identifies multiple valid readings — and a validator is never even invoked on an ambiguous attempt.
BOOTH uses the model's reported confidence as an acceptance signal.
The default threshold is:
0.7
You can configure it:
result = booth.check(
call_llm,
prompt,
threshold=0.8
)
The confidence value is self-reported by the model. BOOTH does not calibrate or independently validate that probability.
When an answer is not ambiguous but its confidence is below the configured threshold, BOOTH can ask the model to reconsider its previous answer.
For example:
Previous answer: "Lyon"
Previous confidence: 0.3
Reconsider carefully. If that answer is correct, restate it.
If it is wrong, give the corrected answer.
If the reconsidered answer reaches the threshold, BOOTH returns:
REPAIRED
You can control the number of retries:
result = booth.check(
call_llm,
prompt,
max_retries=2
)
max_retries=0 means only the initial model call is made.
LLM responses do not always follow the requested format.
BOOTH handles parse failures separately from low-confidence answers.
When a response cannot be parsed, the retry prompt tells the model that its previous response failed to meet the required format rather than simply repeating the original request.
You can determine whether an UNCERTAIN result occurred because no response could ever be parsed:
result.all_parse_failed
A value of True means that every attempt failed to produce a valid BOOTH response.
validatorBOOTH can run a caller-supplied validation rule against each attempt's answer, in addition to (and checked separately from) ambiguity and confidence.
def is_valid_order_id(answer: str) -> bool:
return answer.strip().upper().startswith("ORD-")
result = booth.check(
call_llm,
"What is the order ID for this request?",
validator=is_valid_order_id,
)
validator receives the attempt's answer and returns either:
True / False — a plain pass/fail. False produces a generic failure message.(bool, str) — pass/fail plus a specific reason, shown to the model verbatim on the retry prompt:def validate_amount(answer: str):
if not answer.replace(".", "", 1).isdigit():
return False, "The answer must be a plain numeric amount, e.g. 42.50"
return True, ""
result = booth.check(call_llm, prompt, validator=validate_amount)
Ordering: an attempt is only run through validator if it parsed successfully and was not flagged ambiguous — a question that's ambiguous as asked isn't something a validator should be judging, and there is nothing to validate if the response never parsed. A validation failure is checked before the confidence gate: an answer that fails your validator does not get a chance to pass purely on high self-reported confidence.
An exception raised inside validator, or a return value that isn't bool or (bool, str), is treated as a failed validation — it never propagates out of check()/acheck().
validator=None (the default) is a true no-op: every code path this parameter introduces is unreachable if you never pass it, so existing calls are unaffected.
validator must be synchronous, for both check() and acheck(). If your validation logic needs to await something (an API call, a DB lookup), resolve it yourself first and pass a plain sync closure in.
result.methodTells you which of BOOTH's mechanisms actually determined a result, derived entirely from existing fields:
result.method
# "ambiguity" — status is AMBIGUOUS
# "evidence" — result came from check_with_evidence()
# "parse_failure" — UNCERTAIN because every attempt failed to parse
# "validation" — UNCERTAIN because the last attempt parsed fine
# and was confident enough, but failed your validator
# "confidence" — the ordinary case: VERIFIED / REPAIRED, or
# UNCERTAIN from persistent low confidence on an
# attempt that did parse and did pass validation
This is most useful for UNCERTAIN results, where it distinguishes genuinely different problems that call for different fixes:
if result.status == booth.UNCERTAIN:
if result.method == "parse_failure":
print("Model never produced a parseable response — check call_fn / prompt formatting.")
elif result.method == "validation":
print("Model was confident, but never satisfied the custom validator.")
else:
print("Model tried, but confidence never reached the threshold.")
method reflects the last attempt's determining factor for a mixed history (e.g. a parse failure followed by a validation failure reports "validation"), the same rule all_parse_failed already follows — it is not a full history of every attempt's outcome.
result.parsedExposes the model's raw JSON response object, exactly as returned — before any of BOOTH's own coercion (str(answer), float(confidence), forcing interpretations into a list of strings, and so on).
result = booth.check(call_llm, "What is the order ID for this request?")
result.answer # BOOTH's normalized answer, e.g. "ORD-4471"
result.parsed # the raw dict the model returned, unmodified
This is a transparency layer, not a second validation layer — BOOTH already decided answer/confidence/ambiguous from this same object; parsed exists so you can see exactly what the model said, including fields BOOTH itself doesn't use:
def call_llm(prompt: str) -> str:
# imagine the model, prompted to also include a source field,
# returns:
# {"answer": "30 days", "confidence": "0.95", "ambiguous": false,
# "source_document": "policy.pdf"}
...
result = booth.check(call_llm, prompt)
result.confidence # 0.95 (float — BOOTH coerced it)
result.parsed["confidence"] # "0.95" (str — the raw value, untouched)
result.parsed["source_document"] # "policy.pdf" — a field BOOTH never asked for or uses
That type divergence between result.confidence and result.parsed["confidence"] is deliberate, not a bug — parsed is meant to show you precisely what came back, even where BOOTH normalized it for its own use.
parsed reflects the attempt that determined the result: the winning attempt for VERIFIED/REPAIRED/AMBIGUOUS, the last successfully parsed attempt for UNCERTAIN, or None if no attempt ever parsed. check_with_evidence() results always have parsed=None — there is no LLM JSON parse involved on that path at all.
BOOTH provides both:
booth.check()
and:
await booth.acheck()
The synchronous version accepts:
Callable[[str], str]
The asynchronous version accepts:
Callable[[str], Awaitable[str]]
Example:
import asyncio
import booth
async def call_llm(prompt: str) -> str:
response = await async_client(...)
return response
async def main():
result = await booth.acheck(
call_llm,
"What is the capital of France?"
)
if result.ok:
print(result.answer)
asyncio.run(main())
Both APIs use the same decision logic, including validator and parsed. The difference is how the supplied LLM function is called.
BOOTH also provides:
booth.check_with_evidence()
This checks whether an answer agrees with evidence that your application has already retrieved.
Example:
result = booth.check_with_evidence(
answer="Paris is the capital of France.",
evidence=[
"France's capital city is Paris."
],
compare_fn=compare_answer_to_evidence,
)
The comparison function belongs to the caller:
def compare_answer_to_evidence(answer, evidence):
...
BOOTH does not choose a retrieval system or comparison algorithm for you.
The comparison function can return either True/False for a simple pass/fail comparison, or a float between 0.0 and 1.0:
0.87
When a float is returned, BOOTH compares it with evidence_threshold:
result = booth.check_with_evidence(
answer=answer,
evidence=evidence,
compare_fn=compare_answer_to_evidence,
evidence_threshold=0.8,
)
A score of 0.87 passes. A score of 0.62 does not.
Boolean comparison results are treated as strict pass/fail values. evidence_threshold is not applied to boolean results.
check_with_evidence() has no validator concept, and its results always have parsed=None — it is a standalone, single-purpose comparison gate, untouched by either the validator or parsed additions.
check_with_evidence() checks agreement with the evidence supplied to it.
It does not establish that the evidence itself is true.
For example, if your application retrieves an incorrect document:
Digital downloads are never eligible for refunds.
and your comparison function determines that the answer agrees with that document, BOOTH can return:
VERIFIED
That means the answer passed the supplied evidence comparison. It does not mean BOOTH independently established that the evidence is correct.
The quality, relevance, completeness, freshness, and correctness of retrieved evidence remain the responsibility of the application. This applies with equal force when evidence is baked into a prompt as RAG context and then separately checked — the model can produce a highly confident, unambiguous, evidence-agreeing answer that is still simply wrong, if the retrieved evidence itself was wrong. Neither check()'s confidence check nor check_with_evidence()'s agreement check can catch that; only the quality of retrieval can.
booth.check()booth.check(
call_fn,
prompt,
threshold=0.7,
max_retries=1,
on_attempt=None,
*,
validator=None,
)
Checks an LLM response using ambiguity detection, confidence checking, reconsideration, and (if supplied) a custom validator.
call_fnA synchronous function, Callable[[str], str], that receives a prompt and returns the model's raw response.
promptThe original application or user prompt.
thresholdMinimum self-reported confidence required to accept an unambiguous, validator-passing answer. Default 0.7. Must be between 0.0 and 1.0.
max_retriesNumber of retries after the initial attempt. Default 1.
on_attemptOptional callback invoked after each attempt.
validator (keyword-only, 0.4.2+)Optional Callable[[str], bool | tuple[bool, str]]. Runs on an attempt's answer only if that attempt parsed successfully and was not ambiguous. See Custom validation with validator above for the full contract. Must be synchronous. Default None — a true no-op.
booth.acheck()await booth.acheck(
call_fn,
prompt,
threshold=0.7,
max_retries=1,
on_attempt=None,
*,
validator=None,
)
Asynchronous equivalent of check(), including full validator support (still required to be synchronous itself). The supplied call_fn must be asynchronous:
async def call_llm(prompt: str) -> str:
...
booth.check_with_evidence()booth.check_with_evidence(
answer,
evidence,
compare_fn,
evidence_threshold=0.7,
)
Checks an answer against caller-supplied evidence. It:
BoothResultvalidator parameter — it is a standalone comparison gateparsed=None — there is no LLM JSON response involvedcompare_fnanswerThe answer being checked.
evidenceA sequence of evidence strings already retrieved by the application.
compare_fnA caller-supplied comparison function, Callable[[str, Sequence[str]], bool | float]. Receives answer and evidence, returns either a boolean or a score from 0.0 to 1.0.
evidence_thresholdMinimum score required when compare_fn returns a float. Default 0.7. Separate from check()'s threshold because the two values represent different things.
BOOTH returns a BoothResult.
Important fields include:
result.answer
result.status
result.confidence
result.evidence_agreement
result.attempts
result.n_attempts
result.ok
result.ambiguous
result.interpretations
result.all_parse_failed
result.method
result.parsed
answerThe answer produced by the model or supplied to the evidence checker. May be None when no usable answer exists.
statusOne of VERIFIED, REPAIRED, AMBIGUOUS, UNCERTAIN, BLOCKED.
confidenceFor normal LLM checks, the model's self-reported confidence. For evidence checks, the comparison score when available.
evidence_agreementThe comparison score produced by check_with_evidence(). None for normal check() / acheck() results.
attemptsThe full history of LLM attempts made by check() or acheck(), each including per-attempt passed_validation / validation_error (always True / None if no validator was supplied) and parsed (the raw JSON object for that specific attempt, None if it failed to parse). Evidence checks do not make attempts, so their attempt list is empty.
n_attemptsNumber of recorded attempts.
okTrue only for VERIFIED / REPAIRED. False for AMBIGUOUS, UNCERTAIN, BLOCKED.
ambiguousWhether the model marked the question as ambiguous.
interpretationsThe interpretations reported when the model marks a question as ambiguous.
all_parse_failedTrue if every LLM attempt failed to produce a parseable BOOTH response. Useful for distinguishing a formatting/integration problem from persistent model uncertainty or validation failure.
method (0.4.2+)Which mechanism produced the result — "ambiguity", "evidence", "parse_failure", "validation", or "confidence". See result.method above.
parsed (0.4.3+)The model's raw, uncoerced JSON response object. See result.parsed above.
VERIFIEDThe result passed BOOTH's acceptance condition on the relevant check. For normal LLM checking, the answer was not ambiguous, passed validation (if supplied), and met the confidence threshold on the initial attempt. For evidence checking, the supplied comparison passed. VERIFIED does not mean independently proven true.
REPAIREDThe initial LLM answer did not meet the confidence or validation requirement, but a reconsideration attempt produced an acceptable result.
AMBIGUOUSThe model identified multiple valid interpretations of the question. BOOTH returns this immediately rather than using a confidence retry or a validator to resolve it.
UNCERTAINBOOTH could not obtain an acceptable result. This can occur because:
validator on every attemptcheck_with_evidence() was emptyCheck result.method to tell these apart.
BLOCKEDThe supplied evidence comparison did not pass — a float score below evidence_threshold, or a boolean False from compare_fn.
BOOTH currently does not:
validator is itself correct — a validator can pass a wrong answer or reject a correct one, same as any other application-supplied ruleresult.parsed — it is exposed as-is, entirely unvalidatedvalidator gives you a documented hook to plug your own logic into BOOTH's retry loop rather than reimplementing that loop yourself)BOOTH is a checkpoint library, not an LLM framework, search engine, RAG framework, or autonomous verification system.
A model can report {"confidence": 0.99} and still be wrong. BOOTH does not independently calibrate that number.
Ambiguity detection depends on the model recognizing the ambiguity. BOOTH can detect useful structural ambiguities, but it cannot guarantee every possible interpretation is identified. A model can also mistake its own uncertainty for ambiguity.
validator is exactly as reliable as the logic you give it. BOOTH enforces that a validator's decision is respected consistently in the retry loop — it does not, and cannot, check whether the validator's own logic is actually correct for your use case.
result.parsed is unvalidatedresult.parsed is the model's raw JSON object, exposed as-is. BOOTH does not validate its shape, enforce a schema on it, or guarantee any field beyond the five it extracts for itself (answer, confidence, ambiguous, interpretations, chosen_interpretation) is present or well-typed. Reading extra fields from parsed is entirely the caller's responsibility, including handling missing keys or unexpected types.
Evidence checking is only as useful as the evidence and comparison function supplied by the application. If the evidence is wrong, incomplete, outdated, or unrelated, BOOTH does not independently detect that. Likewise, a weak compare_fn can produce a misleading result. This includes the case where retrieved evidence is baked into the model's own prompt as RAG context — a wrong document can make the model's answer both more confident and more evidence-consistent, without becoming more correct.
check_with_evidence() deliberately does not retrieve documents. The application owns retrieval:
Application
↓
Retrieve evidence
↓
BOOTH.check_with_evidence()
↓
VERIFIED / BLOCKED / UNCERTAIN
This keeps BOOTH small and provider-agnostic.
check_with_evidence() is a standalone evidence checkpoint. It does not automatically consume or modify the result of check() or acheck(). If an application wants to use multiple BOOTH checks together — including building a reconsideration loop that runs check() again after a BLOCKED evidence result — the application decides how those results should be combined and how many extra attempts that composition is allowed to cost. BOOTH's own max_retries only bounds a single check()/acheck() call; it has no visibility into, or control over, retries you build on top across multiple calls.
For example:
b_result = booth.check(call_llm, prompt)
if b_result.ok:
a_result = booth.check_with_evidence(
b_result.answer,
evidence,
compare_fn,
)
if a_result.ok:
print(a_result.answer)
The composition logic remains under application control.
Future BOOTH development may explore:
result.parsed — deferred deliberately until there's a clear, minimal shape for it, rather than reaching for a general schema/validation dependencyThese are future directions, not capabilities currently guaranteed by the library.
pip install boothpy
BOOTH is also installable directly from GitHub:
pip install git+https://github.com/Vedantgitbot/booth.git
Clone the repository and install the development dependencies:
pip install -e ".[dev]"
Run the test suite:
pytest
The test suite covers the core checkpoint behavior, asynchronous API, parsing behavior, ambiguity handling, reconsideration, custom validation, evidence checking, and raw-response exposure via result.parsed. CI runs the full suite on push/PR across Python 3.9–3.12.
The parsed tests specifically verify: the raw object matches the model's actual JSON (including cases where a field's type diverges from BOOTH's own coerced field, and cases where the model includes extra fields BOOTH doesn't use), correct behavior across VERIFIED/REPAIRED/AMBIGUOUS/UNCERTAIN, parsed reflecting the last successfully parsed attempt in a mixed history, parsed=None on total parse failure, parsed=None for every check_with_evidence() outcome, and check()/acheck() parity.
result.parsed is a transparency layer over data BOOTH already has, not a second parser or validator — it shows the model's raw response rather than deciding what it should mean.This is the official BOOTH repository — Vedant Brahmbhatt
BOOTH is released under the MIT License.
See LICENSE for the full license text.
18 commits
Python
100.0%