Skip to main content

BOOTH

A lightweight checkpoint library for LLM outputs.

BOOTH sits between your application and an LLM call and provides structured checkpoints for deciding whether an output should pass through, be reconsidered, be flagged as ambiguous, or be checked against evidence supplied by your application.

BOOTH does not claim to know the truth. It checks whether an output meets a defined acceptance condition.

The name comes from the idea of a ticket booth, toll booth, or parking/payment booth: a booth doesn't need to know everything about what is happening beyond it. It checks whether the required condition has been met before allowing something to pass.


Current Status

v0.4.0

BOOTH currently provides:

  • ambiguity detection
  • self-reported confidence checking
  • reconsideration retries for low-confidence answers
  • separate retry handling for unparseable model responses
  • synchronous and asynchronous LLM checkpoint functions
  • evidence-agreement checking against evidence supplied by the caller
  • configurable confidence and evidence thresholds
  • structured result objects
  • attempt history
  • explicit VERIFIED, REPAIRED, AMBIGUOUS, UNCERTAIN, and BLOCKED statuses

BOOTH is provider-agnostic. It does not require a particular LLM provider, retrieval system, vector database, or framework.


What BOOTH Does

A normal LLM call might look like:

answer = call_llm(prompt)

BOOTH adds a checkpoint around the model call:

import booth

result = booth.check(
    call_llm,
    "What is the capital of France?"
)

if result.ok:
    print(result.answer)
else:
    print(f"BOOTH returned {result.status}")

BOOTH asks the model to provide structured information about its response, including:

{
  "ambiguous": false,
  "interpretations": [],
  "chosen_interpretation": null,
  "answer": "Paris",
  "confidence": 0.95
}

BOOTH then applies its configured acceptance rules to that response.

If the question is identified as ambiguous, BOOTH returns AMBIGUOUS.

If it is not ambiguous but the reported confidence is below the configured threshold, BOOTH can ask the model to reconsider its previous answer.

If the model reaches the threshold after reconsideration, the result is REPAIRED.

If BOOTH cannot obtain an acceptable result, it returns UNCERTAIN.

BOOTH also provides check_with_evidence() for applications that already have evidence from their own RAG, search, database, or tool pipeline.


Features

Ambiguity detection

BOOTH asks the model to identify whether the question has multiple valid interpretations before accepting the answer.

For example:

What is the capital of Georgia?

could refer to:

Georgia (the country) -> Tbilisi
Georgia (the US state) -> Atlanta

BOOTH can return:

AMBIGUOUS

with the detected interpretations available through:

result.interpretations

Ambiguity takes priority over confidence. A highly confident answer can still be returned as AMBIGUOUS if the model identifies multiple valid readings.


Confidence checking

BOOTH uses the model's reported confidence as an acceptance signal.

The default threshold is:

0.7

You can configure it:

result = booth.check(
    call_llm,
    prompt,
    threshold=0.8
)

The confidence value is self-reported by the model. BOOTH does not calibrate or independently validate that probability.


Reconsideration retries

When an answer is not ambiguous but its confidence is below the configured threshold, BOOTH can ask the model to reconsider its previous answer.

For example:

Previous answer: "Lyon"
Previous confidence: 0.3

Reconsider carefully. If that answer is correct, restate it.
If it is wrong, give the corrected answer.

If the reconsidered answer reaches the threshold, BOOTH returns:

REPAIRED

You can control the number of retries:

result = booth.check(
    call_llm,
    prompt,
    max_retries=2
)

max_retries=0 means only the initial model call is made.


Parse-failure handling

LLM responses do not always follow the requested format.

BOOTH handles parse failures separately from low-confidence answers.

When a response cannot be parsed, the retry prompt tells the model that its previous response failed to meet the required format rather than simply repeating the original request.

You can determine whether an UNCERTAIN result occurred because no response could ever be parsed:

result.all_parse_failed

A value of:

True

means that every attempt failed to produce a valid BOOTH response.


Synchronous and asynchronous APIs

BOOTH provides both:

booth.check()

and:

await booth.acheck()

The synchronous version accepts:

Callable[[str], str]

The asynchronous version accepts:

Callable[[str], Awaitable[str]]

Example:

import asyncio
import booth


async def call_llm(prompt: str) -> str:
    response = await async_client(...)
    return response


async def main():
    result = await booth.acheck(
        call_llm,
        "What is the capital of France?"
    )

    if result.ok:
        print(result.answer)


asyncio.run(main())

Both APIs use the same decision logic. The difference is how the supplied LLM function is called.


Evidence agreement checking

BOOTH also provides:

booth.check_with_evidence()

This checks whether an answer agrees with evidence that your application has already retrieved.

Example:

result = booth.check_with_evidence(
    answer="Paris is the capital of France.",
    evidence=[
        "France's capital city is Paris."
    ],
    compare_fn=compare_answer_to_evidence,
)

The comparison function belongs to the caller:

def compare_answer_to_evidence(answer, evidence):
    ...

BOOTH does not choose a retrieval system or comparison algorithm for you.

The comparison function can return either:

True

or:

False

for a simple pass/fail comparison.

It can also return a float between 0.0 and 1.0:

0.87

When a float is returned, BOOTH compares it with:

evidence_threshold

For example:

result = booth.check_with_evidence(
    answer=answer,
    evidence=evidence,
    compare_fn=compare_answer_to_evidence,
    evidence_threshold=0.8,
)

A score of 0.87 passes.

A score of 0.62 does not.

Boolean comparison results are treated as strict pass/fail values. evidence_threshold is not applied to boolean results.


Important: What Evidence Checking Means

check_with_evidence() checks agreement with the evidence supplied to it.

It does not establish that the evidence itself is true.

For example, if your application retrieves an incorrect document:

Digital downloads are never eligible for refunds.

and your comparison function determines that the answer agrees with that document, BOOTH can return:

VERIFIED

That means:

The answer passed the supplied evidence comparison.

It does not mean:

BOOTH independently established that the evidence is correct.

The quality, relevance, completeness, freshness, and correctness of retrieved evidence remain the responsibility of the application.


API

booth.check()

booth.check(
    call_fn,
    prompt,
    threshold=0.7,
    max_retries=1,
    on_attempt=None,
)

Checks an LLM response using ambiguity detection, confidence checking, and reconsideration.

call_fn

A synchronous function:

Callable[[str], str]

It receives a prompt and returns the model's raw response.

prompt

The original application or user prompt.

threshold

Minimum self-reported confidence required to accept an unambiguous answer.

Default:

0.7

Must be between 0.0 and 1.0.

max_retries

Number of retries after the initial attempt.

Default:

1

on_attempt

Optional callback invoked after each attempt.


booth.acheck()

await booth.acheck(
    call_fn,
    prompt,
    threshold=0.7,
    max_retries=1,
    on_attempt=None,
)

Asynchronous equivalent of check().

The supplied call_fn must be asynchronous:

async def call_llm(prompt: str) -> str:
    ...

booth.check_with_evidence()

booth.check_with_evidence(
    answer,
    evidence,
    compare_fn,
    evidence_threshold=0.7,
)

Checks an answer against caller-supplied evidence.

It:

  • makes no LLM calls
  • makes no network calls
  • performs no retrieval
  • performs no retries
  • does not modify a previous BoothResult
  • uses the caller's compare_fn

answer

The answer being checked.

evidence

A sequence of evidence strings already retrieved by the application.

compare_fn

A caller-supplied comparison function:

Callable[[str, Sequence[str]], bool | float]

It receives:

answer
evidence

and returns either a boolean or a score from 0.0 to 1.0.

evidence_threshold

Minimum score required when compare_fn returns a float.

Default:

0.7

It is separate from check()'s threshold because the two values represent different things.


Result Object

BOOTH returns a BoothResult.

Important fields include:

result.answer
result.status
result.confidence
result.evidence_agreement
result.attempts
result.n_attempts
result.ok
result.ambiguous
result.interpretations
result.all_parse_failed

answer

The answer produced by the model or supplied to the evidence checker.

May be None when no usable answer exists.

status

One of:

VERIFIED
REPAIRED
AMBIGUOUS
UNCERTAIN
BLOCKED

confidence

For normal LLM checks, this contains the model's self-reported confidence.

For evidence checks, it contains the comparison score when available.

evidence_agreement

The comparison score produced by check_with_evidence().

It is None for normal check() / acheck() results.

attempts

The full history of LLM attempts made by check() or acheck().

Evidence checks do not make attempts, so their attempt list is empty.

n_attempts

Number of recorded attempts.

ok

Returns:

True

only for:

VERIFIED
REPAIRED

It is False for:

AMBIGUOUS
UNCERTAIN
BLOCKED

ambiguous

Whether the model marked the question as ambiguous.

interpretations

The interpretations reported when the model marks a question as ambiguous.

all_parse_failed

Indicates that all LLM attempts failed to produce a parseable BOOTH response.

This is useful for distinguishing a formatting/integration problem from persistent model uncertainty.


Result Statuses

VERIFIED

The result passed BOOTH's acceptance condition on the relevant check.

For normal LLM checking, this means the answer was not marked ambiguous and met the confidence threshold on the initial attempt.

For evidence checking, this means the supplied comparison passed.

VERIFIED does not mean independently proven true.


REPAIRED

The initial LLM answer did not meet the confidence requirement, but a reconsideration attempt produced an acceptable result.


AMBIGUOUS

The model identified multiple valid interpretations of the question.

BOOTH returns this immediately rather than using a confidence retry to resolve it.


UNCERTAIN

BOOTH could not obtain an acceptable result.

This can occur because:

  • the model remained below the confidence threshold
  • every response failed to parse
  • the answer or evidence supplied to check_with_evidence() was empty
  • the evidence comparison function raised an exception
  • the evidence comparison function returned an invalid score

BLOCKED

The supplied evidence comparison did not pass.

For example, a float comparison score below the configured evidence_threshold produces:

BLOCKED

A boolean False from compare_fn also produces:

BLOCKED

What BOOTH Does Not Do

BOOTH currently does not:

  • guarantee factual correctness
  • independently establish truth
  • automatically browse the web
  • automatically perform RAG
  • automatically retrieve evidence
  • automatically choose a vector database
  • automatically choose an evidence-comparison method
  • retry evidence retrieval
  • manage a tool-calling loop
  • compare multiple independent LLMs
  • provide calibrated confidence probabilities
  • guarantee that retrieved evidence is correct, complete, relevant, or current
  • replace application-specific validation or safety systems

BOOTH is a checkpoint library, not an LLM framework, search engine, RAG framework, or autonomous verification system.


Current Limitations

Self-reported confidence

Confidence in normal LLM checking comes from the model itself.

A model can report:

{
  "confidence": 0.99
}

and still be wrong.

BOOTH does not independently calibrate that number.


Self-reported ambiguity

Ambiguity detection also depends on the model recognizing the ambiguity.

BOOTH can detect useful structural ambiguities, but it cannot guarantee that every possible interpretation is identified.

A model can also mistake its own uncertainty for ambiguity.


Evidence quality

Evidence checking is only as useful as the evidence and comparison function supplied by the application.

If the evidence is wrong, incomplete, outdated, or unrelated, BOOTH does not independently detect that.

Likewise, a weak compare_fn can produce a misleading result.


No automatic retrieval

check_with_evidence() deliberately does not retrieve documents.

The application owns retrieval:

Application
    ↓
Retrieve evidence
    ↓
BOOTH.check_with_evidence()
    ↓
VERIFIED / BLOCKED / UNCERTAIN

This keeps BOOTH small and provider-agnostic.


No automatic reconciliation

check_with_evidence() is a standalone evidence checkpoint.

It does not automatically consume or modify the result of check() or acheck().

If an application wants to use multiple BOOTH checks together, the application decides how those results should be combined.

For example:

b_result = booth.check(call_llm, prompt)

if b_result.ok:
    a_result = booth.check_with_evidence(
        b_result.answer,
        evidence,
        compare_fn,
    )

    if a_result.ok:
        print(a_result.answer)

The composition logic remains under application control.


Future Plans

Future BOOTH development may explore:

  • stronger evidence adequacy checks
  • better handling of evidence completeness
  • improved detection of convention-based ambiguity
  • methods for distinguishing genuine ambiguity from model uncertainty
  • additional evidence-comparison strategies
  • richer composition of multiple checkpoint results
  • better evaluation and calibration tooling
  • additional integrations with retrieval and tool systems

These are future directions, not capabilities currently guaranteed by the library.


Installation

pip install boothpy

BOOTH is also installable directly from GitHub:

pip install git+https://github.com/Vedantgitbot/booth.git

Development

Clone the repository and install the development dependencies:

pip install -e ".[dev]"

Run the test suite:

pytest

The test suite covers the core checkpoint behavior, asynchronous API, parsing behavior, ambiguity handling, reconsideration, and evidence checking.

The evidence-checking tests include cases for:

  • passing float scores
  • failing float scores
  • boolean comparison
  • boolean False with a zero threshold
  • empty answers
  • empty evidence
  • comparison exceptions
  • invalid comparison scores
  • non-numeric comparison results
  • invalid evidence thresholds
  • result-field behavior
  • independent evidence thresholds

Design Principles

  1. Keep the checkpoint small. BOOTH should provide a reusable decision layer rather than become another full LLM framework.

  2. Make uncertainty explicit. When an output does not meet the configured acceptance condition, return a structured status instead of silently passing it through.

  3. Treat ambiguity separately from confidence. A confident answer can still be ambiguous if the question has multiple valid interpretations.

  4. Reconsider instead of blindly resampling. Retries give the model an opportunity to examine its previous response.

  5. Keep evidence retrieval outside BOOTH. Applications remain free to use their own RAG, search, database, or tool infrastructure.

  6. Do not pretend agreement is truth. Agreement with an answer, confidence value, or retrieved evidence is not the same as independently proving the claim.

  7. Stay provider-agnostic. BOOTH works with different LLM providers because the application supplies the model-calling function.


License

This is the official BOOTH repository — Vedant Brahmbhatt

BOOTH is released under the MIT License.

See LICENSE for the full license text.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

boothpy-0.4.0.tar.gz (29.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

boothpy-0.4.0-py3-none-any.whl (17.2 kB view details)

Uploaded Python 3

File details

Details for the file boothpy-0.4.0.tar.gz.

File metadata

  • Download URL: boothpy-0.4.0.tar.gz
  • Upload date:
  • Size: 29.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.12.7

File hashes

Hashes for boothpy-0.4.0.tar.gz
Algorithm Hash digest
SHA256 bfd5ee7256035a6ca35276751f1aacae9b57173b4b6bab67b63016e0f1bc034f
MD5 692b988b6b2a274d633eec1230e9bd8c
BLAKE2b-256 21993a6dd7ef43374a60c4aa30407abc25a00878494b2be87cbc6136deee4928

See more details on using hashes here.

File details

Details for the file boothpy-0.4.0-py3-none-any.whl.

File metadata

  • Download URL: boothpy-0.4.0-py3-none-any.whl
  • Upload date:
  • Size: 17.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.12.7

File hashes

Hashes for boothpy-0.4.0-py3-none-any.whl
Algorithm Hash digest
SHA256 2338eedc0012dd4766fbc27dd96355cb71a48cb3788555bc0ca57e353bbe20d0
MD5 0b63c468da034ccfe3acae6997da02fe
BLAKE2b-256 245f0d81f9a6d67f826b2c7989dec74b01b1b54247bd9a08c368c8ebd1b52643

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

0.4.0 This release

2 files

0.3.1

2 files

0.3.0

2 files

0.2.0

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page