Hypothesis Generation
Generate novel, testable hypotheses without an LLM.
Hypothesis Generation is a deterministic, programmable engine for scientists, engineers, mathematicians, and research software teams. It constructs candidate hypotheses from explicit observations, concepts, relations, constraints, mechanisms, analogies, and domain knowledge.
The engine runs locally. It has no runtime model dependency, makes no network requests, and does not sample from a pretrained distribution. The same request, engine version, extension set, domain pack, and corpus snapshot produces the same result.
Why it is different
Language models generate text from learned statistical patterns. This project uses typed transformations, finite formal reasoning, constraint checks, bounded search, and quality-diversity selection to construct hypotheses that were not supplied as templates.
That makes each result reproducible and inspectable:
- domain facts come only from declared inputs, traceable corpora, formal rules, or reviewed domain packs;
- invalid candidates can be rejected by explicit type, logic, graph, dimensional, contradiction, and mechanism checks;
- requested portfolios are selected for structural coverage instead of padded with paraphrases;
- novelty diagnostics are always relative to an identified corpus; and
- derivations, scores, provenance, tests, and archives are available when a caller requests them.
The engine proposes candidates for investigation. Experiments, prior-art review, safety review, and scientific judgment remain external responsibilities.
Install
Python 3.10 or later is required. The core has no runtime dependencies.
python -m pip install -e .
The repository version can also be installed directly:
python -m pip install "hypothesis-generation @ git+https://github.com/cdimurro/hypothesis-generation.git"
Quick start
The common API requires a phenomenon and a count:
from hypothesis_generation import generate_hypotheses
hypotheses = generate_hypotheses(
"urban heat islands remain hotter overnight after heat waves",
count=8,
)
for hypothesis in hypotheses:
print(hypothesis)
The default output is a list of English hypotheses. Backend artifacts remain hidden unless explicitly requested.
The equivalent command-line request is:
{
"phenomenon": "urban heat islands remain hotter overnight after heat waves",
"count": 8
}
python -m hypothesis_generation --input request.json --pretty
Add scientific structure
More explicit inputs produce more specific hypotheses:
from hypothesis_generation import ConceptSpec, DiscoveryRequest, RelationSpec, discover
request = DiscoveryRequest(
phenomenon="nighttime cooling slows after a heat wave",
count=4,
concepts=(
ConceptSpec("impervious surface fraction", "DRIVER"),
ConceptSpec("stored subsurface heat", "MEDIATOR"),
ConceptSpec("nighttime cooling rate", "OUTCOME"),
),
relations=(
RelationSpec(
"impervious surface fraction",
"increases",
"stored subsurface heat",
),
),
)
result = discover(request)
print(result.as_dict())
Requests can also declare known facts, observations, contexts, constraints, prior hypotheses, analogy sources, causal structures, quantities, domain packs, and a frozen prior-art corpus. Undeclared synonyms and domain facts are never guessed.
See the discovery API for request fields, response projection, search budgets, persistence, and exhaustion behavior.
How generation works
- Normalize the request. Convert declared knowledge into typed concepts, relations, constraints, mechanisms, quantities, and provenance records.
- Invent premises. Apply typed transformations and bounded multi-step programs to propose mechanisms, missing bridges, alternatives, boundary conditions, analogies, interventions, and counterexamples.
- Generate candidates. Explore fourteen complementary reasoning modes.
- Test and constrain. Reject candidates that fail applicable formal or declared scientific checks.
- Check prior art. Compare candidates with a frozen local corpus using exact, lexical, BM25, relation, and structural diagnostics.
- Select a portfolio. Optimize quality and structural coverage under a declared ranking profile.
- Render the result. Return English hypotheses by default and only the requested artifact fields when details are needed.
The fourteen public modes cover causal, abductive, temporal, systems, regime, heterogeneity, measurement, invariant, contrastive, intervention, null-model, relational, analogical, and compositional reasoning. They are search strategies over one integrated premise-invention system, not sentence templates.
Large hypothesis campaigns
Persistent campaigns can generate exact portfolios of up to 10,000 unique hypotheses when the admitted search pool is large enough. Results are returned through deterministic, tamper-evident pages. If the requested total cannot be produced without duplicates or cosmetic rewrites, the campaign reports search exhaustion before returning a partial result.
Portfolio profiles include:
BALANCEDDIVERSITY_FIRSTNOVELTY_FIRSTPLAUSIBILITY_FIRSTTESTABILITY_FIRST- bounded caller-declared custom weights
Profiles change transparent ranking objectives; they do not load a learned ranking model.
Prior-art comparison
The optional local prior-art index supports exact overlap, phrase containment, token similarity, Okapi BM25, declared aliases, relation overlap, metadata filters, date-bounded corpus views, and SQLite/FTS5 storage.
These diagnostics answer whether a candidate overlaps the supplied corpus. They do not claim exhaustive or universal novelty.
Program it for an application
There are two extension paths:
- Domain packs add reviewed, data-only concepts, relations, rules, sources, aliases, constraints, and generative vocabulary.
- Extension bundles add deterministic application code through content-addressed manifests that bind source files, dependencies, and runtime policy.
Starter domain-pack scaffolds are available for engineered-system reliability, materials/process engineering, and systems biology:
hypothesis-generate --scaffold-domain-pack materials-process-engineering `
--output-directory packs/materials-process
Scaffolds contain no facts or approvals and remain AWAITING_REVIEW until the
completed digest receives the required reviews. See domain packs
and extensions.
Optional scientific-computing and formal-solver adapters can be supplied by a caller. They are never discovered, installed, or invoked implicitly.
Evaluation
The repository includes structural, historical, cross-domain, external, and negative-control benchmarks. Selected pinned results are:
| Evaluation | Result | Scope |
|---|---|---|
| Release qualification | 9/9 controls | Runtime boundary, campaigns, ranking, novelty, extensions, packs, evaluation, and legacy behavior |
| HypoSpace | 441/441 admissible targets | 88 Boolean, causal-graph, and voxel cases |
| HypoBench synthetic | 1,450/1,500 held-out rows | Nine classification cases |
| HypoBench real-world | 2,003/3,100 held-out; 1,431/2,593 OOD | Seven text-classification tasks |
| DiscoveryBench expanded | 8/24 strong; median sealed-test R-squared 0.727 | Systematic synthetic task sample |
| SRSD-Feynman portfolio | 37/50 strong; median test R-squared 0.876 | Recursive symbolic search with sealed test rows |
| External Premise Challenge | 11/12 generated and admitted; 5/12 public top-10 | Dated historical diagnostic across six fields |
These measurements apply only to the pinned files, grammars, budgets, and evaluation rules. They are not estimates of general scientific truth or expert preference. External metrics that require an LLM judge are reported as not run.
See the benchmark guide, evaluation kit, and scientific limitations.
Release status
Version 0.5 is a public beta. The deterministic core, packaging, compatibility, and benchmark contracts are qualified; the complete suite passes 538 tests. Version 1.0 is reserved for stable contracts informed by independent users, reviewed reference domain packs, blind expert evaluation, and prospective evidence in multiple fields.
The legacy v0.3 typed mathematical-model API remains supported. See release readiness, release qualification, and the release plan.
Documentation
Start with the documentation index. Important references are:
- Architecture and runtime boundary
- Discovery API
- Domain packs
- Application extensions
- Agent protocol
- Scientific capabilities and limitations
- Release qualification
Contributing
Read CONTRIBUTING.md before adding an operator, family, domain pack, adapter, or benchmark. Security guidance is in SECURITY.md.
Hypothesis Generation is released under the Apache License 2.0.
Metadata
Release files for hypothesis-generation 0.5.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| hypothesis_generation-0.5.0.tar.gz | 640.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| hypothesis_generation-0.5.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 1.2 MB
Release files / hypothesis_generation-0.5.0.tar.gz
| Download URL | hypothesis_generation-0.5.0.tar.gz |
|---|---|
| Size | 640.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
2d9bba1960e61ab97b3ba6263a905cfa89bfa2262557a0645da6fce511875808
|
|
BLAKE2b-256 checksum How to use checksums |
c91a44ccf0266c162569aec5f4feea6b33136cdbbaa3a98d76cb5aaa893f6a71
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 2, 2026.
Transparency logRelease files / hypothesis_generation-0.5.0-py3-none-any.whl
| Download URL | hypothesis_generation-0.5.0-py3-none-any.whl |
|---|---|
| Size | 540.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
8b197101439c37d14929715c5eed9904b5a969fb6ecc2d0e8bb8c67937785eb9
|
|
BLAKE2b-256 checksum How to use checksums |
82431c524fd7409f4c7ab693b49fd32c5058f2b48ff77df455064b5aca18177e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 2, 2026.
Transparency log