Open-ended Scientific Discovery via Bayesian Surprise
Asta Autodiscovery is an autonomous agent that performs data exploration on arbitrary datasets. The agent will generate hypotheses and run experiments to test each one. Surprising outcomes generate follow-up hypotheses in a recursive exploration.
Link to our NeurIPS 2025 paper: AutoDiscovery: Open-ended Scientific Discovery via Bayesian Surprise
Upgrading from 0.2.x
1.0.0 changes two things that every existing command line depends on:
- Model flags now name their provider.
--model gemini-3.1-pro-previewbecomes--model vertex_ai/gemini-3.7-flash. A bare model name is rejected at startup. See Selecting models below. - Vertex AI authenticates via Application Default Credentials only.
VERTEX_ACCESS_TOKEN,GOOGLE_OAUTH_ACCESS_TOKENandVERTEX_OPENAI_BASE_URLare no longer read.
Full notes: CHANGELOG.
Installation
Requires Python 3.13 or newer.
pip install asta-autodiscovery
This installs the auto-discovery command-line tool.
Quick start
Point auto-discovery at one or more dataset files and describe what you want explored:
auto-discovery \
--name "Plant growth study" \
--description "Field trial measurements of plant height under varying fertilizer dosage" \
--intent "Focus on dose-response relationships" \
--n_experiments 20 \
--out_dir ./results \
data/measurements.csv data/treatments.csv
CSV/TSV column headers are detected automatically. Datasets can also be directories — every file under them will be included.
Dataset files/directories can have different descriptions for each one listed. Use a repeated --dataset_description
parameter in place of the overall --description.
When the run finishes, a static HTML report is written to <out_dir>/report.
Common options
| Flag | Description |
|---|---|
--n_experiments |
Number of experiments to run (required). |
--out_dir |
Output directory for results and the HTML report (required). |
--name |
Short title for the run. |
--description |
Context about the dataset: provenance, collection method, known gaps. |
--domain |
Research domain (e.g. Genomics). |
--intent |
High-level exploration guidance for the agent. |
--dataset_description |
Per-dataset description; repeat once per dataset, in order. |
--exploration_weight |
Higher = broader exploration (default 2.0). |
--surprisal_width |
Surprise threshold; lower = more sensitive (default 0.2). |
Run auto-discovery --help to see the full set of options.
Authentication
All model traffic goes through litellm, and every model flag names
its provider explicitly as <provider>/<model> — for example
vertex_ai/gemini-3.7-flash, openai/o4-mini, github_copilot/claude-haiku-4.5. You only
need to configure the providers you actually name in --model, --belief_model,
--vision_model and --embedding_model.
Vertex AI
Used when a model flag names vertex_ai/.... The defaults are Vertex models, so this is required
unless you override them all.
Pick one of the following. In all cases, set the project and location so the agent knows which Vertex endpoint to call — these are litellm's own variable names, read by litellm itself:
export VERTEXAI_PROJECT=your-gcp-project-id
export VERTEXAI_LOCATION=global # `global` serves the default models
Both are required for a vertex_ai/ model, and a run that names one without them stops at
startup naming whichever is missing. Neither is inferred: Application Default Credentials carry
a project of their own, which is frequently not the project you meant, and litellm's fallback
location is us-central1, which does not serve the Gemini models the defaults name. Relying on
either turns a missing setting into a confusing 404 mid-run.
Service account key file (recommended for non-interactive use):
Create a service account in your GCP project, grant it the Vertex AI User role, download a
JSON key, and point Google's standard ADC env var at it:
export GOOGLE_APPLICATION_CREDENTIALS=/path/to/service-account-key.json
User credentials via gcloud (recommended for local development):
gcloud auth application-default login
Vertex authenticates via Application Default Credentials, so either
GOOGLE_APPLICATION_CREDENTIALS or gcloud auth application-default login is
required. Raw bearer tokens (VERTEX_ACCESS_TOKEN) are no longer accepted.
OpenAI
Used when a model flag names openai/... (e.g. openai/gpt-4o).
export OPENAI_API_KEY=sk-...
GitHub Copilot
Used when a model flag names github_copilot/.... litellm reads a GitHub OAuth token from a
file. On an interactive terminal it runs GitHub's device-code login on first use and
caches the token itself, so no setup is needed. For non-interactive runs, pre-seed it:
export GITHUB_COPILOT_TOKEN_DIR=/path/to/dir # must contain a file named `access-token`
There is no TTY detection, so a headless run without a cached token prints a device code and blocks for about three minutes before failing.
Selecting models
| Flag | What it controls | Default |
|---|---|---|
--model |
Primary reasoning model used for hypothesis generation and analysis. | vertex_ai/gemini-3.7-flash |
--belief_model |
Model used for belief updates over experimental outcomes. | vertex_ai/gemini-3.7-flash |
--vision_model |
Model used to interpret plots and figures emitted by experiments. | vertex_ai/gemini-3.7-flash |
--embedding_model |
Model used for deduplication embeddings. | openai/text-embedding-3-large |
Because the provider travels with each flag, mixing providers is supported — for example,
--model openai/gpt-4o --belief_model vertex_ai/gemini-3.7-flash uses OpenAI for the main
loop and Vertex AI for belief updates, with both OPENAI_API_KEY and the Vertex variables set.
Each flag is validated against litellm's offline model registry at startup, before the first model call.
Citation
If you find this work useful, please cite:
@inproceedings{
agarwal2025autodiscovery,
title={AutoDiscovery: Open-ended Scientific Discovery via Bayesian Surprise},
author={Dhruv Agarwal and Bodhisattwa Prasad Majumder and Reece Adamson and Megha Chakravorty and Satvika Reddy Gavireddy and Aditya Parashar and Harshit Surana and Bhavana Dalvi Mishra and Andrew McCallum and Ashish Sabharwal and Peter Clark},
booktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},
year={2025},
url={https://openreview.net/forum?id=kJqTkj2HhF}
}
Metadata
Release files for asta-autodiscovery 1.0.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| asta_autodiscovery-1.0.1.tar.gz | 86.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| asta_autodiscovery-1.0.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 182.0 kB
Release files / asta_autodiscovery-1.0.1.tar.gz
| Download URL | asta_autodiscovery-1.0.1.tar.gz |
|---|---|
| Size | 86.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
ea8d92f8602fb7684ad1f28893de8c178d56465ba6349013276b9d32d07ce588
|
|
BLAKE2b-256 checksum How to use checksums |
331ceae75bdef1889e34eb6eae25aa9ec0276f94b640e399162641782ffe7760
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 4, 2026.
Transparency logRelease files / asta_autodiscovery-1.0.1-py3-none-any.whl
| Download URL | asta_autodiscovery-1.0.1-py3-none-any.whl |
|---|---|
| Size | 95.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
a47e376f5962ac70e74c83c3a5cf324256b84c8545657be05fdb5fa487188919
|
|
BLAKE2b-256 checksum How to use checksums |
17a1c1f181a0b68882114b2aeba4e22684449501ea5e923dde89ce212a8a0cae
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 4, 2026.
Transparency log