Skip to main content

passk-inference

Do small-budget gains after RL persist at larger sampling budgets? Compare the same prompts across base and RL checkpoints with a simultaneous confidence band over the sampling-budget grid. Identify supported gains, supported losses, and a confidence set for the first loss budget.

Paper · Code and API · Real example

This is a tool for comparing checkpoints and sampling budgets on a specified task population and decoding setup. It does not measure general capability or include uncertainty across independent training runs.

Real DeepScaleR comparison

Install from PyPI

python -m pip install "passk-inference[report]"
passk-inference --base base.jsonl --rl rl.jsonl --output output/comparison

Use python -m pip install passk-inference for the API and JSON-only CLI. The wheel contains the software; the repository and source release include the real and synthetic example datasets and scripts.

Run the paper example on a laptop

With Python 3.10 or later:

git clone https://github.com/cafferychen777/passk-inference.git
cd passk-inference
python -m venv .venv
source .venv/bin/activate  # Windows: .venv\Scripts\activate
python -m pip install '.[example]'
python examples/deepscaler32k/reproduce.py

No GPU, model download, cluster account, or new generation is needed. The input contains real sampled counts for 1,060 paired prompts. The command checks data hashes, recomputes the analysis, checks the paper result, and writes output/deepscaler32k/{result.json,curves.csv,comparison.png}.

Budget k Evidence from the 95% simultaneous band
1–10 RL improves pass@k
11–60 Evidence is insufficient to determine the sign
61–128 RL reduces pass@k

The first-loss confidence set is [11, 61]. It is not a precise crossover point or a universal deployment threshold. This is the paper's post hoc dense grid analysis; the original sparse-grid test and its p-value are distinct. See data provenance and paper mapping.

Use your own counts

Provide two JSONL files. Each row has a unique prompt ID, integer successes c, and integer trials n, for example {"id":"problem-1","c":5,"n":128}. Both files must contain exactly the same IDs; row order does not matter.

python -m pip install '.[report]'
passk-inference --base base.jsonl --rl rl.jsonl --output output/my-comparison

This writes result.json, curves.csv and comparison.png, using the same report implementation as the paper example. --base-label and --rl-label set figure labels. Files with these names in the output directory are replaced. JSON is also printed to stdout; without --output, the CLI remains JSON-only and does not require Matplotlib:

passk-inference --base base.jsonl --rl rl.jsonl --bootstrap 4000 --seed 0 > result.json

Every result records the package version, result format version and analysis settings. File-based comparisons also record SHA-256 hashes of the exact input bytes consumed, by base/RL role, without local paths. See the result contract.

By default, the grid is every integer from 1 through the smallest trial count. Differences are RL minus base. Output includes both curves, the simultaneous band, supported gain/loss budgets, inconclusive budgets, and the first-loss set. Supplying a sparse --ks grid disables the all-integer first-loss interval. The procedure assumes independent prompt rows and the sampling assumptions in the API reference. Insufficient evidence is not equivalence.

The eight-prompt files examples/base.jsonl and examples/rl.jsonl are synthetic teaching data, separate from the real example above. The optional response-kernel API is a working model for prediction; it is not needed for the model-free paper result.

Check statistical behavior under known truths

The synthetic coverage example varies prompt count, rare success and heterogeneity, and reports whole-grid coverage, crossing power or false detection, and Monte Carlo intervals. It is independent of the real-data example and is not run on every CI build:

python examples/coverage/simulate.py

Development and release

python -m pip install -e '.[test,example]' build
python -m pytest
python -m build
python scripts/check_sdist.py dist/passk_inference-0.3.1.tar.gz
python scripts/export_release.py

Repository layout and publication policy explains what is included, what stays private, and how to create a versioned release. Code and original aggregate-count artifacts are under the MIT license; third-party benchmark text and model weights are not redistributed.

Metadata

Release files for passk-inference 0.3.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for passk-inference 0.3.1
File Size Uploaded
passk_inference-0.3.1.tar.gz 286.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for passk-inference 0.3.1
File Interpreter ABI Platform
passk_inference-0.3.1-py3-none-any.whl Python 3 none any Details

Total release size: 300.0 kB

Release files / passk_inference-0.3.1.tar.gz

Download URL passk_inference-0.3.1.tar.gz
Size 286.0 kB
Tags Source
SHA-256 checksum
How to use checksums
e279dcd98a910624fb06624a6b5c8a9937cb34d176987495bada4f5ff2aaa418
BLAKE2b-256 checksum
How to use checksums
a0c48a87345d2bcf60f63ae6afa02424dcad4ed78ebc64f6888574ecd6e7d42c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.

Transparency log

Release files / passk_inference-0.3.1-py3-none-any.whl

Download URL passk_inference-0.3.1-py3-none-any.whl
Size 14.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f32892f7eed22de2abc43b9adb471d1ada20cfa46999d4022697df36ff9c3e6e
BLAKE2b-256 checksum
How to use checksums
b99d198076157c847d04dc4cbcf6ef2f654a77f12f562cfd5d18b2b4ee031d88
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.3.1 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page