Test whether an AI agent can complete your quickstart using only your docs
Project description
quickstarted
Test whether an AI agent can complete your quickstart using only your docs.
Roughly half the traffic to documentation sites is now AI agents and AI-assisted workflows (Mintlify's measurement). When an agent reads your quickstart and fails, you lose the developer behind it and you never find out. quickstarted turns that failure into a CI signal: a sandboxed agent gets your goal, your docs, and nothing else, and a script you wrote decides whether it succeeded.
Full documentation: https://snehankekre.com/quickstarted/
Status: early. The interface will change.
Try it in one command
uvx quickstarted run --example streamlit --agent replay
No API key, no cost, and nothing to clone:
[ 0s] started on seatbelt
[ 2s] read https://docs.streamlit.io/get-started
[ 19s] blocked from the shell: checkip.amazonaws.com
[ 20s] check exited 0
[PASS] streamlit-quickstart (replay)
classification: passed
turns: 2, duration: 19.6s
backend: seatbelt
docs pages read: 1
That ran the commands Streamlit's documentation tells a reader to type, booted
the app, and polled its health endpoint. The third line is the sandbox refusing
Streamlit's own call home, which is what enforcement looks like from the
outside. quickstarted examples lists the others.
How it works
You write a task, in YAML. quickstarted init <docs-url> scaffolds one:
name: streamlit-quickstart
goal: >
Install Streamlit into the virtualenv in this workspace and build a small app
in app.py that shows a title and a slider whose value is displayed on the
page. Do not start a long-running server yourself; just leave app.py ready to
run and confirm streamlit is installed.
docs:
entrypoint: https://docs.streamlit.io/get-started
allow:
- docs.streamlit.io
- streamlit.io
setup:
- python3 -m venv .venv
success:
script: |
set -e
test -f app.py || qs_fail "no app.py, so the quickstart produced nothing"
.venv/bin/python -c "import streamlit"
Then you run an agent against it:
pip install "quickstarted[claude]"
export QUICKSTARTED_ANTHROPIC_API_KEY=...
quickstarted run tasks/streamlit-quickstart.yaml --agent claude
[PASS] streamlit-quickstart (claude:claude-opus-5)
classification: passed
turns: 6, duration: 63.5s
backend: docker
tokens: 12 in / 1150 out, cache 11204 written / 33181 read
docs pages read: 4
The agent works in a throwaway workspace with two tools: a sandboxed bash,
and read_docs, which is the only way it can read documentation. It gets no
browser, no search engine, and no network of its own. The system prompt also
forbids leaning on prior knowledge of the project under test, because a model
that already knows Streamlit would sail through docs that teach a newcomer
nothing.
The harness enforces the boundary rather than requesting it. All shell traffic
leaves through a proxy the harness owns, and documentation hosts are
unreachable from the shell, so an agent that tries
curl https://docs.streamlit.io/... is refused and the attempt is recorded.
The pages in a report are therefore the pages the agent really read, and when a
run fails, the last page it read is a fact rather than a guess.
Keys are read from QUICKSTARTED_* names first so they can sit in your shell
without other tooling on the same machine picking them up and spending them.
The vendor-standard names still work as a fallback, which is what CI sets.
The success script is the whole verdict
When the agent stops, the harness runs your check in the same workspace. Exit
code 0 is a pass. Nothing else is. The agent's own claim of success is recorded
and never trusted, and there is no LLM judge anywhere in scoring, because a
model asked to grade its own work will tell you the app is ready when app.py
does not parse.
Most checks are the size of the one above. You are asserting whatever your
tutorial already promises the reader: the file exists, the import works, the
command runs, the output contains the number on the page. If your quickstart
ends with "you should see 200", the check is grep -q 200.
Say what you saw. qs_fail is defined for every check, and a failure that
reports check failed: httpx is not a dependency in pyproject.toml is a bug
report, where exit code: 1 is a page to go and guess about.
Keep a long check in a file. success.file: checks/mine.sh puts it beside
the task, where shellcheck and syntax highlighting work on it. It is read when
the task loads and never written into the workspace, so the agent cannot read
its own success criteria.
Let the harness boot the server. If the goal is a running application, do not ask the agent to leave a process running and do not take its word for it:
success:
serve: .venv/bin/fastapi run app.py --host 127.0.0.1 --port $QS_PORT
wait_http:
path: /items/42
json:
item_id: 42
script: test -f app.py
The harness backgrounds it, polls, keeps the last error, dumps the server log when it gives up, and kills it. Every criterion is still yours: a task that serves and asserts nothing is a validation error, because a server that answers every request with a 500 also boots.
Iterate on a check for free
quickstarted run tasks/mine.yaml --agent claude --keep-sandbox
quickstarted check tasks/mine.yaml --sandbox /tmp/quickstarted-8ilw9l6v/workspace
check re-runs only the success script against the workspace the run left
behind, under the same backend. About a second per iteration, instead of a paid
run per iteration.
quickstarted validate catches the rest before you spend anything: a check that
requires a .venv nothing creates, a check that can fail in silence, an
entrypoint that 404s (--check-urls). Every one of those produces a number that
is wrong rather than low.
Replay mode, the free precondition
--agent replay runs your replay commands, the literal commands your docs
tell a reader to type, and stops at the first failure. No model, no API key, no
cost.
quickstarted run tasks/streamlit-quickstart.yaml --agent replay
Treat it as a floor to check on every push. If the documented commands are broken, no reader stands a chance and an agent run only adds noise on top of a failure you already know about. It cannot tell you whether a reader could have found those commands, understood their order, or guessed the prerequisite you left out. That is what agent mode is for.
Pass rates
The same docs and the same model can pass at 10:00 and fail at 10:05. One run
is a sample. Use --repeat and read the rate:
quickstarted run --repeat 5 --workers 3 --agent claude
With no paths, that runs every task in tasks/.
Every run is classified, and only two classifications say anything about your documentation:
| Classification | Meaning | Counts toward pass rate |
|---|---|---|
passed |
success script exited 0 | yes |
docs_gap |
the agent finished, the check failed | yes |
budget_exhausted |
out of turns, time, or tokens | no |
infra_error |
rate limit, upstream 5xx, network failure | no |
harness_error |
misconfigured task, our bug | no |
agent_refusal |
model declined | no |
Runs that produced no evidence are excluded from the numerator and the denominator, and reported separately. When nothing produced evidence the suite says so. Reporting 0% there would be a different claim, and a false one.
Does llms.txt actually help?
Whether a project ships llms.txt is a checklist item anyone can curl, and
scoring it would be the proxy metric this tool exists to replace. So
quickstarted never scores affordances. It measures them:
quickstarted run tasks/streamlit-quickstart.yaml --agent claude --repeat 10
quickstarted run tasks/streamlit-quickstart.yaml --agent claude --repeat 10 --affordances none
Same prompt, same task, same model. The only difference is whether the
agent could read llms.txt and .md variants. The difference in pass rate is
a measurement of the affordance itself.
--probe-affordances records what exists without changing anything, including
the size. Presence and usefulness come apart quickly: Streamlit's llms.txt is
67 KB, while its llms-full.txt is 1.8 MB, which is more than most agents can
read in one go.
Sandboxing
Agent-authored commands from somebody else's quickstart are untrusted code.
| Backend | Filesystem | Network |
|---|---|---|
docker |
container | internal network; the proxy sidecar is the only route out |
seatbelt (macOS) |
home directory unreadable, writes confined to the workspace | all egress denied except the harness proxy and loopback |
local |
none | none; proxy variables are advisory |
auto picks the best available. quickstarted run refuses local unless you
pass --allow-unenforced. Run quickstarted doctor to see what your machine
can enforce, which SDKs and keys it can find, and whether your tasks parse.
Details in SECURITY.md.
Install
uvx quickstarted run --example httpx --agent replay # nothing to install
pip install "quickstarted[claude]" # agent mode with Claude
pip install "quickstarted[all-agents]" # + OpenAI and Gemini
pip install quickstarted # replay mode only, no SDK
Python 3.9+. One runtime dependency: PyYAML. Every adapter is a plain tool-use loop against the same two harness-owned tools and the same shared prompt, so cross-model numbers compare like with like.
CI
Two jobs, on two schedules. Replay on every push, because it is free:
- run: pip install quickstarted
- run: quickstarted run --agent replay --junit junit.xml
Agent mode on a schedule, or when a model ships, because it costs real tokens:
on:
schedule:
- cron: "0 6 * * 1"
workflow_dispatch:
jobs:
agent:
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12"
- run: pip install "quickstarted[claude]"
- run: quickstarted run --agent claude --repeat 3
--out results --junit junit.xml
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
- uses: actions/upload-artifact@v4
if: always()
with:
name: quickstarted-results
path: results/
quickstarted run exits 1 when any task fails, so either job can gate
merges. results.json is versioned (schema 2.0). JUnit XML reports a docs gap
as a failure and infrastructure trouble as an error, so a rate limit does not
read as a broken quickstart.
Repeated flags belong in quickstarted.yaml instead:
run:
backend: docker
cache_dir: .cache
tasks:
setup:
- python3 -m venv .venv
It refuses agent, model, repeat and affordances. A file that quietly
changed which model served a task would make two runs incomparable for a reason
invisible in the command you typed.
Fetching other people's docs
Benchmarking means requesting pages from organisations who did not ask to be
measured. quickstarted sends a truthful User-Agent, honours robots.txt, and
waits a second between requests to the same host. --cache-dir makes reruns
read the same bytes. --refresh flags pages whose content changed since the
last run, which is a finding in its own right.
What this does and does not tell you
It tells you whether a given model, on a given day, could get from your docs to a working result, and where it was reading when it could not. That is a lower bound on documentation quality. Nothing here measures whether your docs are pleasant, complete, or accurate beyond the task you wrote, and there is no UI testing at all: if your product's button moved, this will not notice. Doc Detective is the tool for that, and the two answer different questions.
Write success scripts that assert local, deterministic facts. A pass means the floor holds, and says nothing about the ceiling.
License
MIT.
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file quickstarted-0.4.0.tar.gz.
File metadata
- Download URL: quickstarted-0.4.0.tar.gz
- Upload date:
- Size: 161.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1d8db4f62b9213bb8ae09d45edd68292604672962b9cf4e494a6139ea59ced28
|
|
| MD5 |
8291a88f7821f8e52df1447ecea4e341
|
|
| BLAKE2b-256 |
b77d9e3fe602146fce6564e30c12d6225885297721c2b6f8d7ba310f0ce351c6
|
Provenance
The following attestation bundles were made for quickstarted-0.4.0.tar.gz:
Publisher:
release.yml on snehankekre/quickstarted
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
quickstarted-0.4.0.tar.gz -
Subject digest:
1d8db4f62b9213bb8ae09d45edd68292604672962b9cf4e494a6139ea59ced28 - Sigstore transparency entry: 2310041624
- Sigstore integration time:
-
Permalink:
snehankekre/quickstarted@d080623fe164b9b146cac89576650863f55d54de -
Branch / Tag:
refs/tags/v0.4.0 - Owner: https://github.com/snehankekre
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@d080623fe164b9b146cac89576650863f55d54de -
Trigger Event:
release
-
Statement type:
File details
Details for the file quickstarted-0.4.0-py3-none-any.whl.
File metadata
- Download URL: quickstarted-0.4.0-py3-none-any.whl
- Upload date:
- Size: 95.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
fb258ed888d41ad7f2fc494c76b3e0157daad43625c374e5ae349013568795b3
|
|
| MD5 |
7940a84e9edc2fb89781d243fd84bbc3
|
|
| BLAKE2b-256 |
9f611291915f0ddb4f59aa819797f7996a9a6134c09928b4141ece67fc9f4309
|
Provenance
The following attestation bundles were made for quickstarted-0.4.0-py3-none-any.whl:
Publisher:
release.yml on snehankekre/quickstarted
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
quickstarted-0.4.0-py3-none-any.whl -
Subject digest:
fb258ed888d41ad7f2fc494c76b3e0157daad43625c374e5ae349013568795b3 - Sigstore transparency entry: 2310041668
- Sigstore integration time:
-
Permalink:
snehankekre/quickstarted@d080623fe164b9b146cac89576650863f55d54de -
Branch / Tag:
refs/tags/v0.4.0 - Owner: https://github.com/snehankekre
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@d080623fe164b9b146cac89576650863f55d54de -
Trigger Event:
release
-
Statement type: