Skip to main content

APST Starter Kit

APST stands for Accelerated Prompt Stress Testing. It is a depth-oriented LLM safety and reliability evaluation workflow that repeatedly samples the same prompts under controlled conditions to estimate empirical failure probability under repeated inference.

This v0.1 starter kit is local-first and conference-friendly. You can run the mock demo without an API key, then swap in OpenAI, Together.ai, or an OpenAI-compatible local endpoint when you are ready to test a real model.

Quickstart

Install from PyPI after the first release:

pip install apst-starter-kit
apst init my-apst-demo
cd my-apst-demo

apst run --config configs/demo_mock.yaml
apst report --results outputs/demo_results.csv --lang both

Or run from a Git checkout:

git clone <repo-url>
cd apst-starter-kit
pip install -e .

apst run --config configs/demo_mock.yaml
apst report --results outputs/demo_results.csv --lang both

The demo writes:

  • outputs/demo_results.csv
  • outputs/demo_results.json
  • outputs/demo_results_report_both.md

What The Demo Does

The mock run:

  • loads a small prompt set from data/prompts/demo_prompts.json
  • repeatedly samples each prompt at two temperatures
  • judges every response with the local rule judge mode
  • computes APST reliability and repeated-use risk metrics
  • exports CSV and JSON result files
  • generates English, Chinese, or bilingual Markdown reports

No API key or network access is required for configs/demo_mock.yaml.

Configuring Models

Use the mock provider for local demos:

models:
  - name: mock-apst-model
    model_id: mock-apst-model
    provider: mock
judge_mode: rule

Use OpenAI with an LLM judge:

export OPENAI_API_KEY=...
apst run --config configs/openai_example.yaml

Use an OpenAI-compatible local endpoint, such as Ollama or vLLM exposing /v1/chat/completions. APST uses the same client path for both: model, base_url, and a placeholder api_key. With a local base_url, prompts, outputs, and labels are sent only to the local server you configured.

models:
  - model: llama3.1
    base_url: http://localhost:11434/v1
    api_key: local-not-needed
judge_mode: rule

The OpenAI-compatible client is installed with the base package, so Ollama and vLLM only need the local server running:

apst run --config configs/ollama_local.yaml
apst run --config configs/vllm_local.yaml

For local LLM-as-judge, point judge_model at another OpenAI-compatible local server:

judge_mode: llm
judge_model:
  model: llama3.1
  base_url: http://localhost:11434/v1
  api_key: local-not-needed

See docs/local_model_servers.md for Ollama and vLLM server commands.

Together.ai models are also supported through the existing provider extra:

pip install -e ".[providers]"
export TOGETHER_API_KEY=...
apst run --config configs/openai_example.yaml --models meta-llama/Llama-3.3-70B-Instruct-Turbo

Judge Modes

  • rule: local deterministic checks for refusal, harmful operational guidance, crisis-support handling, and gibberish. Good for demos, local LLMs, and fast smoke tests.
  • heuristic: local malformed-output check only. Good when you want a very conservative no-network sanity pass.
  • llm: LLM-as-judge classification using judge_model. Good for richer audits and real model comparisons.

APST Metrics

Each prompt/model/temperature config reports:

  • empirical_failure_probability: observed failures divided by repeated samples
  • reliability: 1 - empirical_failure_probability
  • apst_risk_at_10: estimated chance of at least one failure across 10 independent attempts
  • failure_probability_ci_low and failure_probability_ci_high: Wilson interval bounds
  • failure_mode_distribution: counts by judge label

The repeated-use estimate is intentionally simple:

Risk@N = 1 - (1 - empirical_failure_probability)^N

Reports

apst report --results outputs/demo_results.csv --lang en
apst report --results outputs/demo_results.csv --lang zh
apst report --results outputs/demo_results.csv --lang both

Use --audit-contact to put a real intake link, email address, or conference note in the report. See docs/enterprise_audits.md for the enterprise audit handoff text.

Legacy Commands

The older llm-eval entry point and APST/AIRBench commands are still available:

llm-eval run-apst --config configs/smoke.yaml
llm-eval freeze-airbench --region us --per-l4 5 --output data/prompts/airbench_us_v1.json

Local Checks

pytest

Publishing

See docs/publishing.md for GitHub release and PyPI publishing steps. The package includes an apst init command so PyPI users can scaffold the runnable configs, prompt files, and docs without cloning the repository.

Release files for apst-starter-kit 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for apst-starter-kit 0.1.0
File Size Uploaded
apst_starter_kit-0.1.0.tar.gz 47.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for apst-starter-kit 0.1.0
File Interpreter ABI Platform
apst_starter_kit-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 100.2 kB

Release files / apst_starter_kit-0.1.0.tar.gz

Download URL apst_starter_kit-0.1.0.tar.gz
Size 47.8 kB
Tags Source
SHA-256 checksum
How to use checksums
804e1123622829df92bf4a7c45e92399ce3f37af6095e4b6d4d72ddd7a1a6e6d
BLAKE2b-256 checksum
How to use checksums
6f25eb13af02ba24b5a99f2f43d06dbeb0ef17cd8c49a4d55a0ef9872e12ba16
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.11.11

Release files / apst_starter_kit-0.1.0-py3-none-any.whl

Download URL apst_starter_kit-0.1.0-py3-none-any.whl
Size 52.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
5a92241362c7b090f70bc73c1ae7719f80fea708b6957afa5b0a7a916c641ddf
BLAKE2b-256 checksum
How to use checksums
1ab6adf855149aef383fd96c33493f32c4f581c40801d0b4dd4702d8e0fa977a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.11.11

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page