Skip to main content

prettyplay

UI tests written as plain sentences. Each step sentence is turned into executable code once — by an LLM, against the live page — and cached in the repository. Every later run replays the cached code with no LLM involvement at all.

Installation

pip install prettyplay
playwright install            # browser binaries for the driver

Requires Python 3.10+.

Quick start

from prettyplay import PrettyTest


def test_login():
    t = PrettyTest("login-flow")
    t.action("открыть страницу логина")
    t.action("ввести логин и пароль")
    t.action("нажать «Войти»")
    t.assertion("появилась надпись «Добро пожаловать»")
    t.close()

Or with the context manager:

with PrettyTest("login-flow") as t:
    t.action("открыть страницу логина")

Each PrettyTest is fully self-contained: it owns its settings, its attempt budgets and its own browser session. close() (or leaving the with block) closes the page and stops the whole browser of that test, and every test starts with fresh attempt budgets. The old process-wide prettyplay.get_runtime() singleton was removed — build prettyplay.PrettyplayRuntime(config) directly if you composed objects over it.

The constructor arguments form the cache address: cache_key (mandatory) and cache_path (optional subdirectory). Equal keys in the shared root reuse one cached step across tests; a different language, step type or key is a different step.

The third argument overrides settings per test — only the fields you pass count: explicitly set values win over pyproject.toml and the environment, everything else resolves from the file layer as before:

from prettyplay import PrettyConfig

t = PrettyTest("login-flow", config=PrettyConfig(browser="firefox"))

Screenshots belong to the author — nothing is captured automatically. Both methods need a step to have run (the page opens lazily):

png = t.get_screenshot()                  # full-page PNG bytes
t.save_screenshot("artifacts/home.png")   # write full-page PNG to a file

What happens on a step

  • cache hit — the cached code runs; no LLM is contacted
  • cache miss — the step code is generated (a candidate that must actually work on the page), then cached; only successes are cached. A candidate assertion that legitimately fails stops the retries at once and is classified: a real defect fails as product_defect instead of burning the attempt budget
  • cached failure — the failure is classified:
    • rot (the UI changed) — the step is regenerated and the cache rewritten
    • product_defect — the test fails loudly; nothing is regenerated
    • incurable — the step fails with an explanation and a recommendation

Seeing the scenario

Step sentences go to the prettyplay logger at info level. The library configures no handlers — enable logging to see the scenario in the output:

import logging

logging.basicConfig(level=logging.INFO)          # plain unittest runs
# pytest: --log-cli-level=INFO

Configuration

The [tool.prettyplay] section of pyproject.toml:

[tool.prettyplay]
provider = "openai"              # openai | anthropic
browser = "chromium"             # chromium | firefox | webkit | chrome | msedge
model = "gpt-5"
generation_model = ""            # optional: empty -> model
classification_model = ""        # optional: empty -> model
base_url = ""
cache_root = ""                  # empty -> <repo>/.prettyplay/cache/
generation_prompt = ""           # user instructions for generation; empty -> no instructions block
browser_endpoint = ""            # ws:// endpoint of a remote browser; empty -> local launch
generation_attempts = 3
healing_attempts = 2
send_screenshots = false
headless = true                  # false -> run with a visible browser window

chrome and msedge launch the locally installed browser through the chromium engine; the browser must be installed on the machine.

A non-empty generation_prompt is sent verbatim as a USER INSTRUCTIONS block with every generation and regeneration request — it steers the style of the generated code (e.g. prefer data-test-id attributes), never the failure classification. The instructions are not part of the cache address: changing them never invalidates cached steps — a cached step runs unchanged.

A non-empty browser_endpoint (e.g. ws://ci-grid:3000/playwright/chromium) connects to a remote Playwright Server or browser grid instead of launching locally: headless does not apply to a connect (window visibility belongs to the endpoint server) and chrome/msedge map to the chromium engine — channels are a local-launch concept. A non-empty endpoint must be a valid ws/wss URL (otherwise ConfigurationError names the setting), and a failed connect fails loudly with the endpoint in the message.

Every setting has a PRETTYPLAY_<SETTING_UPPER> environment override for CI, except the browser (PRETTYPLAY_BROWSER_NAME) and headless (PRETTYPLAY_BROWSER_HEADLESS). The removed legacy name PRETTYPLAY_BROWSER fails immediately with a hint to use PRETTYPLAY_BROWSER_NAME.

An invalid setting fails loudly with a ConfigurationError: one line per setting — the name, the received value and the allowed values.

LLM API keys are never stored in the config file: they come only from the environment — OPENAI_API_KEY for openai, ANTHROPIC_API_KEY for anthropic — and are read lazily on the first request.

Failure taxonomy

Every library failure derives from PrettyplayError:

Exception Meaning Recommended reaction
ProductDefectError real product regression treat as a bug: this failure is the value of the suite
IncurableStepError the step cannot be (re)generated follow recommendation: reword the step or refresh the cache
LlmUnavailableError LLM infrastructure down restore provider access; cached steps are unaffected
ConfigurationError invalid [tool.prettyplay] settings fix the named setting — the message lists received and allowed values

ProductDefectError and IncurableStepError carry a verdict — category, explanation, recommendation from the failure classification — appended to the exception message and delivered through the on_step_verdict hook. ProductDefectError also derives from AssertionError, so any runner counts it as a failed test, never an error. Tracebacks of library failures are folded at the t.action(...) / t.assertion(...) call site: internal engine frames never appear in what the runner shows.

A healed run never turns a ProductDefectError into a green test.

Hooks

from prettyplay.reporting import StepHooks


class Reporter(StepHooks):
    def on_step_started(self, step_text: str, step_type: str) -> None: ...
    def on_step_passed(self, step_text: str, step_type: str) -> None: ...
    def on_step_failed(self, step_text: str, step_type: str, error: str) -> None: ...
    def on_step_verdict(self, step_text: str, category: str, explanation: str, recommendation: str) -> None: ...
    def on_generation_started(self, step_text: str, attempt: int) -> None: ...
    def on_healing_started(self, step_text: str, category: str) -> None: ...
    def on_healed(self, step_text: str, explanation: str) -> None: ...
    def on_cache_saved(self, step_text: str, filename: str) -> None: ...
    def on_cache_skipped(self, step_text: str, reason: str) -> None: ...


t = PrettyTest("login-flow")
t.add_hooks(Reporter())

A raising hook never fails the run; the failure is logged. on_step_verdict fires after on_step_failed, only when the terminal failure carried a verdict.

Cache and CI workflow

The cache lives under .prettyplay/cache/ as plain Python files — one per step, carrying its metadata (step sentence, cache key, step type, creation date) and the step code.

Generate locally where the LLM is reachable → commit the cache directory → CI runs the whole suite from the cache with no LLM keys at all.

Limitations

Step sentences land in the repository cache, the logs and the LLM requests: never put secrets or personal data into a step.

Metadata

Release files for prettyplay 0.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for prettyplay 0.0.0
File Size Uploaded
prettyplay-0.0.0.tar.gz 389.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for prettyplay 0.0.0
File Interpreter ABI Platform
prettyplay-0.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 482.8 kB

Release files / prettyplay-0.0.0.tar.gz

Download URL prettyplay-0.0.0.tar.gz
Size 389.0 kB
Tags Source
SHA-256 checksum
How to use checksums
c1fc1df7dd877f117d9f022906154d2a146915e12f39d7a48787c6be2674ada6
BLAKE2b-256 checksum
How to use checksums
4ed64a68a727d1dca271e810fc413432c266d141c5e5085cc9bb18d2ab62662b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / prettyplay-0.0.0-py3-none-any.whl

Download URL prettyplay-0.0.0-py3-none-any.whl
Size 93.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
bdf3bb788c52828db5fb7eed68d1b4270f8f0e609bb571a1d2954cb41642af66
BLAKE2b-256 checksum
How to use checksums
db54db2e6934e818c589db5d6091aebc7b276bcef142a597baa93653c3ebdc62
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

0.0.2

2 release files

0.0.1

2 release files

This release

0.0.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page