twinbox
Check that a rewritten HTTP service behaves exactly like the original.
twinbox is a Python library for black-box parity testing of HTTP services. It is meant for teams that rewrite a service — for example, from Python to Go or Rust — and need to show that the new implementation behaves the same way from the outside as the old one.
You describe the expected behaviour once, as a corpus of YAML scenarios (request → expected response). twinbox starts each implementation in a Docker container on a fresh, isolated environment, runs the same scenarios against it and compares the live responses with the expectations. Differences you accept on purpose are written into the scenario together with the reason, so they stay visible instead of being silently ignored.
Status: early development. There is no release yet, and the API and the scenario format may still change.
What works today
- Corpus check without containers —
twinbox checkvalidates every scenario (structure, tags, masks, substitutions, accepted differences, coverage matrix) and reports each problem with the field path and an error code. - Run against one implementation —
twinbox run <name>starts a PostgreSQL database, a mock of external services and the implementation itself, runs the scenarios and writes a JSON result document and, optionally, an HTML report. - pytest integration — every scenario is an ordinary pytest test, so
-k,-mand parallel runs withpytest-xdistwork as usual. - Load comparison —
twinbox loadruns locust against each implementation in turn, with the same CPU and memory limits, and writes one Markdown report: latency and RPS by handle, CPU and memory of the container, CPU cost of single calls, and a verdict by handle from the project's own criterion.
Planned after the first release: code coverage of the reference implementation, through integration plugins.
How it works
For every scenario twinbox builds a clean environment:
- a fresh copy of a PostgreSQL database with the schema applied;
- a mock of the external services the implementation calls — it echoes requests by default and can be programmed per scenario;
- the implementation under test, started from its Docker image.
Then it sends the scenario's requests, compares each response with the expectation and tears the environment down. Nothing leaks from one scenario into the next.
Requirements
- Python 3.13+
- uv
- Docker; for
twinbox loadalso cgroup v2 on the host andcatin the implementation's image (see docs/usage.md)
Installation
Until the first release, install from the repository:
uv add git+https://gitlab.com/twinbox-group/twinbox.git
Quick start
1. Enable twinbox in your project's pyproject.toml. Every field has a
default; set only what differs:
[tool.twinbox]
plugin = "project_plugin" # module that describes your implementations
scenarios_dir = "scenarios" # where scenario.yaml files live
2. Describe your implementations in that module — one Profile per
implementation, exactly one of them marked as the reference:
from twinbox.harness.profile import Endpoint, Profile, ReadinessProbe
_API = (Endpoint(name="api", port=8080, default=True),)
_READY = ReadinessProbe(path="/health", expected_status=200)
profiles = [
Profile(
name="legacy",
is_reference=True,
default_image="shop/legacy:dev",
named_endpoints=_API,
readiness_probe=_READY,
),
Profile(
name="rewrite",
default_image="shop/rewrite:dev",
named_endpoints=_API,
readiness_probe=_READY,
),
]
3. Write a scenario, e.g. scenarios/users/create_and_fetch/scenario.yaml:
description: create a user and fetch it by the returned id
tags: [] # tags must be declared in the plugin's `labels`; none are used here
steps:
- name: create user
request:
method: POST
path: /users
json: {name: Ann}
expect:
status: 201
json_exact:
id: "{{any_uuid}}" # mask: any UUID matches
created_at: "{{any_iso8601}}" # mask: any ISO 8601 timestamp matches
name: Ann
capture:
- {name: user_id, path: id}
- name: fetch user by captured id
request:
method: GET
path: "/users/{{capture.user_id}}" # value captured in the previous step
expect:
status: 200
json_subset: {id: "{{capture.user_id}}"}
4. Check and run:
uv run twinbox check # validate the scenarios
uv run twinbox run legacy --image shop/legacy:dev --html reports/legacy.html
uv run twinbox run rewrite --image shop/rewrite:dev --html reports/rewrite.html
Run twinbox help or twinbox help <command> for the full reference of
every command and option.
Results are written to artifacts/results/ (JSON) and, when --html <path>
is passed, to the specified file path (HTML).
Load comparison
twinbox load answers the question "is the rewrite more expensive under
load?". For each implementation, one after another, it starts a fresh stand
with the CPU and memory limits from the profile, calls the project's
load_dataset to create data, waits for the CPU to calm down and warms the
implementation up. Then it measures: a locust ladder with the read locust
file, the same with the write locust file, and the CPU cost of single calls
to chosen handles. The ladder adds users step by step until the container
reaches its CPU limit, requests start to fail, or the steps run out; in the
last case the report says limit not reached in <max_steps> steps. The
peak number of database connections is a column of the report and does not
stop the ladder: the connection pool size is set inside the service, and
twinbox does not know it. A service that hits its pool shows it in this
column together with a growing p95.
1. Give every profile CPU and memory limits:
from twinbox.harness.profile import ResourceLimits
_LIMITS = ResourceLimits(cpu_nanocores=1_000_000_000, memory_bytes=512 * 1024 * 1024)
# in every Profile(...): resource_limits=_LIMITS
2. Point twinbox at the locust files in pyproject.toml. Every other
field of [tool.twinbox.load] has a default:
[tool.twinbox.load]
read_locustfile = "load/read.py"
write_locustfile = "load/write.py" # optional
step_duration = "20s" # default: 30s
max_steps = 10 # default: 20
Settings of locust itself go to [tool.locust] or after --. Users, spawn
rate, run time, host and report files are set by twinbox on every step, so
-u, -r, -t, -H, -f, --csv, --html and --headless are rejected
there.
3. Write the locust file. twinbox.load.binding() returns the data that
load_dataset created on the current stand, twinbox.load.endpoints() — the
addresses of the implementation's endpoints:
from locust import HttpUser, task
from twinbox.load import binding
class ReadUser(HttpUser):
def on_start(self) -> None:
self.user_id = binding()["user_id"]
@task
def get_user(self) -> None:
self.client.get(f"/users/{self.user_id}", name="/users/[id]")
4. Add two functions to the plugin module: load_dataset creates the
data, acceptance_criterion decides what is acceptable. The numbers come as
dataclasses from twinbox.load:
from collections.abc import Mapping, Sequence
import httpx
from twinbox.load import HandleVerdict, ImplementationNumbers, NetworkNumbers
def load_dataset(endpoints: Mapping[str, str], profile: str) -> Mapping[str, str]:
response = httpx.post(f"{endpoints['api']}/users", json={"name": profile}, timeout=10)
response.raise_for_status()
return {"user_id": response.json()["id"]}
def acceptance_criterion(numbers: Sequence[ImplementationNumbers]) -> Sequence[HandleVerdict]:
reference = next(item for item in numbers if item.is_reference)
if not isinstance(reference.read, NetworkNumbers):
return [] # a failed measurement already makes the exit code 2
return [
HandleVerdict(
handle=handle,
implementation=item.implementation,
accepted=own.p95_ms <= reference.read.handles[handle].p95_ms * 1.2,
note=f"p95 {own.p95_ms:.1f} ms",
)
for item in numbers
if not item.is_reference and isinstance(item.read, NetworkNumbers)
for handle, own in item.read.handles.items()
if handle in reference.read.handles
]
5. Run:
uv run twinbox load # every implementation, then the verdict
uv run twinbox load rewrite # one implementation, numbers only
uv run twinbox load -- --only-summary # everything after -- goes to locust
The report is written to artifacts/load/<time>/summary.md, next to
locust's HTML and CSV files of every step; its path is printed at the end.
Exit code 0 means every measurement ran and no verdict is rejected, 1
means at least one verdict is rejected, 2 means a check before the start or
a measurement failed, the criterion failed, or the report was not written.
Errors are printed as [<code>] <message>. While the run goes, lines like
[load] rewrite read: step 2: starting, 20 users for 30s on stderr show which
implementation, kind and step is running now.
The CPU and memory limits apply to twinbox run too. The full reference —
every setting and its default, the isolated handles, the forms of the
numbers, the report and the errors — is in
docs/usage.md.
Examples
A runnable example lives in its own repository, gitlab.com/twinbox-group/example: one calculator service in two implementations, Python and Rust, and nine lessons that introduce twinbox's capabilities one at a time, starting from a plain check.
Documentation
- docs/usage.md — settings, extension points, the scenario format, commands and their results, the load comparison.
- CONTRIBUTING.md — development setup, checks, the Docker test suite and supported platforms.
License
MIT.
Metadata
Release files for twinbox 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| twinbox-0.1.0.tar.gz | 459.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| twinbox-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 629.6 kB
Release files / twinbox-0.1.0.tar.gz
| Download URL | twinbox-0.1.0.tar.gz |
|---|---|
| Size | 459.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
c1da2d1ad60856703104aa73b5e036abbc092892f8b83cf342a2a373b2454fc9
|
|
BLAKE2b-256 checksum How to use checksums |
de680ec48dc11bf16366eacc830674b2a6a78aa476b4f626413a297c533ae5ea
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.11.26 {"installer":{"name":"uv","version":"0.11.26","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"13","id":"trixie","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|
Release files / twinbox-0.1.0-py3-none-any.whl
| Download URL | twinbox-0.1.0-py3-none-any.whl |
|---|---|
| Size | 169.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
bdeb64d2216c277f3682d3edb4a9b15d7dd9629bb7df3efdec044b215dd99177
|
|
BLAKE2b-256 checksum How to use checksums |
6c14fd2e31d52b4c223d483d5020b2154649fd63583053f3f6da1298d436ec5c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.11.26 {"installer":{"name":"uv","version":"0.11.26","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"13","id":"trixie","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
|