Skip to main content

tapio

CI Coverage Ruff Typed: mypy strict License: Apache 2.0

A Pekko-inspired actor toolkit for Python. Typed, asyncio-native actors with supervision, and Pydantic models throughout.

Status: pre-alpha. The runtime core runs: actor systems, spawning, typed tell, bounded mailboxes, dead letters, a deadline-based shutdown, supervision with backoff, death watch, ask, timers, stash, message adapters, run_blocking and a round-robin pool router. Two systems can also talk over a TCP link, with a handshake, optional TLS and messages that carry refs across, plus watching and asking over the link, a failure detector, quarantine, an explicit reconnect, and starting an actor on another node. Nodes can form a cluster and agree on who is in it, watch one another for silence, and, given a downing strategy, resolve a partition by downing the losing side rather than blocking on it for ever. Applications react to membership through cluster events delivered to an actor's mailbox, and can run a singleton on the oldest member of a role or route work over a group of them.

Why this exists. tapio is a testbed for AI agentic development. The point of the project is to find out what coding agents can carry on their own over a real codebase: a non-trivial design, invariants that span modules, and a gate that has to stay green. Nearly all of the code, tests and documentation here is written by agents under human review.

The library itself is meant to work, and the design decisions are argued rather than generated. Treat it as an experiment that happens to be working software, not as something to depend on in production yet.

Named for the Finnish god of the forest, because supervision hierarchies are trees. It keeps the mythological lineage of Akka (Sámi) and Apache Pekko (Finnish) without borrowing anyone's trademark. See the note at the bottom.

What it is

tapio gives you the concurrency structure of Apache Pekko (itself the ASF fork of Akka) inside a single Python process:

  • Actors: isolated state, one mailbox, no locks
  • Supervision: restart with backoff, escalate, stop. Failure policy is a first-class thing rather than scattered try/except
  • Death watch: learn when a child dies, without polling
  • Ask, timers, stash, routers: the patterns you would otherwise hand-roll
  • Remoting: two systems on a TCP link, sending each other typed messages
  • Membership: nodes gossip a view of the cluster that they all converge on, with a leader that is computed rather than elected
  • Downing: nodes watch one another for silence, and, with a strategy configured, resolve a partition by downing the losing side (keep the majority, a static quorum, the oldest, down everything, or hand an even split to an outside lease), the losing side downing itself
  • Cluster surface: react to membership as events on an actor's mailbox, run a singleton on the oldest member of a role, or spread work over a group router of them

It is a library, not infrastructure. Pip-install it into the service you already have. Nodes find each other from a seed list you deploy with them, so there is no broker, no coordinator and no separate cluster process to run.

What it is not

This is inspired by Pekko, not a port, and shares no code with it. Deliberately out of scope, permanently:

  • Sharding and distributed data. Placing an entity on one node, moving it when that node goes away, or replicating state so that two nodes can write it, is a much larger problem than agreeing on who is in the cluster. If you need either, use Ray or put a broker between your nodes.
  • Competing with the JVM on throughput. See below.

Clustering as a whole used to be on that list, and the reason given was that reimplementing gossip and split-brain resolution in Python is a multi-year project with a high bug-severity floor. Membership is here because the merge two nodes run to agree is a pure function with three laws behind it, so it can be tested as one, and it is. Downing, deciding what to do about a member that has stopped answering, is the part with the high bug-severity floor, so it is off unless you configure it: a node watches its peers, and a partition is resolved by a strategy you choose, with the losing side downing itself and, if you ask, shutting its own system down. What stays out of scope is agreement that needs consensus, which is why an even split is handed to a lease you hold elsewhere rather than to a Paxos or Raft this library ships.

Remoting does give you location transparency in the narrow sense: an actor holding a ref just sends, and does not need to know which node the target is on. What it does not change is the failure model. A network is in the middle, delivery over a link is at-most-once, and a message that crossed one is equal to what was sent rather than the same object.

Where it fits

The target is I/O-bound orchestration: many independent, long-lived, stateful things, each mostly waiting on something external, each able to fail on its own.

Good fits:

  • A session actor per user, holding conversation state and calling an LLM API
  • Saga orchestration: payment, then inventory, then shipping, compensating on failure
  • One actor per websocket, with behavior-switching as the protocol state machine
  • Rate limiting and circuit breaking, where the mailbox is the mutex

Bad fits, use something else:

  • High-volume per-record stream processing → Bytewax, Quix Streams
  • Anything CPU-heavy inside a handler. One blocking call stalls every actor sharing the event loop.

Actors are not microservices. A microservice is a unit of deployment. An actor is a unit of concurrency. You will have tens of thousands of actors inside one service, each an asyncio.Task, so the ceiling is memory rather than a cluster: an idle actor costs about 15 KB, so a hundred thousand of them fit in 1.6 GB. tapio structures the inside of that service. It competes with asyncio.Queue + TaskGroup + a dict + hand-rolled retries, not with your API framework.

A first actor

import asyncio

from tapio import ActorSystem, Behavior, Behaviors, Message
from tapio.actor import ActorContext, ActorRef


class Greeted(Message):
    whom: str


class Greet(Message):
    whom: str
    reply_to: ActorRef[Greeted]


async def on_greet(ctx: ActorContext[Greet], message: Greet) -> Behavior[Greet]:
    ctx.log.info("hello, %s!", message.whom)
    message.reply_to.tell(Greeted(whom=message.whom))
    return Behaviors.same()


async def on_greeted(message: Greeted) -> Behavior[Greeted]:
    print(f"{message.whom} has been greeted")
    return Behaviors.same()


async def main() -> None:
    async with ActorSystem("hello") as system:
        listener = system.spawn(Behaviors.receive_message(on_greeted), name="listener")
        greeter = system.spawn(Behaviors.receive(on_greet), name="greeter")
        greeter.tell(Greet(whom="world", reply_to=listener))
        await asyncio.sleep(0.1)


asyncio.run(main())

Runnable versions of this and every other example live in examples/:

uv run python -m tapio_examples.hello_world

Design notes

Sending never blocks. ref.tell(msg) is synchronous and fire-and-forget, so it works from sync callbacks and signal handlers too. Backpressure belongs to the mailbox rather than the send call. Bounded mailboxes take an overflow strategy (fail, drop_new, drop_oldest). drop_new and drop_oldest publish what they discard as a dead letter; fail discards nothing and raises MailboxFullError in the sender instead, so the sender decides. (A fail send from another thread has nobody to raise into, so it dead-letters.) await ref.offer(msg) waits for capacity when you want to be throttled.

Every message is a validated Pydantic model. Messages subclass tapio.Message, which is frozen and re-validated on delivery rather than only at construction. What that costs depends on the model: about 30% more per message for a one-field message, and about 3x for a ten-field one with nested models. In an actor that spends 50 ms on an HTTP call it is invisible. In a tight per-record loop it dominates, which is why that workload is out of scope above. The check sits behind a single validate_on_tell setting. The numbers are below.

Undeliverable messages go to dead letters. An ActorRef stays valid after its actor dies, so tell never raises for a dead target. The message is published as a DeadLetter you can subscribe to and assert on. To know when something died, watch it with ctx.watch rather than asking whether it is alive. A liveness answer is out of date as soon as you have it.

Requires Python 3.11+ (for asyncio.timeout() and typing.Self).

Numbers

Measured with make bench, make bench-scale and make bench-cluster, which anybody can run. Each prints what it ran on, and this is what they printed:

Apple M1 Pro, 8 cores, 17 GB RAM
Darwin 25.5.0
CPython 3.13.2, pydantic 2.13.4
measured 2026-08-20

The figures below are from twenty rounds, on a laptop that was doing other things at the time. Take the ratios seriously and the absolute numbers as an order of magnitude: a server core running nothing else does better on the absolute figures, so the ratios are the durable part.

Messages per second, one sender to one actor, ten thousand at a time:

Message validate_on_tell=True validate_on_tell=False What validation costs
one int field 450,000/s 637,000/s 1.4x
ten fields, two of them nested models 261,000/s 618,000/s 2.4x

That is the design bet of this library, priced. Revalidation on delivery is not free and it is not 10x either: it is a property of your message, and the setting is there for the case where you have measured it and it matters.

Everything else, locally:

Starting an actor 24 us
ask round trip 77 us

Across a link, two systems over loopback, so the network itself is nearly free and what is left is the codec and the socket:

Local Remote Ratio
Messages per second 450,000/s 24,000/s 19x slower
ask round trip 77 us 410 us 5x slower

JSON on the wire is the cost there, and it is the reason the codec sits behind one module with the frame format versioned: a binary codec is an additive change rather than a rewrite.

Resident actors, idle but able to answer, each measured in its own process:

Actors RSS Per actor ask p50 ask p99
1,000 60 MB 15.2 KB 54 us 164 us
10,000 197 MB 14.9 KB 48 us 132 us
100,000 1,583 MB 15.0 KB 50 us 216 us

Memory per actor is flat, and the median round trip to one actor among a hundred thousand is the same as to one among a thousand: a mailbox nobody is sending to costs nothing to have. The p99 column is the honest part of this table. It does not grow with the number of actors, it wanders, because what puts a tail on a round trip here is the garbage collector walking a large heap rather than anything in the runtime.

A cluster, at scale, measured with make bench-cluster on the same machine:

Nodes Converges in Gossip frame Per node
5 ~10 rounds 1.0 KB 1.6 KB/s
20 ~19 rounds 3.2 KB 4.0 KB/s
50 ~22 rounds 9.1 KB 9.0 KB/s

A gossip frame carries a node's whole view, so it grows with the cluster, and the table shows the growth is linear: about 180 bytes a member, whether there are five or fifty. Per-node bandwidth is a few kilobytes a second, one gossip frame and a handful of heartbeats each. That is the scale claim, as a number: adding a node costs every other node a couple of hundred bytes a second, not a new connection to everyone.

Convergence is counted in gossip rounds rather than seconds, and it varies from run to run because a node gossips to a peer it picks at random. Seconds are only rounds times the one-second gossip interval anyway, and the fifty-node figure is the pessimistic end: all the nodes share one event loop here, which no real deployment does, so at that size the loop is the bottleneck rather than anything in the runtime.

Development

Managed with uv. The Makefile is the entry point and the single source of truth for what CI runs.

make            # list every target
make install    # create the venv, install deps
make check      # pre-push gate: lint + types + tests
make ci         # exactly what GitHub Actions runs

Trademark note

Apache Pekko and Apache Kafka are trademarks of the Apache Software Foundation. This project is not affiliated with, endorsed by, or derived from the Apache Pekko codebase. It is an independent implementation inspired by its design, and references to Pekko above are descriptive only.

License

Apache-2.0. Copyright 2026 Carmelo Polito.

Release files for tapio-py 0.14.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tapio-py 0.14.1
File Size Uploaded
tapio_py-0.14.1.tar.gz 579.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for tapio-py 0.14.1
File Interpreter ABI Platform
tapio_py-0.14.1-py3-none-any.whl Python 3 none any Details

Total release size: 859.7 kB

Release files / tapio_py-0.14.1.tar.gz

Download URL tapio_py-0.14.1.tar.gz
Size 579.9 kB
Tags Source
SHA-256 checksum
How to use checksums
22f509605316c22464fc9ddbfae9f9ff2bfbc2106b72e7e03525dff13ffeac0f
BLAKE2b-256 checksum
How to use checksums
08999c078242e3acebdce3b5938ddda63a1182ac1c4abafec45d163ff6c71660
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.19 {"installer":{"name":"uv","version":"0.12.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / tapio_py-0.14.1-py3-none-any.whl

Download URL tapio_py-0.14.1-py3-none-any.whl
Size 279.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
23294eb7d966d8844cab3d7600dfeeb41ef5bdf66650f0a8eb2a03b20ece170a
BLAKE2b-256 checksum
How to use checksums
fb121dd21a28d0a83d85d8bda0389f1bfbf96a18878ac7944c1b0500dbdb06a3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.19 {"installer":{"name":"uv","version":"0.12.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

0.14.2

2 release files

This release

0.14.1 This release

2 release files

0.14.0

2 release files

0.13.9

2 release files

0.13.8

2 release files

0.13.7

2 release files

0.13.6

2 release files

0.13.5

2 release files

0.13.4

2 release files

0.13.3

2 release files

0.13.2

2 release files

0.13.1

2 release files

0.13.0

2 release files

0.12.2

2 release files

0.12.1

2 release files

0.12.0

2 release files

0.11.0

2 release files

0.10.0

2 release files

0.9.0

2 release files

0.8.2

2 release files

0.8.1

2 release files

0.8.0

2 release files

0.7.0

2 release files

0.6.2

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page