abjt-limitter
Rate limiting for Python. No dependencies, no frameworks.
pip install abjt-limitter
from limitra import TokenBucket
limiter = TokenBucket(rate=10, capacity=100)
result = limiter.allow()
if result.allowed:
print("ok")
else:
print(f"slow down, retry in {result.retry_after:.1f}s")
Picking an algorithm
Five of them, all behind the same interface:
| class | behaviour | pick it when |
|---|---|---|
TokenBucket |
refills steadily, lets you spend the whole bucket at once | you want to tolerate bursts |
LeakyBucket |
the same policy as TokenBucket, expressed as a queue filling and draining |
you'd rather think in "how full is it" than "how much credit is left" |
FixedWindow |
counter resets every N seconds | you need a hard "100 per minute" cap and nothing subtler |
SlidingWindow |
weights the previous window as it decays | you want fixed-window cost without the boundary burst |
SlidingLog |
keeps a timestamp per request | you need exact limits and can afford O(limit) memory per key |
Two things worth knowing before you choose:
LeakyBucket and TokenBucket admit the same requests. As admission
control the water-level leaky bucket is the token bucket's dual —
capacity - water is the token count — so for the same rate and
capacity they accept and reject identically. The smoothing a leaky bucket
is famous for belongs to the queueing variant, which delays requests rather
than refusing them. If you want output genuinely paced, set capacity=1 so
no burst is possible, or use wait() below.
FixedWindow can admit 2 * limit across a boundary. Its counter
clears all at once, so limit requests at the end of one window and limit
at the start of the next land within one window-length of each other. Use
SlidingWindow or SlidingLog if that matters.
Swapping algorithms only changes the constructor — the windows take limit
and window, the buckets take rate and capacity, and everything after
that is identical:
from limitra import SlidingWindow
limiter = SlidingWindow(limit=100, window=60)
result = limiter.allow()
The result
Every allow() call returns the same object:
| field | type | what it is |
|---|---|---|
allowed |
bool | whether the request went through |
remaining |
int | units left after this decision |
limit |
int | the max |
reset_after |
float | seconds until back at full capacity, 0.0 if already there |
retry_after |
float | seconds to wait if denied, 0.0 if allowed |
It's frozen and hashable, and if result: means if result.allowed:.
retry_after is exact: wait that long and the next attempt gets through.
It isn't rounded up to the next window, so a client that obeys it isn't
throttled to less than the rate you configured, and a denial never
advertises a zero-second wait.
For HTTP, hand it straight to your framework:
result = limiter.allow()
response.headers.update(result.as_headers())
That sets RateLimit-Limit, RateLimit-Remaining and RateLimit-Reset,
the older X-RateLimit-* spellings, and Retry-After when the request was
denied.
The two spellings are read differently, so each carries the unit its
readers expect: RateLimit-Reset and Retry-After are seconds from
now, while X-RateLimit-Reset is a Unix timestamp, which is how
GitHub-style clients parse that name. Times round up, so a client that
obeys them never comes back early.
The rest of the API
limiter.allow(cost=5) # spend more than one unit
limiter.peek() # what would allow() say? doesn't spend anything
limiter.remaining() # units available right now
limiter.reset_after() # seconds until back to full
limiter.reset() # wipe it back to a fresh limiter
limiter.refund(cost=5) # give back capacity you didn't end up using
limiter.limit # the largest cost this limiter could ever admit
refund() is what makes stacked limits work. Almost every real API has
two — a burst tier and a sustained one — and without a refund the first
tier is charged for requests the second one rejected:
if burst.allow().allowed:
if sustained.allow().allowed:
serve()
else:
burst.refund() # never served it, don't charge for it
Limiters copy cleanly too: copy.copy() gives you an independent limiter
with its own lock, rather than a second handle on the same one.
peek() returns exactly what the next allow() would, including the
remaining you'd be left with.
cost must be between 1 and limit. Asking for more than limit raises
ValueError rather than denying you, because no amount of waiting would
make that request succeed.
Limiters also report the settings they were built with — rate and
capacity on the buckets, limit and window on the windows.
Waiting instead of refusing
If you're the client rather than the server — pacing yourself against
someone else's API — wait() blocks until there's room:
for item in queue:
limiter.wait() # sleeps until there's capacity
send(item)
if not limiter.wait(timeout=5).allowed:
print("gave up after 5s")
It sleeps for the limiter's own retry_after and never holds the lock
while sleeping, so other threads keep working. In asyncio use
wait_async(), which is identical but yields to the event loop:
async def handler(request):
await limiter.wait_async()
return await upstream.fetch(request)
allow() and peek() never block, so they're safe to call from a
coroutine directly.
Multiple keys
Use RateLimitManager for per-user or per-IP limits:
from limitra import RateLimitManager, TokenBucket
manager = RateLimitManager(TokenBucket, rate=10, capacity=50)
manager.allow("user_123")
manager.allow("user_456") # completely separate bucket
Keys are usually attacker-controlled, so the map needs a bound. Sweeping idle keys is the safer option, because it only drops keys that have gone quiet:
manager.cleanup(max_idle=300) # drop keys idle for 5 minutes, returns the count
max_keys caps the number of entries instead:
manager = RateLimitManager(TokenBucket, max_keys=100_000, rate=10, capacity=50)
Be aware of the trade: when the map is full, adding a key evicts one, and
an evicted key loses its state and comes back with a full budget. Eviction
prefers keys already at full capacity, which have nothing to lose, but a
large enough flood of fresh keys will eventually drop a throttled one.
Size max_keys well above your expected number of active keys, and prefer
cleanup() when you can.
Only allow(), wait() and wait_async() create a key. peek(),
remaining(), reset_after() and get() answer for an unknown key
without starting to track it, so a metrics scrape can't fill the map. The
manager also has reset, remove, clear, keys, len() and in, and
its constructor arguments are checked when you build it rather than on the
first request.
Threads and processes
Every limiter is thread-safe — a threading.Lock guards each one, and
timing uses time.monotonic(), so changing the system clock can't move a
window. The suite runs on free-threaded Python (3.14t) as well, where the
GIL isn't there to paper over a missing lock.
State lives in memory, in one process. Run four gunicorn workers and you get four independent limiters, so the effective limit is four times what you configured. If you need one limit across processes or machines, you need a shared backend; this library doesn't provide one.
Development
pip install -e ".[dev]"
pytest # the suite
pytest -m benchmark -s # throughput and scaling, opt-in
pytest --doctest-modules src/limitra # the examples in the docstrings
ruff check . && ruff format --check .
mypy src/limitra
Ships with type hints and a py.typed marker. Requires Python 3.10+.
Changes are in CHANGELOG.md.
License
MIT
Metadata
Release files for abjt-limitter 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| abjt_limitter-0.2.0.tar.gz | 54.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| abjt_limitter-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 84.7 kB
Release files / abjt_limitter-0.2.0.tar.gz
| Download URL | abjt_limitter-0.2.0.tar.gz |
|---|---|
| Size | 54.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
913cf110b55a3a4da701d9ecc92bdc2d5dec700269cf4a28895c8cbb24c92efd
|
|
BLAKE2b-256 checksum How to use checksums |
ece47089ab37a06b62fc2f9b9a5064d7ab261c05d7802de6e913d5d0a7a1383a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 17, 2026.
Transparency logRelease files / abjt_limitter-0.2.0-py3-none-any.whl
| Download URL | abjt_limitter-0.2.0-py3-none-any.whl |
|---|---|
| Size | 30.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
b63234a2104762023c338ac4c2747b888760dfefae9ba2d90a74c3ee78c40e63
|
|
BLAKE2b-256 checksum How to use checksums |
76d41ff6f0c06dd88e5f3f345a6b160e178d90ca346f80020322b888abd7bf2a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 17, 2026.
Transparency log