Rate Limiting Algorithms
This project adheres to Semantic Versioning
Algorithms
| Algorithms | Sync | Async |
|---|---|---|
| Leaky Bucket | Yes | Yes |
| Token Bucket | Yes | Yes |
| Generic Cell Rate Algorithm | Yes | Yes |
| LLM-Token | Yes | Yes |
[!NOTE]
Implementations will be single-threaded, blocking requests (or the equivalent) with burst capabilities. With asyncio, we use non-blocking cooperative multitasking, not preemptive multi-threading
Development
Setup uv-based virtual environment
# Install uv
# for a mac or linux
brew install uv
# OPTIONAL: or
curl -LsSf https://astral.sh/uv/install.sh | sh
# python version are automatically downloaded as needed or: uv python install 3.12
uv venv rate --python 3.12
# to activate the virtual environment
source .venv/bin/activate
# to deactivate the virtual environment
deactivate
Create lock file + requirements.txt
# after pyproject.toml is created
uv lock
uv export -o requirements.txt --quiet
Upgrade dependencies
# can use sync or lock
uv sync --upgrade
or
# to upgrade a specific package
uv lock --upgrade-package requests
Usage
[!IMPORTANT] These are special use cases. The general use cases are in the
examples/folder
Token & Variable-Amount Rate Limiting (LLM TPM, Batch Sizes)
limitor natively supports variable-capacity and LLM token-based rate limiting via decorators and context managers using acquire_ctx(), reconcile(), and optional token_estimate / token_reconcile decorator hooks.
1. Dynamic Token Estimation with Decorators (LLM Prompt Tokens)
import random
import time
from limitor import rate_limit
from limitor.token_bucket.core import SyncTokenBucket
# Rate limit of 100,000 tokens per second
@rate_limit(
capacity=100_000,
seconds=1,
bucket_cls=SyncTokenBucket,
token_estimate=lambda prompt, **kw: len(prompt) * 4, # Estimate tokens from prompt
)
def generate_response(prompt: str):
print(f"[{time.strftime('%X')}] Generating response for prompt ({len(prompt)*4} tokens)")
return f"Response to: {prompt}"
for i in range(10):
sample_prompt = "x" * random.randint(1_000, 5_000)
generate_response(sample_prompt)
2. Full LLM Prompt Estimation + Response Token Reconciliation
When working with LLM APIs, you can estimate prompt tokens before the request and reconcile with the exact total tokens reported in the API response:
from limitor import rate_limit
from limitor.token_bucket.core import SyncTokenBucket
# Automatically debit excess tokens or refund unused tokens
@rate_limit(
capacity=50_000,
seconds=60,
bucket_cls=SyncTokenBucket,
token_estimate=lambda prompt, **kw: len(prompt) // 4,
token_reconcile=lambda response: response["usage"]["total_tokens"],
)
def query_llm(prompt: str):
# Simulated API call returning usage metadata
return {"content": "Hello!", "usage": {"total_tokens": 120}}
3. Variable Tokens with Context Managers (acquire_ctx & reconcile)
from limitor.configs import BucketConfig
from limitor.token_bucket.core import SyncTokenBucket
bucket = SyncTokenBucket(BucketConfig(capacity=100_000, seconds=60))
estimated_tokens = 500
# Acquire custom amount via context manager and reconcile actual usage
with bucket.acquire_ctx(amount=estimated_tokens):
# response = client.chat.completions.create(...)
actual_tokens = 450 # e.g. response.usage.total_tokens
bucket.reconcile(actual=actual_tokens, estimated=estimated_tokens)
With User-Specific Rate Limits + Cache
from functools import wraps
import time
from typing import Optional
from cachetools import LRUCache, TTLCache
from limitor.base import SyncRateLimit
from limitor.configs import BucketConfig
from limitor.leaky_bucket.core import (
AsyncLeakyBucket,
SyncLeakyBucket,
)
def _get_user_cache(max_users, ttl):
if ttl is not None:
return TTLCache(maxsize=max_users, ttl=ttl)
return LRUCache(maxsize=max_users)
def rate_limit_per_user(capacity=10, seconds=1, max_users=1000, ttl=None, bucket_cls: type[SyncRateLimit] = SyncLeakyBucket):
buckets = _get_user_cache(max_users, ttl)
global_bucket = bucket_cls(BucketConfig(capacity=capacity, seconds=seconds))
def decorator(func):
# optional use_id. if not set, it will default to a regular global rate limiter
# if user_id is not set, this means the max_users / ttl parameters will be ignored
@wraps(func)
def wrapper(*args, user_id=None, **kwargs):
if user_id is None:
bucket = global_bucket
else:
if user_id not in buckets:
buckets[user_id] = bucket_cls(BucketConfig(capacity=capacity, seconds=seconds))
bucket = buckets[user_id]
with bucket:
return func(user_id, *args, **kwargs)
return wrapper
return decorator
@rate_limit_per_user(capacity=2, seconds=1, max_users=3, ttl=600) # TTLCache: 10 min/user
def something_user(user_id):
print(f"User {user_id} called at {time.strftime('%X')}")
for _ in range(20):
try:
x = 1 if _ % 2 == 0 else 0
something_user(user_id=x)
except Exception as error:
print(f"Rate limit exceeded: {error}")
Release files for limitor 0.6.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| limitor-0.6.0.tar.gz | 109.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| limitor-0.6.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 133.1 kB
Release files / limitor-0.6.0.tar.gz
| Download URL | limitor-0.6.0.tar.gz |
|---|---|
| Size | 109.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
0800b9e7af889e7fcaec0af785afe3be9f640f9b29adad7263ff5cabe2f6edc1
|
|
BLAKE2b-256 checksum How to use checksums |
975b4cd9ba31d46aed94f0dfa2f11c7fc45e49f1f3463fba6ea105da0911d4f1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.8.4
|
Release files / limitor-0.6.0-py3-none-any.whl
| Download URL | limitor-0.6.0-py3-none-any.whl |
|---|---|
| Size | 23.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
4f628db36626efe0fd8b08fff71ce1e3d1079ce09e0342c340074a862e2775e3
|
|
BLAKE2b-256 checksum How to use checksums |
d13bc1f8c0f117a35dd8008f51ca6692d1c713517218d5d02d5ec133adbbd409
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
uv/0.8.4
|