Skip to main content

Rate Limiting Algorithms

PyPI Version Build Status Documentation Status Code Coverage PyPI - Python Version

This project adheres to Semantic Versioning

Algorithms

Algorithms Sync Async
Leaky Bucket Yes Yes
Token Bucket Yes Yes
Generic Cell Rate Algorithm Yes Yes
LLM-Token Yes Yes

[!NOTE]
Implementations will be single-threaded, blocking requests (or the equivalent) with burst capabilities. With asyncio, we use non-blocking cooperative multitasking, not preemptive multi-threading

Development

Setup uv-based virtual environment

# Install uv
# for a mac or linux
brew install uv
# OPTIONAL: or
curl -LsSf https://astral.sh/uv/install.sh | sh

# python version are automatically downloaded as needed or: uv python install 3.12
uv venv rate --python 3.12


# to activate the virtual environment
source .venv/bin/activate

# to deactivate the virtual environment
deactivate

Create lock file + requirements.txt

# after pyproject.toml is created
uv lock

uv export -o requirements.txt --quiet

Upgrade dependencies

# can use sync or lock
uv sync --upgrade

or 

# to upgrade a specific package
uv lock --upgrade-package requests

Usage

[!IMPORTANT] These are special use cases. The general use cases are in the examples/ folder

Token & Variable-Amount Rate Limiting (LLM TPM, Batch Sizes)

limitor natively supports variable-capacity and LLM token-based rate limiting via decorators and context managers using acquire_ctx(), reconcile(), and optional token_estimate / token_reconcile decorator hooks.

1. Dynamic Token Estimation with Decorators (LLM Prompt Tokens)

import random
import time
from limitor import rate_limit
from limitor.token_bucket.core import SyncTokenBucket

# Rate limit of 100,000 tokens per second
@rate_limit(
    capacity=100_000,
    seconds=1,
    bucket_cls=SyncTokenBucket,
    token_estimate=lambda prompt, **kw: len(prompt) * 4,  # Estimate tokens from prompt
)
def generate_response(prompt: str):
    print(f"[{time.strftime('%X')}] Generating response for prompt ({len(prompt)*4} tokens)")
    return f"Response to: {prompt}"

for i in range(10):
    sample_prompt = "x" * random.randint(1_000, 5_000)
    generate_response(sample_prompt)

2. Full LLM Prompt Estimation + Response Token Reconciliation

When working with LLM APIs, you can estimate prompt tokens before the request and reconcile with the exact total tokens reported in the API response:

from limitor import rate_limit
from limitor.token_bucket.core import SyncTokenBucket

# Automatically debit excess tokens or refund unused tokens
@rate_limit(
    capacity=50_000,
    seconds=60,
    bucket_cls=SyncTokenBucket,
    token_estimate=lambda prompt, **kw: len(prompt) // 4,
    token_reconcile=lambda response: response["usage"]["total_tokens"],
)
def query_llm(prompt: str):
    # Simulated API call returning usage metadata
    return {"content": "Hello!", "usage": {"total_tokens": 120}}

3. Variable Tokens with Context Managers (acquire_ctx & reconcile)

from limitor.configs import BucketConfig
from limitor.token_bucket.core import SyncTokenBucket

bucket = SyncTokenBucket(BucketConfig(capacity=100_000, seconds=60))

estimated_tokens = 500

# Acquire custom amount via context manager and reconcile actual usage
with bucket.acquire_ctx(amount=estimated_tokens):
    # response = client.chat.completions.create(...)
    actual_tokens = 450  # e.g. response.usage.total_tokens
    bucket.reconcile(actual=actual_tokens, estimated=estimated_tokens)

With User-Specific Rate Limits + Cache

from functools import wraps
import time
from typing import Optional

from cachetools import LRUCache, TTLCache

from limitor.base import SyncRateLimit
from limitor.configs import BucketConfig
from limitor.leaky_bucket.core import (
    AsyncLeakyBucket,
    SyncLeakyBucket,
)


def _get_user_cache(max_users, ttl):
    if ttl is not None:
        return TTLCache(maxsize=max_users, ttl=ttl)
    return LRUCache(maxsize=max_users)

def rate_limit_per_user(capacity=10, seconds=1, max_users=1000, ttl=None, bucket_cls: type[SyncRateLimit] = SyncLeakyBucket):
    buckets = _get_user_cache(max_users, ttl)
    global_bucket = bucket_cls(BucketConfig(capacity=capacity, seconds=seconds))

    def decorator(func):
        # optional use_id. if not set, it will default to a regular global rate limiter
        # if user_id is not set, this means the max_users / ttl parameters will be ignored
        @wraps(func)
        def wrapper(*args, user_id=None, **kwargs):
            if user_id is None:
                bucket = global_bucket
            else:
                if user_id not in buckets:
                    buckets[user_id] = bucket_cls(BucketConfig(capacity=capacity, seconds=seconds))
                bucket = buckets[user_id]
            with bucket:
                return func(user_id, *args, **kwargs)

        return wrapper

    return decorator

@rate_limit_per_user(capacity=2, seconds=1, max_users=3, ttl=600)  # TTLCache: 10 min/user
def something_user(user_id):
    print(f"User {user_id} called at {time.strftime('%X')}")

for _ in range(20):
    try:
        x = 1 if _ % 2 == 0 else 0
        something_user(user_id=x)
    except Exception as error:
        print(f"Rate limit exceeded: {error}")

Release files for limitor 0.6.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for limitor 0.6.0
File Size Uploaded
limitor-0.6.0.tar.gz 109.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for limitor 0.6.0
File Interpreter ABI Platform
limitor-0.6.0-py3-none-any.whl Python 3 none any Details

Total release size: 133.1 kB

Release files / limitor-0.6.0.tar.gz

Download URL limitor-0.6.0.tar.gz
Size 109.3 kB
Tags Source
SHA-256 checksum
How to use checksums
0800b9e7af889e7fcaec0af785afe3be9f640f9b29adad7263ff5cabe2f6edc1
BLAKE2b-256 checksum
How to use checksums
975b4cd9ba31d46aed94f0dfa2f11c7fc45e49f1f3463fba6ea105da0911d4f1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.8.4

Release files / limitor-0.6.0-py3-none-any.whl

Download URL limitor-0.6.0-py3-none-any.whl
Size 23.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
4f628db36626efe0fd8b08fff71ce1e3d1079ce09e0342c340074a862e2775e3
BLAKE2b-256 checksum
How to use checksums
d13bc1f8c0f117a35dd8008f51ca6692d1c713517218d5d02d5ec133adbbd409
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.8.4

Release history Release notifications | RSS feed

This release

0.6.0 This release

2 release files

0.5.0

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.3

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page