retryhop
English | 简体中文
Retries, client-side rate limiting and a circuit breaker for LLM and HTTP API calls. No runtime dependencies. Works on sync and async functions.
Errors are classified by attribute (status_code, status,
response.headers) and by class name, so retryhop does not import any SDK.
The tests use real exception objects from openai, anthropic, httpx and
requests. aiohttp errors go through the same checks but are not covered by
tests yet.
Status: alpha. The API may change before 1.0.
pip install retryhop
Python 3.9+.
Usage
from openai import OpenAI
from retryhop import llm_retry
client = OpenAI(max_retries=0)
@llm_retry() # defaults: 6 attempts, 300 s deadline
def ask(prompt: str) -> str:
r = client.chat.completions.create(
model=MODEL,
messages=[{"role": "user", "content": prompt}],
)
return r.choices[0].message.content
Async functions work the same way:
from anthropic import AsyncAnthropic
from retryhop import llm_retry
client = AsyncAnthropic(max_retries=0)
@llm_retry(attempts=8, deadline=600)
async def ask(prompt: str) -> str:
msg = await client.messages.create(
model=MODEL, max_tokens=1024,
messages=[{"role": "user", "content": prompt}],
)
return msg.content[0].text
Create the SDK client with max_retries=0. The openai and anthropic clients
retry twice by default, so otherwise each retryhop attempt can be up to three
HTTP requests.
When retryhop gives up it re-raises the last exception, so existing handlers
such as except openai.RateLimitError: still work. Pass reraise=False to
get a RetryError instead.
What is retried
is_transient(exc) makes the decision. It checks, in order:
- The
x-should-retryresponse header, if the server sent one. - The HTTP status. 408, 409, 429, 500, 502, 503, 504 and 529 are retried. Any other status is raised immediately.
- If there is no status: connection errors and timeouts are retried. This
covers the built-in
ConnectionErrorandTimeoutErrorand the matching classes in openai, anthropic, httpx, requests/urllib3 and aiohttp.
Anything else is raised immediately.
409 is on the list because the openai and anthropic SDKs retry it (they treat
it as a lock timeout). 529 is Anthropic's "overloaded". If a 409 from your API
means a real conflict, use retry(...) with your own retry_on.
Wait time
- If the error response has
retry-after-msorRetry-After(seconds or an HTTP date), retryhop waits that long plus 0-0.25 s of jitter, capped atmax_retry_after. - Otherwise it uses exponential backoff with full jitter: a random wait in [0, 1 s], then [0, 2 s], [0, 4 s], and so on, with the upper bound capped at 60 s.
deadlineis the total budget across attempts. If the server asks for a wait longer than what is left, retryhop gives up without sleeping. A backoff wait is cut short to fit.
The deadline is checked between attempts. It does not cancel a request that is already running, so also set a request timeout on the client.
llm_retry options
| Option | Default | Description |
|---|---|---|
attempts |
6 |
Total calls, including the first. |
deadline |
300 |
Seconds across all attempts. None disables it. |
max_retry_after |
120 |
Upper bound, in seconds, on a server-requested wait. |
backoff |
exponential, see above | Used when there is no retry-after header. Any f(attempt) -> seconds. |
circuit |
None |
A shared CircuitBreaker. |
on_retry |
None |
Called with a RetryState before each wait. |
reraise |
True |
False raises RetryError instead of the last exception. |
sleep |
None |
Replacement sleep function, for tests. |
Logging
import logging
from retryhop import llm_retry, describe
log = logging.getLogger("llm")
@llm_retry(on_retry=lambda s: log.warning(describe(s)))
def ask(prompt): ...
Output looks like:
attempt 1 failed: RateLimitError (HTTP 429); waiting 1.10s (server Retry-After), elapsed 0.0s
attempt 2 failed: InternalServerError (HTTP 503); waiting 1.64s (backoff), elapsed 1.1s
Rate limiting
RateLimiter keeps one token bucket for requests per minute and one for
tokens per minute. A call that would go over either limit sleeps until there
is room.
from retryhop import RateLimiter, llm_retry
limiter = RateLimiter(requests_per_minute=500, tokens_per_minute=200_000)
def estimate(prompt, **_):
return len(prompt) // 4 + 1024 # rough input estimate + max output
@llm_retry()
@limiter.limit(tokens=estimate)
def ask(prompt): ...
Put @limiter.limit below @llm_retry so retries are limited too.
The token count is your own estimate, and the provider counts tokens its own
way. If the response reports actual usage, call
limiter.consume(actual - estimated). A negative value gives tokens back.
Without the decorator: limiter.acquire(tokens=n) or
await limiter.acquire_async(tokens=n).
Things to know:
- The limiter is thread-safe and can be used from asyncio tasks, but its state is per process. If several processes or hosts share one API key, split the limits between them.
- The buckets start full, so a full minute's quota can go out at once right after startup.
- A single call that needs more tokens than
tokens_per_minuteraisesValueError.
Circuit breaker
from retryhop import CircuitBreaker, CircuitOpenError, llm_retry
breaker = CircuitBreaker(failure_threshold=5, recovery_time=30)
@llm_retry(circuit=breaker)
def ask(prompt): ...
try:
ask("hi")
except CircuitOpenError as e:
... # e.g. fall back to another model; e.retry_in = seconds until the next trial
- After
failure_thresholdconsecutive failed attempts the circuit opens. Calls then raiseCircuitOpenErrorwithout hitting the API. - After
recovery_timeseconds one trial call is let through. Success closes the circuit; failure opens it again. Other calls made during the trial getCircuitOpenError. - Failures are counted per attempt, not per call. With the defaults
(6 attempts, threshold 5) a single call that keeps getting 503 opens the
circuit by itself, and that call ends with
CircuitOpenErrorrather than the SDK exception. - A non-retryable error such as a 400 counts as a success, because the API did respond. It resets the failure count.
CircuitOpenErroris never retried.- State is per process.
Lower-level API
-
retry(...): the decoratorllm_retryis built on. It takesattempts,exceptions,retry_on,retry_if_result,wait_hint,backoff,deadline,on_retry,reraise,circuitandsleep. Its defaults are different: 3 attempts, anyExceptionis retried, no deadline, andRetryErroris raised on give-up.Polling a batch job:
@retry(attempts=60, retry_if_result=lambda job: job.status != "completed", backoff=constant(30)) def wait_for_batch(job_id): return client.batches.retrieve(job_id)
-
retry_call(func, *args, retry_options={...}, **kwargs): same asretrywithout decorating. -
is_transient(exc),status_code_of(exc),retry_after_of(exc): the checks described above. -
describe(state): formats aRetryStatefor logging. -
constant,linear,exponential: backoff functions.
Development
pip install -e ".[sdk-test]" # ".[test]" skips the tests that need the SDKs
pytest
python -m build && twine check dist/*
License
MIT
Metadata
Release files for retryhop 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| retryhop-0.1.0.tar.gz | 21.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| retryhop-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 36.7 kB
Release files / retryhop-0.1.0.tar.gz
| Download URL | retryhop-0.1.0.tar.gz |
|---|---|
| Size | 21.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
4264162a21e46ec513019b3e418aa4bbb1351d0427cc7677cde8d2a07d64a73f
|
|
BLAKE2b-256 checksum How to use checksums |
0bd7c0ecde0923ea4688e42f1442f8181dc8df17f11a2319eb5c2ab24d09590a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.
Transparency logRelease files / retryhop-0.1.0-py3-none-any.whl
| Download URL | retryhop-0.1.0-py3-none-any.whl |
|---|---|
| Size | 15.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
45082d2fa0ee241eef5a8443ad1b2897d509c9888aa37b6494fe94d831da3620
|
|
BLAKE2b-256 checksum How to use checksums |
c6627f2e0e2fbf5c0e7c6b36bc29ca58acbc2b60b94de22cf596b2492f05a64b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.
Transparency log