Skip to main content

Predictive scheduling for autonomous LLM agents: suspend gracefully before rate limits, resume without losing work.

Project description

agentpause

Predictive scheduling for autonomous LLM agents. Suspend an agent gracefully before it hits a provider rate limit, and resume it without redoing work.

The core works on any provider — cloud (OpenAI, Anthropic, Groq) or local — because it only serializes application-level state. True KV-cache warm start is an optional plugin for self-hosted runtimes (llama.cpp, vLLM).

Measured results (from the accompanying research)

Experiment Result
Simulation, 300 runs/config reactive baseline crashes up to 100% → predictive 0% (k=4), token waste −80%
Real provider (Groq free, 6K TPM) multi-window task completed with zero 429 errors
vs LangGraph's MemorySaver (200 runs) LangGraph: 7.8 × 429/run, 9,239 tokens wasted; predictive: 0 and 0
End-to-end A/B (thinking models, 4B/8B) recovery 54×–93× faster; total task time −19%

Reproduce them yourself with a free key: python scripts/benchmark_groq.py.

Status: early alpha (0.1). Core components, the high-level PredictiveScheduler API, the LiteLLM adapter (any provider), the LangGraph adapter, runnable examples, and a full test suite are in place; the optional KV-cache plugin is next.

Quick example

from agentpause import PredictiveScheduler

sched = PredictiveScheduler(backend=my_llm_call, telemetry=my_rate_limit_reader)

with sched.session("task-1") as s:      # resumes automatically if interrupted before
    for question in questions[s.step:]:  # skip steps finished before a suspend
        s.add_user(question)
        if s.should_suspend():           # predictive check, *before* the call
            s.checkpoint()
            break                        # stop cleanly; rerun to resume
        reply = s.call()
    else:
        s.complete()                     # task done: drop the checkpoint

backend is any callable messages -> (reply, tokens_used); telemetry is any callable () -> remaining_tokens. See examples/quickstart.py for a runnable demo (no API keys needed) that suspends mid-task and resumes on the next run.

Real providers via LiteLLM

The LiteLLM adapter supplies both callables for 100+ providers (OpenAI, Groq, Anthropic, local servers, ...), reading the budget from each response's rate-limit headers and refreshing stale readings with a tiny telemetry ping:

from agentpause import PredictiveScheduler
from agentpause.adapters.litellm import LiteLLMAdapter

adapter = LiteLLMAdapter(model="groq/llama-3.1-8b-instant")
sched = PredictiveScheduler(backend=adapter.backend, telemetry=adapter.telemetry)

Install with pip install -e ".[litellm]"; see examples/litellm_groq.py. To validate against your own provider (frontier models included): python scripts/validate_provider.py gpt-4o-mini.

Rate-limit headers by provider (defaults target the OpenAI-style names):

Provider remaining tokens remaining requests reset
OpenAI x-ratelimit-remaining-tokens x-ratelimit-remaining-requests x-ratelimit-reset-tokens (1s, 6m0s)
Groq x-ratelimit-remaining-tokens x-ratelimit-remaining-requests x-ratelimit-reset-tokens (7.66s, 2m59.56s)
Anthropic anthropic-ratelimit-tokens-remaining anthropic-ratelimit-requests-remaining RFC 3339 timestamp — set remaining_header= etc.

Header names differ? Override them: LiteLLMAdapter(model=..., remaining_header="...", requests_header="...", reset_header="...").

Beyond tokens: RPM and wait-vs-suspend

Telemetry can be richer than a token count. adapter.budget reports all three dimensions providers expose — remaining tokens (TPM), remaining requests (RPM), and seconds until the window resets — and unlocks the three-valued decision:

sched = PredictiveScheduler(backend=adapter.backend, telemetry=adapter.budget)

d = session.next_action()   # "continue" | "wait" | "checkpoint"

wait fires when the budget does not fit but the window resets within wait_threshold_s (default 15 s): a short in-place pause is cheaper than a full suspend/resume cycle. Exhausted requests (RPM = 0) block the call even with plenty of tokens left. Plain-int telemetry keeps working unchanged.

Async

Every entry point has an async twin — same rules, same guarantees, never blocks the event loop:

adapter = LiteLLMAdapter(model="groq/llama-3.1-8b-instant")
sched = PredictiveScheduler(backend=None, async_backend=adapter.abackend,
                            telemetry=adapter.telemetry)
reply = await session.acall()          # retry/backoff via asyncio.sleep
await guard.acheck(state["messages"])  # async LangGraph nodes

For sharper input estimates, wire the per-model tokenizer: PredictiveScheduler(..., count_tokens=adapter.count_tokens) (falls back to the ~4 chars/token heuristic if the tokenizer is unavailable).

When prediction fails anyway

Estimates are statistical — a 429 can still slip through. agentpause survives it instead of crashing:

  • typed errors: everything derives from AgentPauseError (RateLimitHit, TelemetryError, CheckpointError, BackendError);
  • retry with backoff: unexpected 429s are retried (provider retry-after honored, else exponential backoff — see RetryPolicy);
  • hits are feedback: each one bumps safety_k up (capped at k_max), so the scheduler grows more cautious on workloads it underestimates;
  • clean failure: a failed call leaves the session state untouched — no phantom steps, resumes stay consistent.

LangGraph integration

AgentPauseGuard adds the predictive gate to any LangGraph node — LangGraph persists reactively (after nodes), the guard pauses before the call that would hit the rate limit, via LangGraph's own interrupt() + checkpointer:

from agentpause.adapters.langgraph import AgentPauseGuard

guard = AgentPauseGuard(telemetry=adapter.telemetry)   # e.g. from LiteLLMAdapter

def agent_node(state):
    guard.check(state["messages"])       # pauses the graph here if needed
    reply = llm.invoke(state["messages"])
    guard.record(state["messages"], reply.usage_metadata["total_tokens"])
    ...

Resume the paused thread with graph.invoke(Command(resume=True), config). On resume the guard re-reads telemetry fresh — never from the checkpoint. Install with pip install -e ".[langgraph]"; see examples/langgraph_quickstart.py.

Why

Current agent frameworks persist state reactively: they checkpoint after a step completes and crash when the provider returns HTTP 429. agentpause adds a predictive layer that estimates the next step's cost, compares it against the remaining rate-limit budget, and suspends cleanly before the error — then resumes from the exact step.

This library is the engineering counterpart of the research preprint "A Resource-Aware Predictive Scheduler for Autonomous LLM Agents".

Components (v0.1)

Module Role
PredictiveScheduler / Session the high-level API: session(), should_suspend(), call(), checkpoint()
Estimator predicts next-step token cost with a moving-average error correction (ε) and tracks σ
should_checkpoint / RiskModel the suspension rule (remaining < estimated + k·σ) and a diagnostic risk score
Budget / decide three-dimensional telemetry (TPM, RPM, reset time) and the three-valued rule: continue / wait / checkpoint
StateStore / Checkpoint atomic logical checkpointing with idempotency keys — works on any provider
adapters.litellm.LiteLLMAdapter backend + telemetry for any LiteLLM-supported provider (headers → budget, stale reading → 1-token ping)
adapters.langgraph.AgentPauseGuard predictive gate for LangGraph nodes: check() interrupts the graph before the fatal call, record() trains the estimator

Install (from source, during development)

git clone https://github.com/<user>/agentpause
cd agentpause
pip install -e ".[dev]"
pytest

Roadmap

  • Core components + test suite
  • PredictiveScheduler high-level API (session(), should_suspend(), call())
  • Runnable quickstart example (no keys)
  • LiteLLM adapter (works with any provider)
  • LangGraph adapter (interrupt + checkpointer)
  • Optional KV-cache plugin for llama.cpp / vLLM (true warm start)

License

MIT

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

agentpause-0.2.0.tar.gz (42.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

agentpause-0.2.0-py3-none-any.whl (29.5 kB view details)

Uploaded Python 3

File details

Details for the file agentpause-0.2.0.tar.gz.

File metadata

  • Download URL: agentpause-0.2.0.tar.gz
  • Upload date:
  • Size: 42.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.7

File hashes

Hashes for agentpause-0.2.0.tar.gz
Algorithm Hash digest
SHA256 f4c7e1948c2b63fc4f084a619f9e7b38bcf8e771d82352c1c77d1be013b61a3b
MD5 99e33396b05c20d41429875567d3f3bc
BLAKE2b-256 44af4e500ef311b1eefac9cac85998ccdf3f1945e127f6ef29cf2524f76f1973

See more details on using hashes here.

File details

Details for the file agentpause-0.2.0-py3-none-any.whl.

File metadata

  • Download URL: agentpause-0.2.0-py3-none-any.whl
  • Upload date:
  • Size: 29.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.7

File hashes

Hashes for agentpause-0.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 83a70297e3eb9e5deefa85b39686333650068a33ca7c34398c5c0da2f87e9b4d
MD5 51ff731012965b099ee12d9a832e7620
BLAKE2b-256 d879f026c7e3848470cbcc30a4d813ce6aa1e047fc95227160ed4b641fdc8c44

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page