This project has been archived by its maintainers, and is no longer receiving any updates.
autourgos-summarizer
Scratchpad compression middleware for Autourgos agents.
Automatically summarizes long reasoning chains so your agent never runs out of token space — even on tasks with dozens of iterations.
Why use this?
LLM agents accumulate a scratchpad — a growing log of thoughts, tool calls, and observations. On long tasks this scratchpad can:
- Hit the LLM's context window limit and crash
- Slow down responses (more tokens = more cost and latency)
- Confuse the LLM with too much irrelevant history
AutoSummarizeMiddleware solves this by compressing the scratchpad in the background every N iterations (or when it exceeds a size limit), keeping only what matters: key findings, tool results, and current status.
Install
pip install autourgos-summarizer
Depends on autourgos-agent. Works with any Autourgos agent.
Quick Start
from autourgos_summarizer import AutoSummarizeMiddleware
from autourgos_agent import Agent
summarizer = AutoSummarizeMiddleware(
summarize_every=5, # compress every 5 iterations
max_scratchpad_chars=15000, # also compress if scratchpad exceeds 15k chars
)
agent = Agent(llm=my_llm, middleware=[summarizer])
result = agent.invoke("Research the latest breakthroughs in quantum computing")
print(result)
When used with a verbose Agent (verbose=True), AutoSummarizeMiddleware also narrates its own actions into the agent's trace, right alongside the Thought/Action/Observation lines, e.g.:
[Summarizer] Compressed scratchpad (iteration 5, was 18342 chars).
How it works
AutoSummarizeMiddleware hooks into on_iteration_start. Before each iteration it checks:
- Is this a multiple of
summarize_every? (e.g. iteration 5, 10, 15…) - Is the scratchpad longer than
max_scratchpad_chars?
If either is true, it calls the LLM (middleware's own llm if set, otherwise agent.llm) with a compression prompt that asks it to distill the scratchpad into:
[Summary of steps 1-N]
Key findings: ...
Tool results: ...
Current status: ...
This summary replaces the full scratchpad before the next LLM call. The agent continues from where it left off, but with a much smaller context.
Summarization runs synchronously (the agent loop is paused at this point), and a threading.Lock prevents duplicate summarizations when parallel tool callbacks fire.
Use a dedicated LLM for summarization
You can pass a separate llm to AutoSummarizeMiddleware. This LLM is used only for compressing the scratchpad — great for keeping costs low by using a cheaper/faster model just for this job.
from autourgos_summarizer import AutoSummarizeMiddleware
from autourgos_openaichat import OpenAIChatModel
# Main agent uses a powerful model
main_llm = OpenAIChatModel(model="gpt-4o")
# Summarizer uses a cheap fast model
cheap_llm = OpenAIChatModel(model="gpt-4o-mini")
summarizer = AutoSummarizeMiddleware(
summarize_every=5,
llm=cheap_llm, # overrides agent.llm for summarization
)
agent = Agent(llm=main_llm, middleware=[summarizer])
result = agent.invoke("Research the latest breakthroughs in quantum computing")
print(result)
If llm is not provided, it falls back to agent.llm automatically.
Parameters
| Parameter | Type | Default | Description |
|---|---|---|---|
summarize_every |
int | None |
5 |
Summarize every N iterations. None = disable iteration-based trigger. |
max_scratchpad_chars |
int |
15000 |
Also trigger if scratchpad exceeds this many characters. |
llm |
any | None |
LLM for summarization. Needs .invoke(prompt). Falls back to agent.llm if not set. |
Combine with other middleware
from autourgos_summarizer import AutoSummarizeMiddleware
from autourgos_history import AgentHistoryMiddleware
history = AgentHistoryMiddleware()
summarizer = AutoSummarizeMiddleware(summarize_every=5)
agent = Agent(llm=my_llm, middleware=[summarizer, history])
Requirements
- Python 3.9+
- Any Autourgos agent that exposes
agent.scratchpad,agent.llm, andagent.current_query(requiresautourgos-agent>=2.0.2)
License
MIT — see LICENSE
Metadata
Release files for autourgos-summarizer 3.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| autourgos_summarizer-3.1.0.tar.gz | 8.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| autourgos_summarizer-3.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 16.0 kB
Release files / autourgos_summarizer-3.1.0.tar.gz
| Download URL | autourgos_summarizer-3.1.0.tar.gz |
|---|---|
| Size | 8.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
a620d0d080dd487be6170527bac950f34d08e3122849a32e4fb057efa5c669dd
|
|
BLAKE2b-256 checksum How to use checksums |
45e37b4f8e09ec82a965a47c1d44e3b39478b4c1f488a945f5a8e2e5f45ac44a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.9
|
Release files / autourgos_summarizer-3.1.0-py3-none-any.whl
| Download URL | autourgos_summarizer-3.1.0-py3-none-any.whl |
|---|---|
| Size | 7.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
1cd4d0db0340687ad2245b9c4eeb6ecce444a11c8ed83513643b436dec880614
|
|
BLAKE2b-256 checksum How to use checksums |
f3b206ce41959cbf90da9a9f742b2f1e86d94f5dff35aad3fb5ce7c5ac18974b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.9
|