Skip to main content
Archived

This project has been archived by its maintainers, and is no longer receiving any updates.

autourgos-summarizer

Scratchpad compression middleware for Autourgos agents.

Automatically summarizes long reasoning chains so your agent never runs out of token space — even on tasks with dozens of iterations.


Why use this?

LLM agents accumulate a scratchpad — a growing log of thoughts, tool calls, and observations. On long tasks this scratchpad can:

  • Hit the LLM's context window limit and crash
  • Slow down responses (more tokens = more cost and latency)
  • Confuse the LLM with too much irrelevant history

AutoSummarizeMiddleware solves this by compressing the scratchpad in the background every N iterations (or when it exceeds a size limit), keeping only what matters: key findings, tool results, and current status.


Install

pip install autourgos-summarizer

Depends on autourgos-agent. Works with any Autourgos agent.


Quick Start

from autourgos_summarizer import AutoSummarizeMiddleware
from autourgos_agent import Agent

summarizer = AutoSummarizeMiddleware(
    summarize_every=5,          # compress every 5 iterations
    max_scratchpad_chars=15000, # also compress if scratchpad exceeds 15k chars
)

agent = Agent(llm=my_llm, middleware=[summarizer])
result = agent.invoke("Research the latest breakthroughs in quantum computing")
print(result)

When used with a verbose Agent (verbose=True), AutoSummarizeMiddleware also narrates its own actions into the agent's trace, right alongside the Thought/Action/Observation lines, e.g.:

[Summarizer] Compressed scratchpad (iteration 5, was 18342 chars).

How it works

AutoSummarizeMiddleware hooks into on_iteration_start. Before each iteration it checks:

  1. Is this a multiple of summarize_every? (e.g. iteration 5, 10, 15…)
  2. Is the scratchpad longer than max_scratchpad_chars?

If either is true, it calls the LLM (middleware's own llm if set, otherwise agent.llm) with a compression prompt that asks it to distill the scratchpad into:

[Summary of steps 1-N]
Key findings: ...
Tool results: ...
Current status: ...

This summary replaces the full scratchpad before the next LLM call. The agent continues from where it left off, but with a much smaller context.

Summarization runs synchronously (the agent loop is paused at this point), and a threading.Lock prevents duplicate summarizations when parallel tool callbacks fire.


Use a dedicated LLM for summarization

You can pass a separate llm to AutoSummarizeMiddleware. This LLM is used only for compressing the scratchpad — great for keeping costs low by using a cheaper/faster model just for this job.

from autourgos_summarizer import AutoSummarizeMiddleware
from autourgos_openaichat import OpenAIChatModel

# Main agent uses a powerful model
main_llm = OpenAIChatModel(model="gpt-4o")

# Summarizer uses a cheap fast model
cheap_llm = OpenAIChatModel(model="gpt-4o-mini")

summarizer = AutoSummarizeMiddleware(
    summarize_every=5,
    llm=cheap_llm,  # overrides agent.llm for summarization
)

agent = Agent(llm=main_llm, middleware=[summarizer])
result = agent.invoke("Research the latest breakthroughs in quantum computing")
print(result)

If llm is not provided, it falls back to agent.llm automatically.


Parameters

Parameter Type Default Description
summarize_every int | None 5 Summarize every N iterations. None = disable iteration-based trigger.
max_scratchpad_chars int 15000 Also trigger if scratchpad exceeds this many characters.
llm any None LLM for summarization. Needs .invoke(prompt). Falls back to agent.llm if not set.

Combine with other middleware

from autourgos_summarizer import AutoSummarizeMiddleware
from autourgos_history import AgentHistoryMiddleware

history    = AgentHistoryMiddleware()
summarizer = AutoSummarizeMiddleware(summarize_every=5)

agent = Agent(llm=my_llm, middleware=[summarizer, history])

Requirements

  • Python 3.9+
  • Any Autourgos agent that exposes agent.scratchpad, agent.llm, and agent.current_query (requires autourgos-agent>=2.0.2)

License

MIT — see LICENSE

Metadata

Release files for autourgos-summarizer 3.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for autourgos-summarizer 3.1.0
File Size Uploaded
autourgos_summarizer-3.1.0.tar.gz 8.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for autourgos-summarizer 3.1.0
File Interpreter ABI Platform
autourgos_summarizer-3.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 16.0 kB

Release files / autourgos_summarizer-3.1.0.tar.gz

Download URL autourgos_summarizer-3.1.0.tar.gz
Size 8.6 kB
Tags Source
SHA-256 checksum
How to use checksums
a620d0d080dd487be6170527bac950f34d08e3122849a32e4fb057efa5c669dd
BLAKE2b-256 checksum
How to use checksums
45e37b4f8e09ec82a965a47c1d44e3b39478b4c1f488a945f5a8e2e5f45ac44a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.9

Release files / autourgos_summarizer-3.1.0-py3-none-any.whl

Download URL autourgos_summarizer-3.1.0-py3-none-any.whl
Size 7.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
1cd4d0db0340687ad2245b9c4eeb6ecce444a11c8ed83513643b436dec880614
BLAKE2b-256 checksum
How to use checksums
f3b206ce41959cbf90da9a9f742b2f1e86d94f5dff35aad3fb5ce7c5ac18974b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.9
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page