maf-compaction
Experimental. Releases before 1.0 may change or remove APIs. Importing this package emits
MafCompactionExperimentalWarning.
Compaction strategies for Microsoft Agent Framework that are shaped around the provider's prompt cache. A prompt cache matches on an exact prefix, and every compaction rewrites history inside that prefix, so compaction breaks the cache by construction. The strategies here break it as rarely and as late as they can: they keep a fixed head and tail of the conversation verbatim, change only the band between them, and change it by rules that produce the same bytes on every later turn.
This is an independent package. It is not affiliated with or endorsed by Microsoft.
What is in it
| Strategy | What it does |
|---|---|
AnchoredCompactionStrategy |
Keeps the first and last groups of the conversation whole and shortens the oldest tool results between them, by position, so the compacted prefix is stable. |
MinimumGainAnchoredCompactionStrategy |
The same, with a floor: a collapse that would save too little to repay the cache it invalidates is declined. |
ToolResultAnchoredSummarizationCompactionStrategy |
Has the model write the facts from its tool results into a record through a recall tool, then drops the tool groups the record covers. Ships with make_recall_tool, RecallGate and ToolResultRecallMiddleware, which ask for the record at the right moment. |
UserTurnAnchoredSummarizationCompactionStrategy |
Summarises the user's turns between a fixed head and tail with a summarizer client, in one of three modes. |
ToolResultAndUserTurnAnchoredSummarizationCompactionStrategy |
Runs the two summarising strategies over one conversation, record phase first, with a bounded last-resort chain that merges and rewrites while the prompt is over budget. When nothing safe is left to remove it stops, and the prompt goes out over the limit: a loud failure rather than a quiet loss of facts. |
All of them implement the framework's CompactionStrategy and attach wherever the framework takes one: create_harness_agent, or a CompactionProvider on a plain Agent.
Install
pip install maf-compaction
The package depends on agent-framework-core alone. The summarising strategies take any framework chat client as their summarizer, so install the client package for your provider separately.
Quickstart
Every strategy sizes its decisions with a tokenizer. The one below counts exactly and needs tiktoken, which this package does not install (pip install tiktoken); the framework's CharacterEstimatorTokenizer needs nothing extra and works too, at the cost of precision.
import tiktoken
class Tokenizer:
def __init__(self, encoding: str = "o200k_base") -> None:
self._encoding = tiktoken.get_encoding(encoding)
def count_tokens(self, text: str) -> int:
return len(self._encoding.encode(text))
tokenizer = Tokenizer()
CONTEXT_WINDOW = 128_000 # the model's input limit
MAX_OUTPUT = 4_000 # reserved for the reply
BUDGET = CONTEXT_WINDOW - MAX_OUTPUT
The anchored strategy needs nothing else:
from agent_framework import create_harness_agent
from agent_framework_openai import OpenAIChatClient
from maf_compaction import AnchoredCompactionStrategy
strategy = AnchoredCompactionStrategy(
max_input_tokens=BUDGET,
tokenizer=tokenizer,
keep_head_groups=3, # task and requirements, never touched
keep_tail_groups=4, # the working set, never touched
band_share=0.25, # the oldest banded result keeps 25% of the budget
)
client = OpenAIChatClient(model_id="gpt-6-luna")
agent = create_harness_agent(
client,
name="assistant",
agent_instructions="Answer from the lookups you make. Quote values exactly.",
tools=[lookup_deployment],
max_context_window_tokens=CONTEXT_WINDOW,
max_output_tokens=MAX_OUTPUT,
before_compaction_strategy=strategy,
after_compaction_strategy=strategy,
tokenizer=tokenizer,
)
The record strategy needs the recall tool registered and its middleware installed, because the model cannot be asked for a record from inside a compaction pass:
from maf_compaction import (
AnchoredCompactionStrategy,
RecallGate,
ToolResultAnchoredSummarizationCompactionStrategy,
ToolResultRecallMiddleware,
make_recall_tool,
)
strategy = ToolResultAnchoredSummarizationCompactionStrategy(
max_input_tokens=BUDGET,
tokenizer=tokenizer,
trigger_fraction=0.6, # ask for a record past 60% of the budget
fallback_fraction=0.9, # stop waiting for one past 90%
fallback=AnchoredCompactionStrategy(max_input_tokens=BUDGET, tokenizer=tokenizer),
)
gate = RecallGate()
recall_tool = make_recall_tool(gate, target_tokens=2_000)
recall = ToolResultRecallMiddleware(
max_input_tokens=BUDGET,
tokenizer=tokenizer, # the same tokenizer as the strategy
arm=gate.arm,
disarm=gate.disarm,
trigger_fraction=strategy.trigger_fraction,
record_max_tokens=4_000,
repeat_records=True, # a record per new batch of tool work
reforce=strategy.take_reforce,
)
agent = create_harness_agent(
client,
name="assistant",
agent_instructions="Answer from the lookups you make. Quote values exactly.",
tools=[lookup_deployment, recall_tool],
middleware=[recall],
max_context_window_tokens=CONTEXT_WINDOW,
max_output_tokens=MAX_OUTPUT,
before_compaction_strategy=strategy,
after_compaction_strategy=strategy,
tokenizer=tokenizer,
)
The summarising strategies, RecallGate and ToolResultRecallMiddleware keep one conversation's decisions on the instance, so build this stack, and the agent that holds it, once per session. The middleware raises if a second session reaches it.
The composed strategy takes a record half and a user-turn half built the same way, the user half with remembered_requests=2 unless it runs in the recompacting mode; find_nested_strategy finds the record half inside it for the middleware's reforce, and record_text reads the record the model wrote.
When to use it
Below the context window, not compacting is cheaper than any strategy here: the whole conversation stays cached. These strategies pay off once a conversation outgrows its window, where the framework's own strategies either lose the facts in the tool results or re-bill the prompt on every turn. The measurements behind that, and what each strategy does at the edges, are in the compaction documentation.
What it depends on
The strategies call agent_framework._compaction, the framework's private compaction helpers, for grouping, token annotation and the exclusion flags. Nothing public exposes those, so the dependency on agent-framework-core is pinned to one minor and moved only after the new minor has been read against.
Licence
MIT. See the LICENSE.
Metadata
Release files for maf-compaction 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| maf_compaction-0.1.0.tar.gz | 113.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| maf_compaction-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 232.4 kB
Release files / maf_compaction-0.1.0.tar.gz
| Download URL | maf_compaction-0.1.0.tar.gz |
|---|---|
| Size | 113.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
19b5a09df650961093332fca3b64c27664d54d767748d5dc7b675f7996fb3b3b
|
|
BLAKE2b-256 checksum How to use checksums |
d21abd0b278dbb8165ef93f65b16445ea1f9a9b595b97860052d9e58975dcd06
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.
Transparency logRelease files / maf_compaction-0.1.0-py3-none-any.whl
| Download URL | maf_compaction-0.1.0-py3-none-any.whl |
|---|---|
| Size | 119.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
51d81481a29e1e731e4a447cce4d31452042086a6e1947c2bdab52440ef4e5a9
|
|
BLAKE2b-256 checksum How to use checksums |
1461f3440bebcceb0de4a1271a35343c6c7cf498170a9afe87a1a18d3af45440
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.
Transparency log