Skip to main content

maf-compaction

PyPI Python License

Experimental. Releases before 1.0 may change or remove APIs. Importing this package emits MafCompactionExperimentalWarning.

Compaction strategies for Microsoft Agent Framework that are shaped around the provider's prompt cache. A prompt cache matches on an exact prefix, and every compaction rewrites history inside that prefix, so compaction breaks the cache by construction. The strategies here break it as rarely and as late as they can: they keep a fixed head and tail of the conversation verbatim, change only the band between them, and change it by rules that produce the same bytes on every later turn.

This is an independent package. It is not affiliated with or endorsed by Microsoft.

What is in it

Strategy What it does
AnchoredCompactionStrategy Keeps the first and last groups of the conversation whole and shortens the oldest tool results between them, by position, so the compacted prefix is stable.
MinimumGainAnchoredCompactionStrategy The same, with a floor: a collapse that would save too little to repay the cache it invalidates is declined.
ToolResultAnchoredSummarizationCompactionStrategy Has the model write the facts from its tool results into a record through a recall tool, then drops the tool groups the record covers. Ships with make_recall_tool, RecallGate and ToolResultRecallMiddleware, which ask for the record at the right moment.
UserTurnAnchoredSummarizationCompactionStrategy Summarises the user's turns between a fixed head and tail with a summarizer client, in one of three modes.
ToolResultAndUserTurnAnchoredSummarizationCompactionStrategy Runs the two summarising strategies over one conversation, record phase first, with a bounded last-resort chain that merges and rewrites while the prompt is over budget. When nothing safe is left to remove it stops, and the prompt goes out over the limit: a loud failure rather than a quiet loss of facts.

All of them implement the framework's CompactionStrategy and attach wherever the framework takes one: create_harness_agent, or a CompactionProvider on a plain Agent.

Install

pip install maf-compaction

The package depends on agent-framework-core alone. The summarising strategies take any framework chat client as their summarizer, so install the client package for your provider separately.

Quickstart

Every strategy sizes its decisions with a tokenizer. The one below counts exactly and needs tiktoken, which this package does not install (pip install tiktoken); the framework's CharacterEstimatorTokenizer needs nothing extra and works too, at the cost of precision.

import tiktoken


class Tokenizer:
    def __init__(self, encoding: str = "o200k_base") -> None:
        self._encoding = tiktoken.get_encoding(encoding)

    def count_tokens(self, text: str) -> int:
        return len(self._encoding.encode(text))


tokenizer = Tokenizer()
CONTEXT_WINDOW = 128_000  # the model's input limit
MAX_OUTPUT = 4_000  # reserved for the reply
BUDGET = CONTEXT_WINDOW - MAX_OUTPUT

The anchored strategy needs nothing else:

from agent_framework import create_harness_agent
from agent_framework_openai import OpenAIChatClient
from maf_compaction import AnchoredCompactionStrategy

strategy = AnchoredCompactionStrategy(
    max_input_tokens=BUDGET,
    tokenizer=tokenizer,
    keep_head_groups=3,  # task and requirements, never touched
    keep_tail_groups=4,  # the working set, never touched
    band_share=0.25,  # the oldest banded result keeps 25% of the budget
)

client = OpenAIChatClient(model_id="gpt-6-luna")
agent = create_harness_agent(
    client,
    name="assistant",
    agent_instructions="Answer from the lookups you make. Quote values exactly.",
    tools=[lookup_deployment],
    max_context_window_tokens=CONTEXT_WINDOW,
    max_output_tokens=MAX_OUTPUT,
    before_compaction_strategy=strategy,
    after_compaction_strategy=strategy,
    tokenizer=tokenizer,
)

The record strategy needs the recall tool registered and its middleware installed, because the model cannot be asked for a record from inside a compaction pass:

from maf_compaction import (
    AnchoredCompactionStrategy,
    RecallGate,
    ToolResultAnchoredSummarizationCompactionStrategy,
    ToolResultRecallMiddleware,
    make_recall_tool,
)

strategy = ToolResultAnchoredSummarizationCompactionStrategy(
    max_input_tokens=BUDGET,
    tokenizer=tokenizer,
    trigger_fraction=0.6,  # ask for a record past 60% of the budget
    fallback_fraction=0.9,  # stop waiting for one past 90%
    fallback=AnchoredCompactionStrategy(max_input_tokens=BUDGET, tokenizer=tokenizer),
)

gate = RecallGate()
recall_tool = make_recall_tool(gate, target_tokens=2_000)
recall = ToolResultRecallMiddleware(
    max_input_tokens=BUDGET,
    tokenizer=tokenizer,  # the same tokenizer as the strategy
    arm=gate.arm,
    disarm=gate.disarm,
    trigger_fraction=strategy.trigger_fraction,
    record_max_tokens=4_000,
    repeat_records=True,  # a record per new batch of tool work
    reforce=strategy.take_reforce,
)

agent = create_harness_agent(
    client,
    name="assistant",
    agent_instructions="Answer from the lookups you make. Quote values exactly.",
    tools=[lookup_deployment, recall_tool],
    middleware=[recall],
    max_context_window_tokens=CONTEXT_WINDOW,
    max_output_tokens=MAX_OUTPUT,
    before_compaction_strategy=strategy,
    after_compaction_strategy=strategy,
    tokenizer=tokenizer,
)

The summarising strategies, RecallGate and ToolResultRecallMiddleware keep one conversation's decisions on the instance, so build this stack, and the agent that holds it, once per session. The middleware raises if a second session reaches it.

The composed strategy takes a record half and a user-turn half built the same way, the user half with remembered_requests=2 unless it runs in the recompacting mode; find_nested_strategy finds the record half inside it for the middleware's reforce, and record_text reads the record the model wrote.

When to use it

Below the context window, not compacting is cheaper than any strategy here: the whole conversation stays cached. These strategies pay off once a conversation outgrows its window, where the framework's own strategies either lose the facts in the tool results or re-bill the prompt on every turn. The measurements behind that, and what each strategy does at the edges, are in the compaction documentation.

What it depends on

The strategies call agent_framework._compaction, the framework's private compaction helpers, for grouping, token annotation and the exclusion flags. Nothing public exposes those, so the dependency on agent-framework-core is pinned to one minor and moved only after the new minor has been read against.

Licence

MIT. See the LICENSE.

Metadata

Release files for maf-compaction 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for maf-compaction 0.1.0
File Size Uploaded
maf_compaction-0.1.0.tar.gz 113.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for maf-compaction 0.1.0
File Interpreter ABI Platform
maf_compaction-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 232.4 kB

Release files / maf_compaction-0.1.0.tar.gz

Download URL maf_compaction-0.1.0.tar.gz
Size 113.2 kB
Tags Source
SHA-256 checksum
How to use checksums
19b5a09df650961093332fca3b64c27664d54d767748d5dc7b675f7996fb3b3b
BLAKE2b-256 checksum
How to use checksums
d21abd0b278dbb8165ef93f65b16445ea1f9a9b595b97860052d9e58975dcd06
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.

Transparency log

Release files / maf_compaction-0.1.0-py3-none-any.whl

Download URL maf_compaction-0.1.0-py3-none-any.whl
Size 119.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
51d81481a29e1e731e4a447cce4d31452042086a6e1947c2bdab52440ef4e5a9
BLAKE2b-256 checksum
How to use checksums
1461f3440bebcceb0de4a1271a35343c6c7cf498170a9afe87a1a18d3af45440
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page