Winnow
Separate the grain from the chaff in your coding agent's tool output.
A task-aware pruner for Headroom, powered by Squeez models.
Coding agents spend most of their context on tool output: test runs, builds, grep hits, logs. Usually only a handful of those lines matter for the agent's next step. Winnow plugs a small model trained on real agent traces into Headroom's compression pipeline. It keeps the lines the task needs, never drops an error or a traceback, and turns everything else into retrievable markers.
41% fewer tokens with 0.85 gold-line recall, zero lost error lines, and 0.29 s median latency on a consumer GPU. At equal compression it keeps 15-27 points more of the relevant lines than Headroom's built-in
relevance_split. Full results →
Example
The agent asks "Find the build output block that reports the missing
GetObject method in the UserHandler implementation of
storage.StorageService" and runs go build ./...:
| Before: 119 lines, 1,838 tokens | After: 32 lines, 479 tokens (−74%) |
|---|---|
$ go build ./...
# .../handlers/user_handler.go:23:12: cannot use h
(type *UserHandler) as type storage.StorageService:
*UserHandler does not implement storage.StorageService
(missing GetObject method)
# .../handlers/order_handler.go:45:9: cannot use orderSvc ...
# .../auth/auth.go:12:5: import cycle not allowed
# .../permissions/perm.go:8:2: import cycle not allowed
# .../api/v2/client.go:67:15: undefined: storage.ObjectMetadata
# .../api/v2/client.go:88:20: cannot assign string literal ...
# .../database/migration.go:33:10: cannot find module ...
# .../cache/lru_test.go:78:5: race detector: data race
Read at 0x00c0000a1230 by goroutine 9:
github.com/example/project/internal/cache.(*LRUCache).Get()
.../internal/cache/lru.go:44 +0x7c
... 100 more lines of goroutine dumps,
unrelated packages and test output
|
$ go build ./...
# .../handlers/user_handler.go:23:12: cannot use h
(type *UserHandler) as type storage.StorageService:
*UserHandler does not implement storage.StorageService
(missing GetObject method)
# .../handlers/order_handler.go:45:9: cannot use orderSvc ...
# .../auth/auth.go:12:5: import cycle not allowed
imports github.com/example/project/internal/permissions
<<ccr:e2e73bc3f2f290a77efc7426 52_lines_offloaded>>
--- FAIL: TestUserHandler_Get (0.00s)
user_handler_test.go:45:
Error: Not equal:
expected: &storage.Object{...}
actual : <nil>
FAIL
exit status 1
<<ccr:0bac82ae58f9e3ba08a67445 15_lines_offloaded>>
--- FAIL: TestConcurrentAccess (0.01s)
panic: runtime error: invalid memory address ...
FAIL
exit status 1
<<ccr:407242e9f3f948833f58dd56 23_lines_offloaded>>
...
|
The block the agent asked for is kept verbatim, every failure line survives,
and each <<ccr:…>> marker can be expanded through Headroom's retrieval tool if
the agent needs the dropped lines after all. (Real model output on an example
from the Squeez test set; long paths and lines shortened to fit.)
How it works
flowchart LR
A[Agent tool call] --> B[Headroom proxy]
B --> C{Content router}
C -- "log / search / diff / text" --> D[WinnowCompressor]
C -- "JSON / code / HTML" --> E[Headroom built-ins]
D --> F[Span model<br/>which lines matter?]
F --> G[Safety rules<br/>errors, tracebacks,<br/>±2 context, edges]
G --> H[Render<br/>kept lines verbatim +<br/>ccr markers]
H --> I[(CCR store<br/>hash → original)]
H --> J[Pruned context to LLM]
D -. "no query / too short /<br/>too large / low savings" .-> E
- Gate. Blocks without a task query, under 40 lines, over the device's token budget, or where pruning would save less than 20% pass through untouched, so Headroom's own compressors handle them.
- Score. A span model reads the task query and the whole output and marks the lines that matter for the next step.
- Protect. Lines with errors, failures, exit codes, panics or pytest
Edetails are always kept, as are whole Python tracebacks, two neighbours of every kept line, and the first and last two lines. - Render. Kept lines stay byte-identical. Each dropped run becomes one
<<ccr:HASH N_lines_offloaded>>marker whose original text goes into the CCR store. Hashes are deterministic, so identical input yields identical output and prompt caches stay warm.
The plugin never raises. Any failure (missing torch, model download error, inference error) falls back to Headroom's own path.
Installation
pip install "headroom-winnow[model]"
Installing changes nothing until you opt in:
headroom proxy --compressor winnow
or in Python:
from headroom.transforms.content_router import ContentRouter, ContentRouterConfig
router = ContentRouter(ContentRouterConfig(active_external_compressors=["winnow"]))
On first use Winnow downloads its model,
rbk4209/winnow-pooled-32m
(128 MB), from the Hugging Face Hub at a pinned commit. A GPU is strongly
recommended; on CPU Winnow only scores shorter outputs and leaves the rest to
Headroom.
Models
| Backend | Model | When to use |
|---|---|---|
pooled (default) |
rbk4209/winnow-pooled-32m, 32M line classifier trained for Winnow |
Fast enough to score any tool output on a GPU |
highlighter |
KRLabsOrg/verbatim-rag-modern-bert-v2, 150M span model |
Squeez's original extractive model; slower, so limited to short outputs |
Configuration
| Variable | Default | Meaning |
|---|---|---|
HEADROOM_WINNOW_BACKEND |
pooled |
pooled or highlighter |
HEADROOM_WINNOW_MODEL |
the backend's published model | Hub id or local path |
HEADROOM_WINNOW_REVISION |
pinned commit of the default model | Model revision |
HEADROOM_WINNOW_DEVICE |
auto |
auto (CUDA if available), cuda or cpu |
HEADROOM_WINNOW_DTYPE |
float32 |
float16 is faster only on GPUs with tensor cores |
HEADROOM_WINNOW_MAX_TOKENS |
per backend and device | Largest output the model scores; larger ones go to Headroom |
Results
On the 618-example test split of the Squeez dataset:
| Method | Recall | Token reduction | Error lines lost | p50 latency |
|---|---|---|---|---|
Headroom relevance_split (BM25) |
0.725 | 58.6%¹ | 7,953 | 1 ms |
Headroom relevance_split (hybrid) |
0.753 | 57.6%¹ | 7,771 | 2.2 s |
| Winnow (32M pooled) | 0.847 | 41.4% | 0 | 0.29 s |
¹ Upper bound: the dropped tail is Kompressed inside Headroom, not removed.
At equal compression (50 / 70 / 90% of lines dropped) the pooled model keeps
0.89 / 0.84 / 0.68 of the gold lines, against 0.73 / 0.63 / 0.46 for
relevance_split. Details, the 150M highlighter numbers and reproduction
commands are in benchmarks/RESULTS.md.
Training your own model
training/kaggle_train_pooled.ipynb
trains the 32M pooled classifier on a free Kaggle T4 in about six hours. It
uses Squeez's own training code at a
pinned commit and evaluates on the same test split as the benchmark. Swap
--base-model to try other encoders.
Project layout
headroom_winnow/
compressor.py WinnowCompressor: the headroom.compressor contract, gates, fail-open
selection.py spans → lines, safety rules, rendering with markers
markers.py CCR marker format and deterministic hashes
backends.py highlighter and pooled backends: lazy, pinned, device-aware
benchmarks/ compare.py and RESULTS.md
training/ Kaggle notebook for the pooled model
upstream/ patch proposed to Headroom (see below)
tests/ unit tests with a fake model, router end-to-end, real-model tests
Limitations
- Headroom's lossless fold runs first. Headroom applies its byte-exact fold
(and, for logs and search results,
relevance_split) before the external compressor hook and returns early when either succeeds, so such blocks never reach this plugin. Seetests/test_router.py::test_known_gap_lossless_fold_preempts_external. - Passthroughs skip Headroom's compressors in 0.39.1. When the plugin
declines a block, the router still adopts the unchanged block. The one-line
fix is in
upstream/router-respect-passthrough.patch; the matching test is a strictxfailuntil it lands.
Development
pip install -e ".[dev,model]"
ruff check . && ruff format --check . && mypy headroom_winnow tests benchmarks
pytest # fast tests, no downloads
pytest -m slow # downloads and runs the real highlighter
Acknowledgements
- Headroom for the proxy, CCR store and the external compressor contract this plugin implements.
- Squeez by KRLabs for the task-conditioned pruning approach, the training code, the highlighter model and the dataset.
- Ettin for the 32M encoder the pooled model is fine-tuned from.
License
Metadata
Release files for headroom-winnow 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| headroom_winnow-0.2.0.tar.gz | 26.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| headroom_winnow-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 49.8 kB
Release files / headroom_winnow-0.2.0.tar.gz
| Download URL | headroom_winnow-0.2.0.tar.gz |
|---|---|
| Size | 26.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
a5ffda0f3ee3d8926a950209cd12b2f81a80201493ad625e2cd0c7442d7f7c39
|
|
BLAKE2b-256 checksum How to use checksums |
5468c6ccf615b861670fb13245e32ed43fbe6714c850fc0e169e72eff971abe6
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.
Transparency logRelease files / headroom_winnow-0.2.0-py3-none-any.whl
| Download URL | headroom_winnow-0.2.0-py3-none-any.whl |
|---|---|
| Size | 23.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
14a6289671732dbbcf5d977b45e155d6a1a16a031cd94a370dc742c4f4ebfb1d
|
|
BLAKE2b-256 checksum How to use checksums |
db3c507c9fdd46b4596777f36f41c1482e75793f64da5643a6a7bf92aeee9feb
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.
Transparency log