incontext is a Hermes Agent plugin that gives each LLM request exactly the
output budget still available below Hermes' context-compression boundary. It
prevents a fixed, oversized max_tokens value from consuming the input space
where Hermes must still be able to compress the conversation.
The package supports CPython 3.8 through 3.15, including free-threaded Python 3.14. Unit tests, formatting, linting, and static type checks run across the complete version matrix in CI.
The plugin reads the effective compression window from the installed Hermes
ContextCompressor, asks the selected inference backend to tokenize the exact
provider-visible prompt, and applies:
max_tokens = max(1, compression_window - prompt_tokens)
This addresses the same output-budget arithmetic discussed in NousResearch/hermes-agent#38652.
Installation
Once a release is available on PyPI, install and enable the package using the
same plugin name, incontext:
python -m pip install incontext
hermes plugins enable incontext
Until the first PyPI release, install the reviewed develop revision:
python -m pip install 'git+https://github.com/pomponchik/incontext.git@develop'
hermes plugins enable incontext
Restart the long-running Hermes gateway after installing or upgrading the
Python package. Hermes discovers it through the official
hermes_agent.plugins entry-point group; no source file has to be copied into
$HERMES_HOME/plugins.
Configuration
The bundled vllm backend is selected by default. Its
INCONTEXT_TOKENIZER_URL setting is required and must point to the /tokenize
endpoint of the same vLLM model Hermes uses:
export INCONTEXT_BACKEND='vllm'
export INCONTEXT_TOKENIZER_URL='https://inference.example/tokenize'
export INCONTEXT_TOKENIZER_TIMEOUT_SECONDS='30'
export INCONTEXT_TOKENIZER_USER_AGENT='incontext/0.0.1'
export INCONTEXT_FALLBACK_MARGIN_TOKENS='1024'
Environment variables are loaded through typed skelet.Storage fields backed
by ordered skelet.EnvSource instances. Primary INCONTEXT_* names take
precedence over the supported legacy aliases. Text normalization and blank
value rejection are implemented by the fields' native conversion and
validation rules.
Hermes' model.default, model.context_length, and compression.threshold
remain the source of truth. The plugin constructs Hermes' installed
ContextCompressor and uses its resolved threshold_tokens; it does not copy
version-sensitive threshold arithmetic.
The optional variables are:
| Variable | Default | Meaning |
|---|---|---|
INCONTEXT_BACKEND |
vllm |
Named pristan backend plugin |
INCONTEXT_TOKENIZER_TIMEOUT_SECONDS |
30 |
/tokenize request timeout |
INCONTEXT_TOKENIZER_USER_AGENT |
incontext/0.0.1 |
HTTP user agent |
INCONTEXT_FALLBACK_MARGIN_TOKENS |
1024 |
Extra reserve only when exact tokenization fails |
INCONTEXT_COMPRESSION_WINDOW_TOKENS |
unset | Explicit emergency override for the resolved Hermes boundary |
The former HERMES_VLLM_TOKENIZER_* and
HERMES_DYNAMIC_BUDGET_FALLBACK_MARGIN_TOKENS names are accepted as migration
aliases. New deployments should use the INCONTEXT_* names.
Replacing the inference backend
The budgeting core depends only on the abstract incontext.Backend contract.
It has no import or construction dependency on vLLM. A backend supplies three
operations: its safe diagnostic source, exact count(...), and
clear_cache().
Backend implementations are named pristan plugins in the
incontext.backends entry-point group. The generic skelet environment has a
typed backend field whose default is vllm. At runtime incontext performs
the single named resolution directly:
backend = backends[environment.backend].one()
The incontext distribution itself publishes the vllm entry point. Loading
that entry point imports incontext.vllm_provider, whose only responsibility is
to construct VllmBackend. All /tokenize payload rules, vLLM response fields,
context-length validation, transport settings, and caching live inside that
class rather than in the budgeting core.
A third-party distribution can provide another backend without changing incontext. Its implementation subclasses the stable abstract contract and its plugin module registers a provider under a new name:
# acme_backend/plugin.py
from __future__ import annotations
from typing import Any
from incontext import Backend, backends
class AcmeBackend(Backend):
@property
def source(self) -> str:
return "acme-tokenizer"
def count(
self,
request: dict[str, Any],
*,
context_length: int,
) -> int:
...
def clear_cache(self) -> None:
...
@backends.plugin("acme")
def provide_acme_backend() -> Backend:
return AcmeBackend()
The third-party package makes that module discoverable in pyproject.toml:
[project.entry-points."incontext.backends"]
acme = "acme_backend.plugin"
After installing the package, select it through the same typed configuration field and restart the Hermes process:
export INCONTEXT_BACKEND='acme'
Only the selected provider is instantiated. An unknown name fails .one();
the unique slot rejects duplicate providers under the same name while loading
entry points. Startup therefore fails instead of choosing a backend implicitly.
Each backend owns and validates its backend-specific configuration; the generic
settings object contains only the compression-window and fallback-budget policy.
Safety properties
- With the bundled backend, vLLM applies its real chat template to messages, tools, and
chat_template_kwargs; local tokenizer approximations are not used. - The returned
max_model_lenmust equal Hermes' configured context length. max_tokens,max_completion_tokens, andmax_output_tokensare normalized to one unambiguousmax_tokensfield.- The incoming request is copied and never mutated.
- Exact counts use a bounded, thread-safe cache.
- If
/tokenizefails, Hermes' own rough estimator is used with an additional safety margin. If both counters fail, the middleware leaves the request unchanged instead of taking Hermes down. - Logs contain counts and exception types, never prompts, credentials, or raw provider errors.
The tokenizer endpoint sees the prompt content by design. Run it on a trusted network path and use the same access controls as the inference endpoint.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file incontext-0.0.1.tar.gz.
File metadata
- Download URL: incontext-0.0.1.tar.gz
- Upload date:
- Size: 15.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
a90b49da40003fa5222b6466b330890490b81c8f078ace37079f23ed5fe760fa
|
|
| MD5 |
6f39a2a1b547f4c4c4fcfe482b837e06
|
|
| BLAKE2b-256 |
9fef6b813f385746da836c108b6143ba84fe320e6fc6d508b1c26235f9e706b6
|
Provenance
The following attestation bundles were made for incontext-0.0.1.tar.gz:
Publisher:
release.yml on pomponchik/incontext
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
incontext-0.0.1.tar.gz -
Subject digest:
a90b49da40003fa5222b6466b330890490b81c8f078ace37079f23ed5fe760fa - Sigstore transparency entry: 2499173305
- Sigstore integration time:
-
Permalink:
pomponchik/incontext@caab93a38f00f63fc3a325f8f63a11431155e9aa -
Branch / Tag:
refs/heads/main - Owner: https://github.com/pomponchik
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@caab93a38f00f63fc3a325f8f63a11431155e9aa -
Trigger Event:
push
-
Statement type:
File details
Details for the file incontext-0.0.1-py3-none-any.whl.
File metadata
- Download URL: incontext-0.0.1-py3-none-any.whl
- Upload date:
- Size: 15.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
87e7f9b8ea608c49147990ffe98c94a20a38b0a0ebea245fe9cb70fc234c62f4
|
|
| MD5 |
2e8d67079eecb2a3596e12f0dd00d8ae
|
|
| BLAKE2b-256 |
189f3df4361108c9b111eee12d029de2cb2e0cd8cc5cc79a6faecbc8e47f8f92
|
Provenance
The following attestation bundles were made for incontext-0.0.1-py3-none-any.whl:
Publisher:
release.yml on pomponchik/incontext
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
incontext-0.0.1-py3-none-any.whl -
Subject digest:
87e7f9b8ea608c49147990ffe98c94a20a38b0a0ebea245fe9cb70fc234c62f4 - Sigstore transparency entry: 2499173309
- Sigstore integration time:
-
Permalink:
pomponchik/incontext@caab93a38f00f63fc3a325f8f63a11431155e9aa -
Branch / Tag:
refs/heads/main - Owner: https://github.com/pomponchik
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@caab93a38f00f63fc3a325f8f63a11431155e9aa -
Trigger Event:
push
-
Statement type: