Per-turn skill routing for LangChain deepagents
A drop-in replacement for the deepagents
SkillsMiddleware: on each user turn a fast judge decides which SKILL.md skills are
needed, and only those are loaded, so a catalog of hundreds stays out of the prompt. Bring any judge: a
hosted model, a self-hosted one, or plain rules. An adapter for
Jev is included.
113.0k → 25.8k input tokens per turn on a 236-skill catalog, and the agent answered 90% of questions correctly against 88% with the whole catalog in the prompt. See the benchmark →
Formerly langchain-loadout.
Quick start · How it works · Your own judge · Results · FAQ · Docs
Why Skill Router
deepagents lists every skill's name and description in the system prompt on every model call. With a handful of skills that is fine. With hundreds, the list takes tens of thousands of tokens per call, and the model has to pick the right procedure from a crowd of similar ones.
Skill Router replaces the built-in SkillsMiddleware with one that decides per user turn:
- Only what the turn needs. A confident pick is loaded with its instructions. When the pick is unsure,
the model gets a short list of up to three candidates to choose from, and it can search the rest of the
catalog with the
find_skilltool. - Any judge, by design. The judge is a small protocol: typed questions in, probabilities out. Use the included Jev adapter, a self-hosted inference model, deterministic rules or anything else that can answer. The routing core depends on neither LangChain nor any provider.
- Calibrated, not guessed. Every question is answered with a probability, so "load it", "offer it" and "skip it" are thresholds you can read and tune. They are not buried in a prompt.
- Conversation-aware. The judge also sees the recent conversation, so a follow-up like "and for April?" still routes to the skill the thread is about.
- Safe by default. A timeout or an outage of the judge gives the agent the full catalog, exactly as it would be without Skill Router. Bad credentials raise instead of hiding a broken setup.
- Cache-friendly. The system prompt is identical on every call, and the turn's skills are written after the user's message. The provider's prompt cache keeps working across turns.
Quick start
pip install "langchain-skill-router[jev]" # with the Jev adapter
pip install langchain-skill-router # with your own judge
Put skills in ./skills/<name>/SKILL.md, with name and description in YAML front matter (the Agent
Skills format deepagents uses). Set TYPESAFE_API_KEY (get a key)
and your model provider's key. Jev is the judge in this example; see below for
your own.
from deepagents import create_deep_agent
from deepagents.backends import FilesystemBackend
from langchain_skill_router.langchain import SkillRouterMiddleware
from langchain_skill_router.providers.jev import JevJudge
backend = FilesystemBackend(root_dir=".", virtual_mode=True)
skill_router = SkillRouterMiddleware(backend=backend, sources=["/skills/"], judge=JevJudge())
agent = create_deep_agent(
model="anthropic:claude-sonnet-5",
backend=backend,
skills=["/skills/"],
middleware=[skill_router], # takes the place of the built-in SkillsMiddleware
)
await agent.ainvoke({"messages": [{"role": "user", "content": "I need a statement for the embassy"}]})
Selection runs on ainvoke and astream. A synchronous run falls back to the ordinary skills
middleware. Requires Python 3.11+.
How it works
On each new user message, Skill Router makes one decision, and the rest of the turn's model calls reuse it:
- Rank and gate, in parallel. The judge ranks the catalog by description against the request and the recent conversation. In the same round it answers whether the request needs a skill at all. A catalog larger than the judge's declared limits (for Jev, 32k tokens and 255 options per call) is split into parts, keeping related skills together, and the part winners are ranked again.
- Verify. The judge reads the start of each top candidate's
SKILL.md, picks one and checks whether each candidate does what the user asked. A ranking that is already sure skips this second call. - Load or offer. A pick verified at 0.9 or higher is loaded with its instructions. Otherwise the model is offered a short list. If no skill is needed, nothing is loaded.
A decision is capped at two seconds by default (Settings.timeout). The routing core has no framework
dependency: SkillRouter works on any list of skills.
Bring your own judge
A judge answers two kinds of question in one call: Pick (a probability for every option) and YesNo
(a probability of yes). It declares its per-call limits, which the router uses to split a large catalog,
and optionally a per-call timeout:
from langchain_skill_router import Answer, Limits, Pick, YesNo
class MyJudge:
limits = Limits(max_tokens=8_000, max_options=100) # what one call can take; Limits() for no limit
timeout = 5.0 # seconds per call, enforced by the router; None for no limit
async def ask(self, state, questions):
# state: {"request": ..., "context": ...}; questions: {key: Pick | YesNo}
# Call a self-hosted model, a classifier or rules here, and answer every key.
return {
key: Answer({option: 1 / len(q.options) for option in q.options}) if isinstance(q, Pick) else Answer({"yes": 0.5})
for key, q in questions.items()
}
Pass it as SkillRouterMiddleware(..., judge=MyJudge()). langchain_skill_router.testing.check_judge checks
an adapter against the contract, and ScriptedJudge answers from a script in your tests. The probabilities
are compared against thresholds, so the closer they are to calibrated, the better the defaults fit.
Results
Benchmark of langchain-skill-router 0.2.2 on a bank-statement assistant built with deepagents: 236 skills, 55 conversations × 5 turns per variant (275 turns each), Jev as the judge, one agent model for all variants. "Perfect selection" always loads the skill the question was written for: the ceiling for any router.
| Metric | Skill Router | Full catalog | Perfect selection |
|---|---|---|---|
| Input tokens per turn | 25.8k | 113.0k | 26.8k |
| Skills in the prompt, characters per call | 2.6k | 89.2k | 2.7k |
| Right skill in front of the model | 85% | 55% | 97% |
| Loaded skill was the right one | 96% (230 of 239) | — | 100% |
| Input from the prompt cache on a new turn | 77% | 95% | 77% |
| Correct answers | 90% | 88% | 88% |
| Answers that depend on a rule inside a skill (21 turns) | 76% | 67% | 62% |
- 4.4× less context per turn, with the right skill in front of the model far more often.
- Accuracy holds. +1.5 points against the full catalog (95% interval −1.5 to +4.7, sign test p = 0.63): the same answers from a fraction of the prompt. Without any skills the agent scored 82%, so the catalog does matter — it just does not have to be in the prompt.
- Where a skill carries a rule the model cannot infer, routing wins: 76% against 67% for the full list.
- A wrong skill costs the most. In the 18 turns where the model worked from a wrong skill, 72% of
answers were correct, against 92% with the right one. Raise
load_atif your catalog has many near-duplicate skills. - The cache holds across turns. 77% of a new turn's first call came from the cache, against 41% before 0.2.2, when the turn's skills were kept out of the conversation.
An earlier run of the same benchmark put Skill Router 5 points below the full catalog. The difference was six skills in the testbed catalog whose instructions contradicted the rule the expected answer was computed from; they dragged down every variant that loads skills, including perfect selection. Skill quality is the ceiling of any router.
Generated data, one judge and one agent model. Fit the thresholds to your own data. Methodology, earlier measurements and known limits →
Configuration
Everything is in Settings, passed as SkillRouterMiddleware(..., settings=Settings(...)):
| Knob | Default | What it does |
|---|---|---|
load_at |
0.9 | How sure verification must be of its pick to load it |
max_suggest |
3 | How many candidates may be offered when the pick is unsure |
need_at |
0.3 | Below this "is a skill needed" probability, nothing is loaded |
skip_verify_at |
0.9 | Ranking this sure skips verification (None turns it off) |
timeout |
2.0 | Seconds for the whole decision, after which the full catalog is used |
need_question, rank_question, … |
general wording | The questions Jev is asked. Naming your domain separates better |
JevJudge(timeout=...) limits each call to Jev. on_decision= on the middleware receives every
decision's trace (probabilities, stage, timing) for logs and metrics.
FAQ
Does Skill Router work without deepagents?
Yes. langchain_skill_router (the core) has no framework or provider dependency:
SkillRouter(catalog, judge).decide(Turn(request, context)) returns what to load and what to suggest.
Do I need Jev?
No. Jev is the included adapter and what Skill Router was measured with, but any Judge works: a self-hosted
inference model behind your own adapter, deterministic rules, or another provider. See
Bring your own judge.
What happens if the judge is slow or down? The turn gets the full catalog, as if Skill Router were not installed, and the trace records why.
What does the judge see?
The request and the recent conversation (the user's messages and the agent's replies, without tool
output) are sent, capped at request_chars and context_chars. So are skill names, descriptions and the
start of the candidates' instructions. Supply your own context= function to send less, or a
self-hosted judge to keep everything in your network.
Does it keep state between turns? No. Every turn is decided from scratch. No checkpointer or extra storage is required.
Documentation
| Read | Covers |
|---|---|
| How it works | Selection flow, judge interface, settings, measurements and known limits |
| Development guide | Architecture, conventions and testing |
| Contributing | Local setup, checks, issues and pull requests |
| Changelog | Release history |
The public API may change in minor releases while the version is 0.x.
Acknowledgements
Skill Router was inspired by Jev, TypeSafe AI's model for typed decisions, and its first round follows TypeSafe's skill suggestion cookbook. Skill Router is an independent open-source project, not affiliated with or endorsed by TypeSafe AI.
License
MIT © 2026 Ivan Deyna
Release files for langchain-skill-router 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| langchain_skill_router-0.3.0.tar.gz | 22.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| langchain_skill_router-0.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 50.3 kB
Release files / langchain_skill_router-0.3.0.tar.gz
| Download URL | langchain_skill_router-0.3.0.tar.gz |
|---|---|
| Size | 22.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
aefa1bdda251015ea781f3d19751349670960d53d22273ceb1ac6efe55205cf5
|
|
BLAKE2b-256 checksum How to use checksums |
f723e084a9f46f16d33a856c9c60a9e01d86eb845aabe0958db5f5e401a051d1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.
Transparency logRelease files / langchain_skill_router-0.3.0-py3-none-any.whl
| Download URL | langchain_skill_router-0.3.0-py3-none-any.whl |
|---|---|
| Size | 27.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
0585ab4e9bcb43b1aafdf09cd30877240eff62c512c3bf5ecc58996a9878224e
|
|
BLAKE2b-256 checksum How to use checksums |
6c2e8ef517f8fb76c9bb082ab86e18ff41e560389a77628a400ce49d323be1da
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 23, 2026.
Transparency log