small-model-harness
Session-level intelligence for small LLMs (2B–4B parameters)
Install · Quick Start · Features · API · Architecture · Research
Small models fail on tool calling in predictable ways. This harness compensates for those failures deterministically — no extra LLM calls, no retries, no cloud dependency.
Why this exists
A 2B–4B parameter model will:
- Skip tool calls entirely — acknowledge uncertainty, then confabulate an answer instead of calling the tool
- Produce malformed JSON — trailing commas, wrong types, camelCase instead of snake_case
- Select the wrong tool — when given 20+ options, small models pick randomly
- Get stuck in loops — call the same tool repeatedly hoping for a different result
- Run out of context — verbose tool responses fill the window before the task completes
This harness fixes all five — deterministically, at the session level, without modifying the model.
Install
pip install small-model-harness
With optional extras:
pip install "small-model-harness[full]" # PDF + EPUB support
pip install "small-model-harness[pydantic-deep]" # Full agent framework
pip install "small-model-harness[dev]" # Testing tools
Quick start
from small_model_harness import (
create_harness_session,
ExecutionRecord,
repair_tool_call,
score_tool_call_confidence,
rank_tools,
compact_tool_response,
)
# 1. Create a session
session = create_harness_session(n_ctx=4096)
# 2. Before showing tools, rank by relevance (show top 5, not all 211)
ranked = rank_tools("play some music", tool_schemas, top_n=5)
prompt = build_compact_tool_prompt("play some music", tool_schemas, top_n=5)
# 3. Model outputs malformed JSON — repair it deterministically
args, name, fixes = repair_tool_call(
'{"tool": "play_audio", "arguments": {"filePath": "/a/b.wav", "volume": "0.8"}}',
tool_schemas,
)
# name = "play_audio"
# args = {"file_path": "/a/b.wav", "volume": 0.8} # renamed + coerced
# fixes = ["Renamed 'filePath' → 'file_path'", "Converted 'volume' from string to number"]
# 4. Score confidence — should we escalate to cloud?
score = score_tool_call_confidence(name, args, tool_schemas[name], harness=session)
if score.should_escalate:
escalate_to_cloud()
# 5. Record the execution
session.record_execution(ExecutionRecord(
timestamp="2026-08-31T12:00:00Z",
tool_name=name,
arguments=args,
status="completed",
result="Playing audio...",
duration_ms=340.0,
))
# 6. Get steering hints for the next system prompt
steering = session.get_steering_prompt()
# → "Avoid: play_audio (failed 3 times this session)"
# 7. Compact verbose responses to save context
compacted, was_compacted = compact_tool_response("list_voices", huge_json, max_tokens=200)
Features
Deterministic tool repair
Fixes malformed output without re-prompting. The model's intent is usually correct — only the serialization drifted.
| What breaks | How it's fixed |
|---|---|
Trailing commas: {"a": 1,} |
JSON repair strips them |
Wrong types: "5" instead of 5 |
Type coercion from schema |
Key naming: maxResults vs max_results |
Key renaming (camelCase/snake_case) |
| Missing defaults: model forgets optional args | Default injection from schema |
Markdown fences: ```json ... ``` |
Stripped before parsing |
from small_model_harness import repair_tool_call
args, name, fixes = repair_tool_call(malformed_output, tool_schemas)
# Always returns (args, tool_name, fixes_applied) or (None, None, error)
Confidence scoring
Multi-signal scoring to know when to trust the model vs escalate to cloud.
| Signal | Weight | What it measures |
|---|---|---|
| Schema match | 35% | Are args valid against the tool schema? |
| History | 25% | Has this tool succeeded before in this session? |
| Completeness | 25% | Are required arguments present? |
| Repair cost | 15% | How much did we have to fix? |
from small_model_harness import score_tool_call_confidence
score = score_tool_call_confidence(
"search_web", args, schema, harness=session, repairs_applied=fixes
)
print(score.overall) # 0.85
print(score.should_escalate) # False
print(format_confidence_summary(score))
# Confidence: 85% (HIGH)
# Schema: 100%
# History: 90%
# Completeness: 100%
# Repair cost: 90%
Progressive tool disclosure
Don't show the model all 211 tools — show the 5 most relevant ones. Reduces selection errors dramatically.
from small_model_harness import rank_tools, build_compact_tool_prompt
# Rank tools by query relevance
ranked = rank_tools("play some music", tool_schemas, top_n=5)
# Returns: [ToolScore(name="play_audio", score=0.85, ...), ...]
# Build a compact prompt with only relevant tools
prompt = build_compact_tool_prompt("play some music", tool_schemas, top_n=5)
# "You have access to the following tools:
# - play_audio: Play an audio file | params: file_path: string*, volume: number
# - speak_text: Convert text to speech | params: text: string*, voice: string
# ..."
Adaptive steering
After repeated failures, inject "avoid this" hints into the system prompt.
session.record_execution(ExecutionRecord(
timestamp="...", tool_name="play_audio", status="failed", ...
))
session.record_execution(ExecutionRecord(
timestamp="...", tool_name="play_audio", status="failed", ...
))
steering = session.get_steering_prompt()
# → "Avoid: play_audio (failed 2 times this session)"
Loop detection
Catch when the model calls the same tool repeatedly.
from small_model_harness import detect_failure_loop
calls = ["search", "search", "search"]
tool = detect_failure_loop(calls)
# → "search" (loop detected)
Context pressure monitoring
Track token usage and trigger compaction before overflow.
session.update_context_pressure(estimated_tokens=3500, n_ctx=4096)
print(session.context_pressure) # 0.85
print(session.is_context_warned) # True (≥75%)
print(session.is_context_critical) # False (<90%)
Response compaction
Shrink verbose tool responses to save context space.
from small_model_harness import compact_tool_response
compacted, was_compacted = compact_tool_response(
"play_audio",
'{"audio_file": "/long/path/file.wav", "duration": 300, "sample_rate": 44100, ...}',
max_tokens=50,
)
# → '{"audio_file": "/long/path/file.wav", "duration": 300}' (compacted)
API
Core functions
| Function | Purpose |
|---|---|
create_harness_session(session_id, budget, n_ctx) |
Create a new session |
repair_tool_call(raw_output, tool_schemas) |
Fix malformed model output |
score_tool_call_confidence(name, args, schema, harness, repairs) |
Score tool call confidence |
rank_tools(query, tool_schemas, harness, top_n) |
Rank tools by relevance |
build_compact_tool_prompt(query, tool_schemas, harness, top_n) |
Build compact tool prompt |
compact_tool_response(tool_name, response, max_tokens) |
Shrink verbose responses |
detect_failure_loop(recent_calls, window) |
Detect same-tool loops |
estimate_tokens(text) |
Rough token estimate |
tool_schema(name, description, properties) |
Generate OpenAI tool schema |
pydantic_tool_schema(model, name, description) |
Generate schema from pydantic model |
HarnessState
| Method | Purpose |
|---|---|
record_execution(record) |
Record a tool call, update stats |
update_context_pressure(tokens, n_ctx) |
Update pressure estimate |
get_steering_prompt() |
Get hints for system prompt |
get_avoided_tools_prompt() |
List tools to avoid |
format_audit_trail() |
Human-readable execution log |
generate_session_summary() |
One-line session summary |
| Property | Type | Meaning |
|---|---|---|
success_rate |
float | 0.0–1.0 |
is_context_warned |
bool | pressure ≥ 0.75 |
is_context_critical |
bool | pressure ≥ 0.90 |
is_budget_exhausted |
bool | no calls remaining |
steering_hints |
list[str] | Current steering hints |
avoided_tools |
set[str] | Tools marked as avoided |
Models (pydantic v2)
| Model | Fields |
|---|---|
HarnessState |
session_id, started_at, budget, execution_history, tool stats, steering, pressure |
ExecutionRecord |
timestamp, tool_name, arguments, status, result, error, duration_ms, confidence |
FileInput |
path, content, mime_type, size_bytes |
ConfidenceScore |
overall, schema_match, history_score, completeness, repair_cost, should_escalate |
ToolScore |
name, score, reason, schema |
Architecture
┌──────────────────────────────────────────────────────────────┐
│ Your application │
│ (streaming loop, agent, CLI, whatever) │
├──────────────────────────────────────────────────────────────┤
│ small-model-harness │
│ │
│ ┌─────────────┐ ┌──────────────┐ ┌────────────────────┐ │
│ │ Tool Repair │ │ Confidence │ │ Tool Disclosure │ │
│ │ JSON fix │ │ Schema │ │ Rank by relevance │ │
│ │ Type coerce │ │ History │ │ Top-N selection │ │
│ │ Key rename │ │ Completeness│ │ Compact prompts │ │
│ │ Defaults │ │ Repair cost │ │ │ │
│ └─────────────┘ └──────────────┘ └────────────────────┘ │
│ │
│ ┌──────────────────────────────────────────────────────┐ │
│ │ Session Intelligence │ │
│ │ Failure tracking · Adaptive steering · Loop detection │ │
│ │ Context pressure · Response compaction · Audit trail │ │
│ └──────────────────────────────────────────────────────┘ │
├──────────────────────────────────────────────────────────────┤
│ llama.cpp / vLLM / OpenAI-compatible endpoint │
└──────────────────────────────────────────────────────────────┘
How the pieces fit together
# The full pipeline — from raw model output to repaired, scored, tracked call
# 1. Model outputs something broken
raw = model.generate(prompt)
# 2. Repair it
args, name, fixes = repair_tool_call(raw, tool_schemas)
# 3. Score confidence
score = score_tool_call_confidence(name, args, tool_schemas[name], session, fixes)
# 4. If confidence is low, escalate
if score.should_escalate:
raw = cloud_model.generate(prompt)
args, name, fixes = repair_tool_call(raw, tool_schemas)
# 5. Execute the tool
result = execute_tool(name, args)
# 6. Compact the response
compacted, _ = compact_tool_response(name, result, max_tokens=200)
# 7. Record execution
session.record_execution(ExecutionRecord(
timestamp=now(), tool_name=name, arguments=args,
status="completed", result=compacted, confidence=score.overall,
))
# 8. Feed back to model
messages.append({"role": "tool", "content": compacted})
Research
This harness is informed by:
- Cho et al. (2026) — "It's Not the Size: Harness Design Determines Operational Stability in Small Language Models" (arXiv:2605.12129). Found that a 4-stage scaffold takes 2B models from 58% to 95% task success.
- ManiFreeBird (2026) — Deterministic tool repair eliminates retry loops. Key insight: "structural failures are not reasoning failures."
- Gorilla (Patil et al., 2023) — Tool retrieval reduces selection hallucination. Show 3–5 relevant tools, not all.
- NVIDIA (2025) — SLM-first routing with confidence-based cloud escalation keeps 80–90% of steps local.
Pydantic-deep integration (optional)
For full agent capabilities (context compaction, subagents, skills, memory, planning):
pip install "small-model-harness[pydantic-deep]"
from small_model_harness.pydantic_deep_integration import (
build_small_model_agent,
run_with_harness,
)
result = build_small_model_agent(
model_url="http://localhost:8080/v1",
model_name="qwen3.5-4b",
n_ctx=4096,
)
run_result = await run_with_harness(result, "Plan a 3-step audio mixing task")
Running tests
pip install "small-model-harness[dev]"
pytest # Run all 142 tests
pytest tests/test_harness.py # Core harness tests only
pytest -v # Verbose output
Contributing
See CONTRIBUTING.md.
License
MIT — see LICENSE.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file small_model_harness-0.2.0.tar.gz.
File metadata
- Download URL: small_model_harness-0.2.0.tar.gz
- Upload date:
- Size: 42.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.12.10
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
392baa0cf1efe0bd0315c38f0abea9355b4ee12c3f230e2ddc2e735aa0145760
|
|
| MD5 |
c36b39dfd95946a9b3a1d48760499dd7
|
|
| BLAKE2b-256 |
b9e3f23550259164637a801ceccd8998b2c0d0424faa48f7a967aab490ad4257
|
File details
Details for the file small_model_harness-0.2.0-py3-none-any.whl.
File metadata
- Download URL: small_model_harness-0.2.0-py3-none-any.whl
- Upload date:
- Size: 28.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.12.10
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
98d4724b177c328b061a37893b7f2338f2a1c264b2287dd02d213e398103cf6d
|
|
| MD5 |
4daee4766f36f6969a91038a8b041406
|
|
| BLAKE2b-256 |
b2c2002f411f40fef9b2586ffe82f94683dc757596b9c80adb88ab48756703f9
|