OntoPrune
Neuro-Symbolic Context Pruning Middleware for Local SLMs and Cloud LLMs
English | Español | 📄 Technical Whitepaper (PDF) | 📄 Whitepaper en Español (PDF)
OntoPrune is an ultra-lightweight (<12ms CPU) neuro-symbolic middleware that transforms multi-file source code into minimal typed dependency contracts. By isolating closed-world functional boundaries before attention computation, OntoPrune slashes input tokens by ~60% in modular enterprise architectures, collapses Time-to-First-Token ($TTFT$) latency by 62% on CPU-bound local models (Ollama), guarantees 100% functional test success (Pass@1), and prevents 100% of proprietary code leakage.
📊 Rigorous Empirical Benchmark: 24 Independent Sandbox Runs
Tested across 3 real-world software archetypes in isolated execution sandboxes with live evaluation via pytest:
- Cloud Frontier: Google Gemini (
gemini-3.8-flashvia native SSE streaming) - Local SLM: Ollama (
qwen2.5-coder:7brunning on AMD Ryzen CPU)
| Software Archetype | Backend | Treatment | Input Tokens (Median) | Token Reduction | TTFT (Local/Cloud) | Pass@1 (pytest) |
Leaked IP (Private Lines) |
|---|---|---|---|---|---|---|---|
| Archetype 1: Isolated Algorithm | gemini |
OntoPrune | 818 | -24.3% | 0.1 ms | 100.0% (5/5) | 0 lines |
| Archetype 1: Isolated Algorithm | gemini |
Naive Full | 1,081 | Baseline | 0.1 ms | 100.0% (5/5) | 0 lines |
| Archetype 1: Isolated Algorithm | ollama |
OntoPrune | 731 | -23.5% | 155 ms | 100.0% (5/5) | 0 lines |
| Archetype 1: Isolated Algorithm | ollama |
Naive Full | 956 | Baseline | 12,985 ms | 100.0% (5/5) | 0 lines |
| Archetype 2: Multi-Module Service | gemini |
OntoPrune | 931 | -59.7% | 0.1 ms | 100.0% (8/8) | 0 lines (100% shielded) |
| Archetype 2: Multi-Module Service | gemini |
Naive Full | 2,313 | Baseline | 0.1 ms | 100.0% (8/8) | 74 private lines exposed |
| Archetype 2: Multi-Module Service | ollama |
OntoPrune | 797 | -59.7% | 10,569 ms | 100.0% (8/8) | 0 lines (100% shielded) |
| Archetype 2: Multi-Module Service | ollama |
Naive Full | 1,979 | Baseline | 28,032 ms | 100.0% (8/8) | 74 private lines exposed |
| Archetype 3: Clean Architecture | gemini |
OntoPrune | 1,147 | -60.1% | 0.1 ms | 100.0% (10/10) | 0 lines (100% shielded) |
| Archetype 3: Clean Architecture | gemini |
Naive Full | 2,873 | Baseline | 0.1 ms | 100.0% (10/10) | 91 private lines exposed |
| Archetype 3: Clean Architecture | ollama |
OntoPrune | 952 | -60.5% | 13,092 ms | 100.0% (10/10) | 0 lines (100% shielded) |
| Archetype 3: Clean Architecture | ollama |
Naive Full | 2,412 | Baseline | 34,376 ms | 100.0% (10/10) | 91 private lines exposed |
Key Scientific Findings:
- Zero Semantic Degradation (100.0% Pass@1): In all 24 sandbox runs, code generated with OntoPrune stubs passed 100% of unit tests, proving that contract interfaces and ontology metadata provide all the context an LLM needs.
- 62% TTFT Speedup on Local CPU: In multi-module and clean architecture projects, time-to-first-token dropped from ~34s to ~13s on consumer CPU.
- 100% Intellectual Property Shield: Naive tools (Cursor/Copilot) send entire method bodies (up to 91 private algorithmic lines). OntoPrune sends exactly 0 lines of internal dependencies.
- Architectural Dynamics: In monolithic single scripts, compression is moderate (~24%). In structured multi-module systems, compression is consistent at ~60% net token reduction.
💡 Why OntoPrune Makes Local AI-Assisted Programming Truly Feasible
Running local code intelligence models (Qwen 2.5 Coder 7B, Llama 3 8B, DeepSeek Coder via Ollama or llama.cpp) on consumer laptops and developer workstations has historically faced an insurmountable barrier. OntoPrune dismantles this barrier through three architectural pillars:
1. Collapsing the "Pre-fill Latency Wall" on CPU ($TTFT$)
- The Bottleneck: On CPUs and consumer laptops without high-end 24GB GPUs, token decoding speed is acceptable (15–25 t/s). However, the Pre-fill phase (Prompt Evaluation) is heavily memory-bandwidth bound and saturates RAM.
- Without OntoPrune: Standard AI assistants inject entire raw files and bloated multi-module contexts (2,500 – 3,000 tokens), causing a 35-second freeze before the first token appears. A half-minute freeze on every code turn shatters the developer's flow state.
- With OntoPrune: By pruning raw context into concise typed contracts (~800 tokens), pre-fill latency collapses from 34.3s down to 13.0s (a 62% to 74% drop in TTFT) on standard CPUs (AMD Ryzen 5600G). Local inference shifts from unusable to completely interactive.
2. Eliminating Attention Dilution in Small Language Models (SLMs)
- Small models (3B to 7B parameters) lack the massive context comprehension of hundred-billion-parameter cloud models. Dumping hundreds of lines of irrelevant database or third-party implementations dilutes attention, leading to hallucinated method calls.
- OntoPrune constructs a mathematically sound closed-world contract. The SLM only sees allowable methods and domain types. In our 24 empirical sandbox runs, Qwen 2.5 Coder 7B achieved 100.0% Pass@1 on
pytestsuites.
3. Absolute Data Sovereignty and Air-Gapped Privacy
- OntoPrune operates in $<12\text{ ms}$ on CPU using local AST visitors and in-memory RDF graphs with zero network calls. It provides an airtight guarantee: 0 lines of internal proprietary code are ever leaked to external APIs.
📦 Installation
# Standard installation (native Python support):
pip install ontoprune
# With multi-language support (Flutter/Dart, Java, TypeScript via Tree-sitter):
pip install "ontoprune[languages]"
# For development, benchmarks, and tests:
pip install "ontoprune[dev,benchmark,languages]"
🌐 Universal Multi-Language Support
OntoPrune automatically detects file types and resolves dependencies across project boundaries:
| Language | Extension | AST Engine | Output Format |
|---|---|---|---|
| Python | .py |
Native Python ast |
def name(args) -> Ret: ... |
| Flutter / Dart | .dart |
tree-sitter-dart |
abstract class ... { Ret method(); } |
| Java / Spring Boot | .java |
tree-sitter-java |
public interface ... { Ret method(); } |
| TypeScript / React | .ts, .tsx, .js |
tree-sitter-typescript |
export interface ... { method(): Ret; } |
🛠️ Usage Modes
1. Native Model Context Protocol (MCP) Server
OntoPrune runs out of the box as an MCP server (ontoprune-mcp) compatible with Claude Desktop, Cursor, Gemini CLI, or Antigravity IDE:
ontoprune-mcp
Configuration in claude_desktop_config.json:
{
"mcpServers": {
"ontoprune": {
"command": "ontoprune-mcp"
}
}
}
Exposed MCP Tools:
prune_context(file_path, target_symbol, format='stubs'): Extracts the minimal dependency contract resolving cross-file imports.verify_response(response_code, contract_or_file): Deterministically validates generated code against authorized contracts.
2. Command Line Interface (CLI)
# Prune a target method across multi-file projects:
ontoprune translate src/services/OrderService.java processOrder --format stubs
# Direct streaming pipeline with local Ollama:
ontoprune translate services/order_service.py procesar_orden | ollama run qwen2.5-coder:3b
# Deterministically verify LLM output against contract:
ontoprune check --file generated_solution.py --contract contract.py
3. Python API
import ontoprune
# 1. Prune a multi-module project to a minimal typed contract
context = ontoprune.translate(
"services/order_service.py",
target="procesar_orden",
fmt="stubs",
multi_module=True,
)
print(context)
# 2. Verify model output against the contract
violations = ontoprune.check(llm_code_response, against=context)
if not violations:
print("Code is 100% compliant and free of hallucinations!")
🧪 Test Suite
pytest tests/
# 46 passed in 1.40s
🔬 Reproduce Empirical Benchmarks
You can independently replicate the benchmark evaluation in isolated sandboxes on your machine:
# 1. Quick deterministic validation in sandbox (no API key needed):
ontoprune benchmark --backend mock --analyze
# or via Make:
make benchmark-mock
# 2. Live evaluation against Google Gemini (requires GEMINI_API_KEY):
ontoprune benchmark --backend gemini --analyze
# or via Make:
make benchmark-gemini
# 3. Live local evaluation against Ollama CPU (requires running Ollama):
ontoprune benchmark --backend ollama --analyze
# or via Make:
make benchmark-ollama
📄 Whitepapers & Publications
- English Technical Whitepaper (PDF)
- Whitepaper Técnico en Español (PDF)
- Community Article / Blog Post
- Artículo Comunitario en Español
👤 Author
Vigmar Carlo
- GitHub: @vigmarcarlo
- Repository: https://github.com/vigmarcarlo/OntoPrune
- License: MIT
Metadata
Release files for ontoprune 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| ontoprune-0.3.0.tar.gz | 388.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| ontoprune-0.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 425.9 kB
Release files / ontoprune-0.3.0.tar.gz
| Download URL | ontoprune-0.3.0.tar.gz |
|---|---|
| Size | 388.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
3fbe2d2eb6d6681c8427602d2a6c31a2fefd58f8ac2c5b821858ef3ad2ae21ef
|
|
BLAKE2b-256 checksum How to use checksums |
bfa1c5d1df85ee21a7d79dcc3c57bc7bd172a414139d32115e06fa83e368b6a3
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.7 {"installer":{"name":"uv","version":"0.12.7","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Linux Mint","version":"21.1","id":"vera","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|
Release files / ontoprune-0.3.0-py3-none-any.whl
| Download URL | ontoprune-0.3.0-py3-none-any.whl |
|---|---|
| Size | 37.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
89236a313797386c78232d4b813026a12eabc746c1f7d7361c89fd32ab259ac4
|
|
BLAKE2b-256 checksum How to use checksums |
9467cd663032e770c8006e89868b9bd0bdb3e332e5edc9a220433ebdb436c19b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.7 {"installer":{"name":"uv","version":"0.12.7","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Linux Mint","version":"21.1","id":"vera","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|