Skip to main content

OntoPrune

Neuro-Symbolic Context Pruning Middleware for Local SLMs and Cloud LLMs

Tests Python License: MIT MCP Server Languages

English | Español | 📄 Technical Whitepaper (PDF) | 📄 Whitepaper en Español (PDF)


OntoPrune is an ultra-lightweight (<12ms CPU) neuro-symbolic middleware that transforms multi-file source code into minimal typed dependency contracts. By isolating closed-world functional boundaries before attention computation, OntoPrune slashes input tokens by ~60% in modular enterprise architectures, collapses Time-to-First-Token ($TTFT$) latency by 62% on CPU-bound local models (Ollama), guarantees 100% functional test success (Pass@1), and prevents 100% of proprietary code leakage.


📊 Rigorous Empirical Benchmark: 24 Independent Sandbox Runs

Tested across 3 real-world software archetypes in isolated execution sandboxes with live evaluation via pytest:

  • Cloud Frontier: Google Gemini (gemini-3.8-flash via native SSE streaming)
  • Local SLM: Ollama (qwen2.5-coder:7b running on AMD Ryzen CPU)
Software Archetype Backend Treatment Input Tokens (Median) Token Reduction TTFT (Local/Cloud) Pass@1 (pytest) Leaked IP (Private Lines)
Archetype 1: Isolated Algorithm gemini OntoPrune 818 -24.3% 0.1 ms 100.0% (5/5) 0 lines
Archetype 1: Isolated Algorithm gemini Naive Full 1,081 Baseline 0.1 ms 100.0% (5/5) 0 lines
Archetype 1: Isolated Algorithm ollama OntoPrune 731 -23.5% 155 ms 100.0% (5/5) 0 lines
Archetype 1: Isolated Algorithm ollama Naive Full 956 Baseline 12,985 ms 100.0% (5/5) 0 lines
Archetype 2: Multi-Module Service gemini OntoPrune 931 -59.7% 0.1 ms 100.0% (8/8) 0 lines (100% shielded)
Archetype 2: Multi-Module Service gemini Naive Full 2,313 Baseline 0.1 ms 100.0% (8/8) 74 private lines exposed
Archetype 2: Multi-Module Service ollama OntoPrune 797 -59.7% 10,569 ms 100.0% (8/8) 0 lines (100% shielded)
Archetype 2: Multi-Module Service ollama Naive Full 1,979 Baseline 28,032 ms 100.0% (8/8) 74 private lines exposed
Archetype 3: Clean Architecture gemini OntoPrune 1,147 -60.1% 0.1 ms 100.0% (10/10) 0 lines (100% shielded)
Archetype 3: Clean Architecture gemini Naive Full 2,873 Baseline 0.1 ms 100.0% (10/10) 91 private lines exposed
Archetype 3: Clean Architecture ollama OntoPrune 952 -60.5% 13,092 ms 100.0% (10/10) 0 lines (100% shielded)
Archetype 3: Clean Architecture ollama Naive Full 2,412 Baseline 34,376 ms 100.0% (10/10) 91 private lines exposed

Key Scientific Findings:

  1. Zero Semantic Degradation (100.0% Pass@1): In all 24 sandbox runs, code generated with OntoPrune stubs passed 100% of unit tests, proving that contract interfaces and ontology metadata provide all the context an LLM needs.
  2. 62% TTFT Speedup on Local CPU: In multi-module and clean architecture projects, time-to-first-token dropped from ~34s to ~13s on consumer CPU.
  3. 100% Intellectual Property Shield: Naive tools (Cursor/Copilot) send entire method bodies (up to 91 private algorithmic lines). OntoPrune sends exactly 0 lines of internal dependencies.
  4. Architectural Dynamics: In monolithic single scripts, compression is moderate (~24%). In structured multi-module systems, compression is consistent at ~60% net token reduction.

💡 Why OntoPrune Makes Local AI-Assisted Programming Truly Feasible

Running local code intelligence models (Qwen 2.5 Coder 7B, Llama 3 8B, DeepSeek Coder via Ollama or llama.cpp) on consumer laptops and developer workstations has historically faced an insurmountable barrier. OntoPrune dismantles this barrier through three architectural pillars:

1. Collapsing the "Pre-fill Latency Wall" on CPU ($TTFT$)

  • The Bottleneck: On CPUs and consumer laptops without high-end 24GB GPUs, token decoding speed is acceptable (15–25 t/s). However, the Pre-fill phase (Prompt Evaluation) is heavily memory-bandwidth bound and saturates RAM.
  • Without OntoPrune: Standard AI assistants inject entire raw files and bloated multi-module contexts (2,500 – 3,000 tokens), causing a 35-second freeze before the first token appears. A half-minute freeze on every code turn shatters the developer's flow state.
  • With OntoPrune: By pruning raw context into concise typed contracts (~800 tokens), pre-fill latency collapses from 34.3s down to 13.0s (a 62% to 74% drop in TTFT) on standard CPUs (AMD Ryzen 5600G). Local inference shifts from unusable to completely interactive.

2. Eliminating Attention Dilution in Small Language Models (SLMs)

  • Small models (3B to 7B parameters) lack the massive context comprehension of hundred-billion-parameter cloud models. Dumping hundreds of lines of irrelevant database or third-party implementations dilutes attention, leading to hallucinated method calls.
  • OntoPrune constructs a mathematically sound closed-world contract. The SLM only sees allowable methods and domain types. In our 24 empirical sandbox runs, Qwen 2.5 Coder 7B achieved 100.0% Pass@1 on pytest suites.

3. Absolute Data Sovereignty and Air-Gapped Privacy

  • OntoPrune operates in $<12\text{ ms}$ on CPU using local AST visitors and in-memory RDF graphs with zero network calls. It provides an airtight guarantee: 0 lines of internal proprietary code are ever leaked to external APIs.

📦 Installation

# Standard installation (native Python support):
pip install ontoprune

# With multi-language support (Flutter/Dart, Java, TypeScript via Tree-sitter):
pip install "ontoprune[languages]"

# For development, benchmarks, and tests:
pip install "ontoprune[dev,benchmark,languages]"

🌐 Universal Multi-Language Support

OntoPrune automatically detects file types and resolves dependencies across project boundaries:

Language Extension AST Engine Output Format
Python .py Native Python ast def name(args) -> Ret: ...
Flutter / Dart .dart tree-sitter-dart abstract class ... { Ret method(); }
Java / Spring Boot .java tree-sitter-java public interface ... { Ret method(); }
TypeScript / React .ts, .tsx, .js tree-sitter-typescript export interface ... { method(): Ret; }

🛠️ Usage Modes

1. Native Model Context Protocol (MCP) Server

OntoPrune runs out of the box as an MCP server (ontoprune-mcp) compatible with Claude Desktop, Cursor, Gemini CLI, or Antigravity IDE:

ontoprune-mcp

Configuration in claude_desktop_config.json:

{
  "mcpServers": {
    "ontoprune": {
      "command": "ontoprune-mcp"
    }
  }
}

Exposed MCP Tools:

  • prune_context(file_path, target_symbol, format='stubs'): Extracts the minimal dependency contract resolving cross-file imports.
  • verify_response(response_code, contract_or_file): Deterministically validates generated code against authorized contracts.

2. Command Line Interface (CLI)

# Prune a target method across multi-file projects:
ontoprune translate src/services/OrderService.java processOrder --format stubs

# Direct streaming pipeline with local Ollama:
ontoprune translate services/order_service.py procesar_orden | ollama run qwen2.5-coder:3b

# Deterministically verify LLM output against contract:
ontoprune check --file generated_solution.py --contract contract.py

3. Python API

import ontoprune

# 1. Prune a multi-module project to a minimal typed contract
context = ontoprune.translate(
    "services/order_service.py",
    target="procesar_orden",
    fmt="stubs",
    multi_module=True,
)
print(context)

# 2. Verify model output against the contract
violations = ontoprune.check(llm_code_response, against=context)
if not violations:
    print("Code is 100% compliant and free of hallucinations!")

🧪 Test Suite

pytest tests/
# 46 passed in 1.40s

🔬 Reproduce Empirical Benchmarks

You can independently replicate the benchmark evaluation in isolated sandboxes on your machine:

# 1. Quick deterministic validation in sandbox (no API key needed):
ontoprune benchmark --backend mock --analyze
# or via Make:
make benchmark-mock

# 2. Live evaluation against Google Gemini (requires GEMINI_API_KEY):
ontoprune benchmark --backend gemini --analyze
# or via Make:
make benchmark-gemini

# 3. Live local evaluation against Ollama CPU (requires running Ollama):
ontoprune benchmark --backend ollama --analyze
# or via Make:
make benchmark-ollama

📄 Whitepapers & Publications


👤 Author

Vigmar Carlo

Metadata

Release files for ontoprune 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for ontoprune 0.3.0
File Size Uploaded
ontoprune-0.3.0.tar.gz 388.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for ontoprune 0.3.0
File Interpreter ABI Platform
ontoprune-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 425.9 kB

Release files / ontoprune-0.3.0.tar.gz

Download URL ontoprune-0.3.0.tar.gz
Size 388.3 kB
Tags Source
SHA-256 checksum
How to use checksums
3fbe2d2eb6d6681c8427602d2a6c31a2fefd58f8ac2c5b821858ef3ad2ae21ef
BLAKE2b-256 checksum
How to use checksums
bfa1c5d1df85ee21a7d79dcc3c57bc7bd172a414139d32115e06fa83e368b6a3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.7 {"installer":{"name":"uv","version":"0.12.7","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Linux Mint","version":"21.1","id":"vera","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / ontoprune-0.3.0-py3-none-any.whl

Download URL ontoprune-0.3.0-py3-none-any.whl
Size 37.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
89236a313797386c78232d4b813026a12eabc746c1f7d7361c89fd32ab259ac4
BLAKE2b-256 checksum
How to use checksums
9467cd663032e770c8006e89868b9bd0bdb3e332e5edc9a220433ebdb436c19b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.7 {"installer":{"name":"uv","version":"0.12.7","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Linux Mint","version":"21.1","id":"vera","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page