Skip to main content

Code Oracle

Sub-50ms neuro-symbolic verification for AI coding agents.
Structural AST topology validated by Tarjan SCC and Tyranid-BERT (164M INT8 decision model).

License Inference Latency Token Waste Paradigm PyPI version GitHub Release


The Problem: Autoregressive Verification Overhead

When coding agents (Claude Code, Cursor, OpenCode, Codex) inspect code modifications, they face three operational bottlenecks:

  1. Confirmation Bias: Generative models reviewing their own diffs frequently rationalize their own logic errors.
  2. Latency and Token Overhead: Re-evaluating complete files through a frontier model introduces 2,000 to 6,000 ms of round-trip latency and consumes output tokens on conversational explanations.
  3. Tool Schema Bloat: Standard MCP servers inject sprawling multi-tool schemas into the prompt context on every turn, reducing effective agent window capacity.

Architecture

Code Oracle decouples code generation from verification. It runs locally as an independent evaluation layer.

Instead of generating conversational critiques, Code Oracle constructs a localized AST graph using Tree-sitter, checks topological invariants symbolically, and evaluates residual drift using Tyranid-BERT (a 164M parameter ModernBERT multi-task decision model quantized to INT8).

Verification Pipeline

flowchart TD
    Agent["🤖 AI Coding Agent / Developer<br/>(Claude Code, Cursor, Antigravity)"]

    subgraph Engine ["⚡ CODE ORACLE ENGINE (&lt; 20-50ms)"]
        direction TD

        Stage1["Stage 1: Tree-sitter &amp; k-Hop TopoSlice<br/>• Multi-language AST parsing (&lt; 8ms)<br/>• Extracts callers, callees &amp; interfaces<br/>• Isolates k-hop neighborhood graph"]

        Stage2{"Stage 2: Deterministic Symbolic Gate<br/>Tarjan's SCC &amp; Contract Invariants"}

        HardVeto["🚫 Hard Veto Early Exit (&lt; 25ms)<br/>Instant rejection on cycles &amp; signature drift"]

        Stage3["🧠 Stage 3: Tyranid-BERT 164M INT8 Head<br/>• Evaluates linearized Micro-DSL subgraph<br/>• Sub-20ms ONNX Runtime CPU inference<br/>• Multi-Task: Risk Regression, 5-Class Taxonomy &amp; Uncertainty"]

        Stage1 --> Stage2
        Stage2 -- "Cycle / Invariant Breach" --> HardVeto
        Stage2 -- "Topologically Valid" --> Stage3
    end

    Agent -->|"(1) Proposes Patch / Refactor"| Stage1
    HardVeto -->|"(2) Fast-Fail Verdict"| Verdict["🎯 Structured Typed Verdict<br/>VERDICT: APPROVED / REJECTED<br/>Risk Score &amp; Invariant Telemetry"]
    Stage3 -->|"(2) Calibrated Verdict"| Verdict

Key Characteristics

  • Sub-50ms Target Latency: In-memory execution using ONNX Runtime or MLX on local CPU and Apple Silicon.
  • Zero Output Token Tax: Emits structured status codes and calibrated probability vectors (Pass, Fail, Risk Score) without text generation.
  • Lean Tool Surface: Exposes a single endpoint (verify_patch), avoiding multi-tool schema overhead in agent context.
  • Offline Execution: Runs without external API calls or network egress.

Architectural Targets vs. Frontier LLM Review

Metric Autoregressive LLM Code Review Code Oracle (Neuro-Symbolic)
Response Latency 2,500 ms to 6,500 ms < 50 ms (Local In-Memory)
Output Token Cost 150 to 500 tokens / check 0 tokens
Monetary Cost $0.003 to $0.02 / call $0.00 (Local / Offline)
Verification Method Probabilistic text generation Deterministic AST + Calibrated Score
Context Consumption Multi-KB schema injection Single-tool lean schema (< 100 tokens)

Operational Scope and Boundaries

Code Oracle operates within explicit technical boundaries:

  1. Evaluator, Not Author: Code Oracle does not generate, autocomplete, or refactor code. It evaluates proposed patches against existing syntax and topology.
  2. Syntax Requirement: Patches must produce a valid Tree-sitter AST. Syntactically invalid inputs fail at Stage 1 before invoking the decision model.
  3. Static Topology Bounds: Focuses on structural invariants, dependency cycles, and interface compatibility. It does not replace dynamic test suites, integration environments, or runtime race condition detectors.
  4. Memory Footprint: Requires only ~150 MB of RAM for the INT8 quantized ONNX model, running on commodity CPUs with standalone onnxruntime (Zero-PyTorch dependency).

Pretrained Model Weights

The fine-tuned Tyranid-BERT (164M INT8) multi-task decision head weights are hosted on Hugging Face:
🤗 wxsys/tyranid-bert

Code Oracle automatically downloads and caches these weights to ~/.cache/code_oracle/weights/ on first invocation when --neural is enabled, or reads from local ./weights_base/ if present.


Installation & Quickstart

1. Install via pip

# Core AST Symbolic Verification Engine (< 8ms, zero neural footprint)
pip install code-oracle

# With Tyranid-BERT ONNX Runtime decision model (< 20ms INT8 CPU)
pip install "code-oracle[neural]"

# Full developer setup with test suites & packaging tools
pip install "code-oracle[all]"

2. FastMCP Server for Coding Agents

Run Code Oracle as an MCP sidecar for Claude Code, Cursor, or Antigravity:

code-oracle serve

3. CLI Verification & Analysis

# Verify proposed patch against current repository state
code-oracle verify --patch /path/to/patch.diff

# Incremental workspace symbol indexing
code-oracle index .

# Dead code & orphan symbol scan (0 in-degree reachability)
code-oracle dead-code .

# Static performance anti-pattern & resource leak audit
code-oracle perf-lint .

# Install Git pre-commit verification hook
code-oracle hook install

Roadmap

  • Architecture Specification & Subgraph Slicing Design
  • TopoSlice AST Slicer & Incremental Workspace Indexer
  • Tarjan's SCC Cycle Detector & Deterministic Symbolic Gate
  • Multi-Task Risk Taxonomy (5 Classes) & Epistemic Uncertainty Estimation (ADR-0003)
  • Embedded Dead Code Semantics Classifier with ModernBERT Representations
  • Standalone ONNX Runtime Inference & Dynamic INT8 Quantization (code-oracle export-onnx)
  • Tyranid-BERT Official Model Release (wxsys/tyranid-bert)
  • Golden Hybrid v3 Multi-Language Dataset (~4,900 balanced samples across Go, Python, TypeScript, Rust)
  • Lean FastMCP Server interface (verify_patch)
  • Agentic SKILL.md distribution for Claude Code, Cursor, and Antigravity
  • Git pre-commit & pre-push verification hook with unblock toggle (code-oracle hook)
  • Multi-language AST extractors for Tier 1 languages (Python, TypeScript, Go, Rust)
  • Dead Code & Orphan Symbol Scanner (code-oracle dead-code via 0-in-degree graph reachability)
  • Static Performance Anti-Patterns & Resource Leak Detector (code-oracle perf-lint: nested loop complexity, unclosed handles)
  • Official Git Tagging & GitHub Release pipeline (v0.1.0)
  • Python Package Wheel Distribution & PyPI Publishing (pip install code-oracle)
  • Multi-Agent Ecosystem Integrations (Claude Code, Cursor, Antigravity, OpenCode, and Cline sidecars)

License & Attribution

Distributed under the Apache-2.0 License. See LICENSE for details.

Architect & Maintainer:
Wahyu Febri Tamtomo (@wahyuzero)
Founder of frugaldev.biz.id (Radical AI Efficiency & Frugal Computing).

Metadata

Release files for code-oracle 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for code-oracle 0.1.0
File Size Uploaded
code_oracle-0.1.0.tar.gz 3.7 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for code-oracle 0.1.0
File Interpreter ABI Platform
code_oracle-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 3.8 MB

Release files / code_oracle-0.1.0.tar.gz

Download URL code_oracle-0.1.0.tar.gz
Size 3.7 MB
Tags Source
SHA-256 checksum
How to use checksums
db7752da09c2484356faca08b668b5c9683987e76ded04c5e8624a35b2d29ef2
BLAKE2b-256 checksum
How to use checksums
36af664de33624d210d80e66b2dfce42e81c9e89487c0f9ea020d108960536a6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 26, 2026.

Transparency log

Release files / code_oracle-0.1.0-py3-none-any.whl

Download URL code_oracle-0.1.0-py3-none-any.whl
Size 145.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
ad3b73239acb6c0e9ec014539996e11064610ad0b436d4d72f8adfb98ead1548
BLAKE2b-256 checksum
How to use checksums
e60ae6ebd72eaab45a19d77ec2f8a34858705cc500f0ffdfa3c0d57767e89fb5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 26, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page