Skip to main content

Prism

License: MIT Python 3.11+ CI Status: Alpha

Turn local models and cloud LLMs into one AI API.

A glass prism splitting white light into a rainbow spectrum

Why Prism · How it works · Quick start · Documentation · Contributing

Prism lets several models work together behind one OpenAI-compatible endpoint. Connect your application once, then choose a profile that combines the models you need: a small model to prepare context, a vision model to read an image, or a larger model to review and finish the answer. Your application receives one assistant response.

Prism is a component of MirrorNeuron, built to make AI workflows useful on infrastructure you control. It helps solve local AI by composing available models into a service your applications can use. You can also run Prism independently: combine local servers with OpenRouter, OpenAI, Claude, Gemini, and other APIs supported by LiteLLM, using local models, cloud models, or a mix of both.

Distribution: mirrorneuron-prism · import: prism · command: prism · Python 3.11+ · MIT

Why Prism

AI applications need different capabilities for different jobs. A coding assistant may benefit from drafting and review; a support tool needs structured answers; an image workflow needs vision before reasoning. Prism gives you a place to compose those capabilities while your application keeps the same API.

  • Keep your application simple. Use your existing OpenAI client and switch workflows by model alias. Define physical models once and reuse them across profiles.
  • Put each model to useful work. Let a small or free model prepare context, a vision model interpret pixels, and a selected final model write the answer. Measure cost, latency, and task quality to find a combination that fits your workload.
  • Bring your own compute and providers. Start with local inference, cloud APIs, or both. Run Prism as a Python package or Docker service; MirrorNeuron is optional for standalone use.
  • Make model choices visible. Test actual JSON, image, and reasoning behavior, route structured output to a capable final model, and inspect stage usage in traces. Set limits on calls, context, output, concurrency, and deadlines.

What Prism does

Prism is a model-composition proxy that serves the OpenAI Chat Completions API. A profile defines the physical models and workflow behind a public alias. You can choose a direct call, vision → text/reasoning, draft → review → synthesis, plain-text preparation → final answer, or source-backed evidence extraction → synthesis.

For example, prism-balanced uses free Nano for ordinary text and Super when JSON is required. prism-vision-reasoning uses Nano to inspect an image, then Super to answer from its observations. Your client selects the alias; Prism executes the configured stages and returns the result through the same endpoint.

Prism is alpha software. Observation-based preparation can lose information; source-backed evidence policies validate quote provenance, which does not prove correctness or recall. Qualify models on your own workload.

How it works

One call from your application. One or more model calls inside Prism. One response back. Prism acts as a transparent proxy at the Chat Completions interface: your client selects a public model alias while Prism handles the configured stages. Models can run locally, behind cloud APIs, or across both.

flowchart TB
    A["Your application<br/>One Chat Completions request"]
    subgraph P["Prism · transparent model proxy"]
        G["Check capabilities, context, and budgets"]
        L{"Select a feasible policy<br/>Laya classification or a fixed profile"}
        D["Direct model<br/>1 model call"]
        W["Small or vision model<br/>Prepare observations"]
        F["Final model<br/>Synthesize the answer"]
        B["Draft model"]
        R["Review model"]
        S["Final model"]
        V["Validate final output<br/>Record a metadata trace"]
        G --> L
        L -->|Direct| D
        L -->|Prepare + synthesize: 2 calls| W
        L -->|Draft + review + synthesize: 3 calls| B
        W --> F
        B --> R --> S
        D --> V
        F --> V
        S --> V
    end
    A --> G
    V --> O["One assistant response<br/>Same public API and model alias"]

The diagram shows three representative paths. Source-backed evidence policies can fan out across multiple workers, then synthesize their results. Every path stays within the profile's declared limits, and traces expose which physical models ran. The proxy keeps the client interface consistent; the answer, latency, and cost depend on the selected workflow.

Request classification with Laya

Prism uses Laya, a local typed-decision engine, to classify requests for routing and choose among eligible execution policies. The default checkpoint is convaiinnovations/laya-typed-decisions, prepared on CPU at server startup. Laya receives a bounded instruction sample and plan metadata; source documents remain outside its decision input.

Prism checks feasible plans before asking Laya to choose. A confident, valid choice selects an existing policy; abstention, low confidence, or decision-inference failure uses the feasible rules fallback. Fixed profiles and requests with a single eligible policy skip decision inference. Laya's confidence is a routing signal, not an answer-quality guarantee. See execution policies for the full decision flow.

Choose your cost–quality tradeoff

Workflow Cost and latency Quality consideration
Direct One model call, without preparation/review overhead The chosen model handles the original request itself
Small/free preparation → final model Can reduce premium input when source context is condensed; adds a worker call Notes can omit facts or qualifications; validate task results
Vision preparation → final model Adds image interpretation before text/reasoning synthesis Enables a text-only final model to use images through potentially lossy observations
Draft → review → synthesis Adds drafting and review calls Review can catch problems, but improvements need workload evidence

Use profiles to decide where to spend model work, then benchmark the result against a direct baseline. Cheaper input, stronger final models, and extra review each change the tradeoff; none establishes equivalent quality by itself. Optional cost/power selection ranks feasible assignments using configured prices and operator ratings. The live pilot shows measured savings alongside quality and latency regressions.

Start with free OpenRouter models

Try the included free-model profiles with an OpenRouter key. Install from this checkout now; use python -m pip install mirrorneuron-prism after publication.

python -m pip install .
mkdir prism-demo && cd prism-demo
prism init --preset openrouter
export OPENROUTER_API_KEY='your-openrouter-key'
export PRISM_API_KEY='your-prism-client-secret'
prism validate
prism profiles
prism serve

First startup may download the required Laya checkpoint and prepares it on CPU. The server defaults to http://127.0.0.1:8080. The OpenRouter preset exclusively uses free Nano Omni, Super, and Ultra models; upstream credentials and free-tier quotas still apply.

curl --fail-with-body http://127.0.0.1:8080/v1/chat/completions \
  -H "Authorization: Bearer $PRISM_API_KEY" -H 'Content-Type: application/json' \
  -d '{"model":"prism-balanced","messages":[{"role":"user","content":"Explain decorators in Python briefly."}],"max_completion_tokens":4096}'

For explicit anonymous serving, use prism serve --no-auth. Upstream credentials remain required. Clients share anonymous trace access in this mode; use it on a trusted network.

Pick a workflow

Free-preset alias Behavior
prism-balanced Nano text; Super when JSON is required
prism-vision-llm / prism-omni-llm Nano interprets images, Ultra writes the answer
prism-vision-reasoning Nano interprets images, Super writes the answer
prism-reasoning-image / prism-reasoning-omni Super drafts/reviews, Nano finishes; Super finishes JSON
prism-llm-omni Ultra drafts, Super reviews, Nano finishes; Super finishes JSON
prism-nano-synthesis Nano prepares text context, Super finishes
prism / prism-evidence Automatic or fixed source-backed evidence routing

Vision means image understanding with text output. Additional omni modalities and image generation are not implemented. Nano's direct alias deliberately rejects JSON requirements; profiles use an explicit structured_output_model instead.

For an image plus JSON output, choose a vision-synthesis alias such as prism-vision-reasoning so pixels reach Nano and structured final output comes from Super.

Free-model setup and live results · Native OpenAI / Claude / Gemini examples

Keep your OpenAI client

import os
from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8080/v1", api_key=os.environ["PRISM_API_KEY"])
answer = client.chat.completions.create(
    model="prism-balanced",
    messages=[{"role": "user", "content": "Return JSON with ok=true."}],
    response_format={"type": "json_object"},
    max_completion_tokens=4096,
)
print(answer.choices[0].message.content)

Chat Completions supports ordinary text, image inputs, tools on direct routes, and SSE. Multi-stage and JSON-constrained streams are delivered after validation. Responses API is not implemented. For a guaranteed requested shape, use JSON Schema and retain task-level checks.

Inspect and qualify

prism --help
prism models
prism profiles --json > profiles.json
prism doctor --probe-backends
prism capacity --model prism-vision-reasoning
prism trace show prism-REQUEST_ID

Terminal output uses tables; redirected results are JSON. Use --json, --output table, or NO_COLOR=1 explicitly. Inventory capabilities are declarations; capacity results come from live challenges and distinguish failures from inconclusive upstream errors.

The live six-task pilot measured coding, copywriting, support, and summaries. Nine matched completed pairs showed 56.3% lower hypothetical premium token cost, using free Nemotron tokens priced like GPT-6 Astra / Claude Opus. Across all attempts, acceptance fell 9/12 → 6/12 and mean latency rose 4.83s → 19.96s. These are workload tradeoffs, not a production savings or frontier-model quality claim. Qualified marketing narrative.

Run with Docker

docker build -t mirrorneuron-prism:0.3.0 .
docker run --rm -p 127.0.0.1:8080:8080 \
  -e OPENROUTER_API_KEY -e PRISM_API_KEY \
  -v prism-huggingface:/home/prism/.cache/huggingface \
  mirrorneuron-prism:0.3.0

Or run docker compose up --build. Custom config, checkpoint caching, and anonymous Docker serving.

Documentation and development

Guide Covers
Usage Installation, auth, CLI, JSON, Docker, local models, streaming, and traces
Model configuration and capacity LiteLLM transports, shared registries, and live probes
Execution policies Stage graphs, limits, and Laya routing
Copyable curl examples Requests exercised by integration tests
Benchmarking Reproducible runs and metric limitations
Release instructions Build, validate, install, and manually publish
Implemented contract Guarantees and boundaries
python -m pip install '.[dev]'
ruff check src tests examples
python -m pytest --cov=prism --cov-report=term-missing -q
python -m pytest tests/standalone -q -m integration -o addopts=''
python -m build
python -m twine check dist/mirrorneuron_prism-0.3.0*

CI exercises Python 3.11–3.13. Live provider qualification is separate from deterministic tests. The release workflow is manual; the prepared distribution ships bundled presets, benchmark fixtures, typing metadata, and the MIT license.

Contributing and support

Bug reports, documentation improvements, provider fixtures, and focused pull requests are welcome. Read the contribution guide for setup, validation, and review expectations. For bugs or feature requests, open an issue with a minimal reproducible example. See the security policy for reporting vulnerabilities privately.

License

Prism is open source under the MIT License. Copyright © 2026 mirrorneuron-prism.

Metadata

Release files for mirrorneuron-prism 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for mirrorneuron-prism 0.3.0
File Size Uploaded
mirrorneuron_prism-0.3.0.tar.gz 308.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for mirrorneuron-prism 0.3.0
File Interpreter ABI Platform
mirrorneuron_prism-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 418.6 kB

Release files / mirrorneuron_prism-0.3.0.tar.gz

Download URL mirrorneuron_prism-0.3.0.tar.gz
Size 308.0 kB
Tags Source
SHA-256 checksum
How to use checksums
ea80f5d5d023ce847d2409c9f45705dd706f18efb77276dd5426874abda76074
BLAKE2b-256 checksum
How to use checksums
6904736e8dfe262371b90fd109b690caf05ba9958d50ff563ce279dd34694a8b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release files / mirrorneuron_prism-0.3.0-py3-none-any.whl

Download URL mirrorneuron_prism-0.3.0-py3-none-any.whl
Size 110.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
78346dab83db29896b7f59f419c87faee36fa6ec44295c2f89142ac36be71c40
BLAKE2b-256 checksum
How to use checksums
3e78d1f6b7df6a52973b885b13cfea55ac685f19d8260a830a75375a4babfde9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release history Release notifications | RSS feed

0.3.1

2 release files

This release

0.3.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page