Skip to main content

Chimeraforge

PyPI version Python CI License: MIT

A local-first, model-agnostic LLM deployment planner. It turns "which model, quantization, GPU, and backend -- how many, will it fit, will it hit my SLO, what will it cost" into a fast, honest, measured answer, from your shell, your Python, or your AI assistant.

uvx chimeraforge plan --model-size 8b --hardware "RTX 4090 24GB"

The trust principle

Every number is labeled measured, estimated, or unknown, and the tool refuses to fake the ones it can't stand behind. VRAM and KV-cache are computed from a model's real architecture (exact). Throughput is a measured lookup when available, otherwise an explicit bandwidth-roofline estimate -- never presented as data it isn't. Quality below the bundled corpus reports unknown, not a made-up score. A 0-result plan names the exact gate that rejected every candidate instead of a generic "nothing found." No telemetry, no phone-home, works air-gapped.

Give it a model -- a size class, a Hugging Face repo, an Ollama tag, or manual overrides for an unreleased model -- and it searches the (model x quantization x backend x GPU count x tensor/pipeline parallelism) space against VRAM, quality, latency, cost, energy, and an opt-in safety gate, then hands back the cheapest config that meets your SLO.

11 commands, one tool: plan - suggest - measure - catalog - safety - bench - eval - compare - refit - report - mcp.

The empirical corpus traces to Technical Reports TR108-TR137 (~204,000 real measurements on consumer GPUs). See the CHANGELOG for the full feature history.


Install

Try it with no install:

uvx chimeraforge plan --model-size 8b --hardware "RTX 4090 24GB"
pipx run chimeraforge plan --model-size 8b --hardware "RTX 4090 24GB"

Install for real:

pip install chimeraforge            # planner + model resolution (HF/Ollama) + suggest/measure/safety/bench
pip install chimeraforge[bench]     # + GPU environment metadata for benchmarks (pynvml)
pip install chimeraforge[mcp]       # + MCP server so Claude/GPT/Cursor can call the planner
pip install chimeraforge[eval]      # + quality evaluation (BERTScore, ROUGE-L)
pip install chimeraforge[refit]     # + coefficient refitting (numpy, scipy)
pip install chimeraforge[all]       # everything

Python 3.10+. The core install covers the planner and network-facing commands (httpx is a core dep). plan / suggest / catalog run fully offline; bench / measure / safety need a running backend (Ollama, vLLM, or TGI). Windows / macOS / Linux.

Quickstart

# Plan a registry size class on your GPU
chimeraforge plan --model-size 8b --hardware "RTX 4090 24GB" --request-rate 2.0

# Plan ANY model -- a Hugging Face repo or an Ollama tag
chimeraforge plan --model Qwen/Qwen2.5-7B-Instruct --hardware "RTX 4090 24GB"
chimeraforge plan --model ollama:qwen3:14b --ollama-url http://localhost:11434

# Split a model too big for one GPU across several (tensor parallelism)
chimeraforge plan --model meta-llama/Llama-3.3-70B-Instruct --hardware "H100 80GB" --tp 4

# Shrink the KV-cache, print the cost/latency/quality trade-off menu
chimeraforge plan --model-size 8b --hardware "RTX 4080 12GB" --kv-quant q8 --pareto

# Benchmark a live model and plan on the MEASURED numbers
chimeraforge plan --model qwen3:14b --measure

# Discover + rank what fits your GPU and budget
chimeraforge suggest --source ollama --hardware "RTX 4090 24GB" --budget 500

MCP server -- give Claude / GPT / Cursor the same numbers

GPU sizing is exactly where assistants fail: training-cutoff hardware prices and specs, plus error-prone KV-cache/batching arithmetic done from memory. chimeraforge mcp runs a stdio MCP server so an assistant calls the real planner against measured data instead of guessing.

pip install "chimeraforge[mcp]"

Claude Code:

claude mcp add --transport stdio chimeraforge -- uvx --from "chimeraforge[mcp]" chimeraforge mcp

Claude Desktop / Cursor (add to your MCP config file):

{
  "mcpServers": {
    "chimeraforge": {
      "command": "uvx",
      "args": ["--from", "chimeraforge[mcp]", "chimeraforge", "mcp"]
    }
  }
}

The --from "chimeraforge[mcp]" pulls in the MCP SDK; uvx runs the server in a self-contained environment. If you have already pip install "chimeraforge[mcp]" into the environment your client launches, you can instead use "command": "chimeraforge", "args": ["mcp"].

Exposes three tools: chimeraforge_plan (the full gate search), chimeraforge_resolve_model (grounds a model id in its real params/architecture), and chimeraforge_list_hardware. Every result carries the same measured / estimated / unknown provenance as the CLI, and the tool descriptions tell the model to prefer them over its own knowledge.


Commands

plan -- predictive capacity planner

chimeraforge plan --model-size 8b --hardware "RTX 4090 24GB" --request-rate 2.0
chimeraforge plan --model Qwen/Qwen2.5-7B-Instruct --hardware "RTX 4090 24GB"   # any HF repo
chimeraforge plan --model ollama:qwen3:14b --ollama-url http://localhost:11434  # any Ollama tag
chimeraforge plan --model meta-llama/Llama-3.3-70B-Instruct --hardware "H100 80GB" --tp 4   # multi-GPU
chimeraforge plan --model-size 3b --kv-quant q4 --pareto                       # smaller KV cache, trade-off menu
chimeraforge plan --model-size 3b --workload agent --safety-target 0.85 --json
  • Plans any model: registry size class, HF repo (org/name), Ollama tag, or manual overrides (--params-b/--n-layers/...).
  • Searches (model x quantization x backend x N-replicas x batch/GPU) through a 5-gate pipeline: VRAM -> quality -> safety (opt-in) -> latency -> budget.
  • Models real serving physics: continuous batching (vLLM/TGI), prefill/decode split (TTFT + TPOT), KV-cache-bound concurrency, and variance-aware queueing (--workload).
  • Fits models too big for one GPU: --tensor-parallel/--tp {N|auto} shards weights + KV across N GPUs (Megatron-style, comms-modelled); --pipeline-parallel/--pp {N|auto} splits layers across N stages instead (cheaper on slow interconnects, needs batching to fill the pipeline). Not combinable yet.
  • KV-cache quantization (--kv-quant {fp16,q8,q4}) shrinks the cache and raises max concurrency -- biggest win at long context.
  • Energy (--electricity-rate): monthly kWh cost, $/1M-tok (+energy), and tok/s-per-watt, reported alongside (not folded into) the budget gate.
  • Per-prediction provenance (measured / estimated / unknown); explains the binding gate when nothing fits.
  • Validated on registry data: VRAM R^2=0.968, throughput R^2=0.859, quality RMSE=0.062, latency MAPE=1.05% (beats analytical M/D/1 by 20.4x, TR133). No ML -- empirical lookup tables with first-principles interpolation (roofline for off-registry models).

suggest -- discover & rank models

chimeraforge suggest --source ollama --hardware "RTX 4090 24GB" --budget 500
chimeraforge suggest --source hf --hf-limit 8 --hardware "RTX 4080 12GB"
chimeraforge suggest --source catalog --hardware "RTX 4080 12GB"   # offline, after `catalog --build`

Pulls candidates from a live Ollama (/api/tags), the HF Hub (top text-generation), and/or the local catalog; resolves each to real params/arch, runs the same gate search, and shows the best config per model.

measure -- benchmark live, plan on real numbers

chimeraforge measure --model qwen3:14b --ollama-url http://localhost:11434
chimeraforge plan --model qwen3:14b --measure   # measure then plan in one step

Benchmarks the live model (real N=1 throughput, service time, concurrency scaling) and folds it into a local corpus. plan / suggest then prefer the measured numbers automatically (provenance flips to measured).

catalog -- local model catalog

chimeraforge catalog --build         # resolve a curated seed (+ --with-ollama) and cache specs
chimeraforge catalog                 # list the cached catalog

Persists resolved specs so suggest --source catalog ranks a known-good set fully offline.

safety -- live refusal screen

chimeraforge safety --model llama3.2-3b --prompts harmful.txt --quant Q4_K_M --safety-target 0.85

Where plan --safety-target decides from bundled TR134/TR142 data, safety measures: it runs your probe prompts against a live model, classifies refusals (rule-based -- the TR134 regex baseline), reports the measured refusal rate vs the bundled gate data (expected, drift, RTSI risk tier), and exits 1 below --safety-target. You provide the prompts (--prompts, one per line) -- no attack corpus ships with the package; point it at HarmBench / AdvBench / your own set. Needs a running Ollama.

bench -- live inference benchmarking

chimeraforge bench --model llama3.2-3b --runs 5
chimeraforge bench --model llama3.2-3b --all-quants --context 512,1024,2048,4096 --json
chimeraforge bench --model llama3.2-3b --backend vllm --base-url http://localhost:8000

Three workload profiles (single / batch / server-Poisson); measures throughput, TTFT, and latency with p50/p90/p95/p99; CV-based stability warnings; JSON output.

eval -- quality evaluation

chimeraforge eval --task general_knowledge --json
chimeraforge eval --predictions preds.txt --references refs.txt --model llama3.2-3b

Metrics: exact match, ROUGE-L (LCS fallback), BERTScore, coherence -> composite (0.2*EM + 0.3*ROUGE + 0.3*BERT + 0.2*coherence). Quality tiers from TR125; 3 built-in tasks (general_knowledge, summarization, code). Pass --fp16-baseline to classify the drop tier.

compare -- diff benchmark runs

chimeraforge compare --baseline run1.json --candidate run2.json,run3.json --json

Matches configs by (model, backend, quant, workload, context_length); computes throughput/TTFT/duration deltas with an aggregate improvement/regression summary.

refit -- update planner coefficients

chimeraforge refit --bench-dir ./results/ --output fitted_models.json --validate

Bayesian blending (per-key confidence weighting), hardware offsets, power-law refitting, and a 10-check validation suite that gates the write (--validate).

report -- generate reports

chimeraforge report --results-dir ./results/ --format markdown --output report.md

Markdown (GitHub-compatible) and self-contained, XSS-safe HTML; statistical analysis (RMSE, MAE, MAPE, R^2) with per-config percentile tables.

mcp -- serve the planner to AI assistants

chimeraforge mcp

Runs the stdio MCP server described above. Requires pip install "chimeraforge[mcp]".


What's modeled

Dimension How it's computed Provenance
VRAM / KV-cache First-principles from real model architecture; KV-quant and TP/PP-aware sharding exact
Max concurrency KV-cache-bound sequences per GPU exact
Throughput (decode) Measured lookup, else bandwidth roofline; continuous-batching curve; TP comms / PP bubble measured / estimated
TTFT (prefill) Compute-bound, GPU FP16 TFLOPS x MFU estimated
Quality Measured composite lookup, family-prior estimate, or unknown measured / estimated / unknown
Cost GPU $/hr x fleet size ($/1M-tok invariant in replica count) exact
Energy TDP-driven monthly kWh, $/1M-tok (+energy), tok/s-per-watt estimated
Safety TR134/TR142 refusal-rate lookup (opt-in gate) measured / unknown

Hardware: 22 GPUs -- consumer Ada + Blackwell (RTX 30/40/50-series), datacenter (A100 40/80GB, H100, H200, B200, L4, T4), and AMD MI300X -- each with VRAM, bandwidth, FP16 TFLOPS, TDP, and interconnect (NVLink/Infinity Fabric/PCIe).

Known limits (honest): MoE active-vs-total-param divergence, reasoning/thinking tokens, speculative decoding, and prefix caching are not yet modeled. Quant coverage for vLLM/TGI is GGUF-only (no FP8/AWQ/GPTQ yet). TP and PP throughput are comms-modelled estimates, not measured, and can't be combined in one plan. Queueing is analytical (variance-aware), not a discrete-event simulator. The bundled corpus is fit primarily on one rig (RTX 4080 12GB); other GPUs scale from bandwidth/compute until you measure on yours. The MCP server is stdio-only (Claude Code/Desktop, local Cursor) -- no hosted remote transport yet.


What the research decided

Phase 2 (TR123-TR133, ~106,000 measurements) distilled into an artifact-backed deployment framework -- the same rules the planner applies:

Decision Recommendation Evidence
Single-agent backend Ollama Q4_K_M Highest throughput/dollar; quality within -4.1pp (TR123-TR125)
Multi-agent backend (N>=4) vLLM FP16 2.25x advantage from continuous batching (TR130-TR132)
Compile policy Prefill only, Linux, Inductor+Triton 24-60% speedup; decode crashes 100% (TR126)
Quantization Q4_K_M default; Q8_0 quality-critical; never Q2_K Universal sweet spot across 5 models (TR125)
Context budget Ollama for >4K tokens on 12 GB VRAM spillover = 25-105x cliffs (TR127)
Capacity planning chimeraforge plan Validated R^2>=0.859; beats M/D/1 by 20.4x (TR133)
Safety screening plan --safety-target (opt-in) Refusal-rate + RTSI risk per config; rejects safety-collapsing cells (TR134/TR142)

Headline findings (full data in the TRs): Rust beats Python single-agent (+15.2% throughput, -58% TTFT, -67% memory -- TR112); dual Ollama reaches near-perfect multi-agent parallelism (~99%) vs 82.2% on one instance (TR110/TR113/TR114); vLLM's continuous batching gives a 2.25x edge at N=8, bottlenecked on GPU memory bandwidth, not the stack (TR130-TR132).

Full research: docs/archive/technical_reports.md indexes all 32 reports; the full archive with methodology and raw-data references lives in outputs/publish_ready/reports/.


How the numbers are made

  • ~204,000 primary measurements across 32 technical reports (TR108-TR137 + the TR142/TR146 safety provenance), on an RTX 4080 Laptop (12 GB). De-duplicated: TR137/TR142 are syntheses of already-counted data.
  • Rigor: fresh-process isolation per run (no warm-cache bias), forced cold starts, 3-5 runs per config for statistical confidence, structured JSON/CSV logging with full provenance. Every claim traces to raw data you can re-run.
  • Program context: ChimeraForge is the actionable CLI splice of the parent Banterhearts program (~1,337,000 primary + judge measurements across 54 TRs); the safety attack-surface and serving-stack research lives in sibling repos.
  • 549 automated tests (pytest tests/) cover the planner models, gate search, resolver, discovery, safety, bench backends, and the MCP server -- GPU-decoupled, no live backend required for the core suite.

Reproduce any number: find the claim in a report under outputs/publish_ready/reports/, follow its reference to the data folder, inspect the CSV/JSON, and re-run the provided scripts or notebooks. See docs/archive/methodology.md.

Repository layout

Path Contents
src/chimeraforge/ The chimeraforge CLI + capacity planner (the pip package)
src/python/banterhearts/ Python agent benchmarking, monitoring, profiling
src/rust/ Rust single- and multi-agent implementations (Tokio + 4 alt runtimes)
outputs/publish_ready/reports/ Canonical TR archive (TR108-TR137) + syntheses -- start here for findings
docs/ Guides, API reference, and the technical-report index -- start here for how-to
experiments/, data/, benchmarks/ Reproduction scaffold, baselines, and raw benchmark artifacts

Documentation

Contributing

Contributions welcome -- see CONTRIBUTING.md. Good areas: additional benchmark configs, new optimization strategies, more models/hardware, docs, and analysis tools.

License

MIT -- see LICENSE.

Acknowledgments

Conducted as part of the Banterhearts LLM Performance Research Program: Phase 1 (TR108-TR122) established the measurement methodology and cross-language comparison, Phase 2 (TR123-TR133) produced the deployment framework and capacity planner, and Phase 3 (TR134-TR137) measured the safety cost of inference optimization -- now the planner's opt-in safety gate.


Repository: https://github.com/Sahil170595/Chimeraforge - PyPI: https://pypi.org/project/chimeraforge/ - Status: Beta, actively developed

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

chimeraforge-0.12.2.tar.gz (160.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

chimeraforge-0.12.2-py3-none-any.whl (133.7 kB view details)

Uploaded Python 3

File details

Details for the file chimeraforge-0.12.2.tar.gz.

File metadata

  • Download URL: chimeraforge-0.12.2.tar.gz
  • Upload date:
  • Size: 160.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for chimeraforge-0.12.2.tar.gz
Algorithm Hash digest
SHA256 9e779d96fd16d4306b511f7962cee4b6ad1a5e8c2c4664b2a9909a22c8e54b6f
MD5 50ef52cd1d3ea2a53edcab6e61e96e78
BLAKE2b-256 1699e30f42fbcc0656cf2069e06b15e5f564bc392924855563c88514dea2f0c1

See more details on using hashes here.

Provenance

The following attestation bundles were made for chimeraforge-0.12.2.tar.gz:

Publisher: publish.yml on Sahil170595/Chimeraforge

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file chimeraforge-0.12.2-py3-none-any.whl.

File metadata

  • Download URL: chimeraforge-0.12.2-py3-none-any.whl
  • Upload date:
  • Size: 133.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for chimeraforge-0.12.2-py3-none-any.whl
Algorithm Hash digest
SHA256 c4068e3a8157065119b1e1dfe12ed34c2c5a49734063b32406504fc2178a000a
MD5 e801d5ba360f45dfcd4c97c0c8652ad1
BLAKE2b-256 72a61f0a47ad96a3d77b7cb1fd8bf4b11f07a59792442f38c90a7a72585e8f25

See more details on using hashes here.

Provenance

The following attestation bundles were made for chimeraforge-0.12.2-py3-none-any.whl:

Publisher: publish.yml on Sahil170595/Chimeraforge

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page