Skip to main content

anvil-serving - local model serving and a thin capability gateway

anvil-serving

Benchmark and serve local models through one explicit capability gateway.

License: MIT Source Version Docs

anvil-serving runs and benchmarks local model serves, then exposes their named capabilities through one authenticated endpoint. It is deliberately a thin gateway: a caller chooses a configured model alias, and that alias maps to one local tier. There is no request classifier, quality-profile router, semantic fallback, cloud escalation, or hidden substitute model.

The reference topology has two equivalent RTX PRO 6000 Blackwell Max-Q GPUs. In split mode, compatible LLM, Omni, voice, purpose-model, and ComfyUI workloads reserve Compute A or Compute B independently. In dual-gpu-exclusive mode, one explicitly declared TP=2 serve owns both cards and every other GPU inference workload is offline. Capability aliases remain independent of that placement. The gateway keeps authentication, dialect translation, streaming, readiness, admission, and decision evidence consistent across those capabilities.

Direct capability contract

[router.model_routes]
llm.primary = "primary-local"
llm.voice = "omni-local"
vision.ocr = "omni-local"
vision.general = "omni-local"
vision.video = "primary-local"

Send one of those aliases as the chat model. Matching is case-insensitive after trimming; compatibility prefixes are not accepted. /v1/models advertises the configured aliases plus each alias's declared context_window and max_output_tokens. Unknown or missing chat aliases return 404. An unavailable selected tier returns an exhaustion error, not an alternate model. The authenticated /v1/models/capacity endpoint joins declared model/GPU capacity with bounded live engine telemetry; it does not operate a serve or grant the router GPU-device access. Related authenticated endpoints expose declared capabilities and fingerprints, router build/config identity, bounded-buffer statistics, request traces, and Prometheus gauges. See the router observability API.

Purpose models and audio are equally explicit: embeddings and reranking use their configured model names on dedicated endpoints, while STT/TTS use operator-configured audio routes. ComfyUI is lifecycle-managed rather than a chat capability.

Quick start

Python 3.11+ is the only runtime prerequisite. Docker and a GPU are required only for real local model serves.

pip install -e .
anvil-serving init
anvil-serving serves groups
anvil-serving serves up SERVE_NAME --dry-run
anvil-serving serves up SERVE_NAME --confirm
anvil-serving serves mode status
anvil-serving router run

init writes the packaged operational manifests to ~/.anvil-serving. It detects NVIDIA GPU UUIDs with nvidia-smi, assigns stable Compute A and Compute B roles, and resolves the host's Tailscale IPv4 address. Capacity is sorted largest-first; equal-capacity cards use canonical UUID ordering so runtime-index changes cannot swap the roles. Use --compute-a-gpu-uuid and --compute-b-gpu-uuid to override discovery, --no-detect-host to keep placeholders, --out-dir to choose another location, or --single-model for a focused one-model scaffold. serves up is the canonical bring-up path for models and other manifest-owned resources. Rerunning init leaves content-identical files untouched; only changed files receive numbered backups before replacement. Preview the resolved operation before confirming it.

With the selected serves running, call the gateway:

curl -s http://127.0.0.1:8000/v1/models
curl -s 'http://127.0.0.1:8000/v1/models/capacity?model=llm.primary&images=1&image_tokens=2048'
curl -s http://127.0.0.1:8000/v1/chat/completions \
  -H 'content-type: application/json' \
  -d '{"model":"llm.primary","messages":[{"role":"user","content":"hello"}]}'

Use anvil-serving eval preflight before mapping a real model to an alias, and record capacity and quality evidence with anvil-serving eval benchmark. A mapping is an exposure decision, not a model promotion claim.

What it provides

Surface Purpose
anvil-serving router run Authenticated Anthropic/OpenAI-compatible capability gateway.
anvil-serving serves Compose-backed lifecycle, GPU reservations, and split/exclusive TP=2 mode transactions.
anvil-serving eval preflight Functional qualification of a concrete endpoint.
anvil-serving eval benchmark Capacity and quality evidence collection.
anvil-serving models Model cache, source, and serve-recipe management.
anvil-serving voice Operator-owned STT/TTS, bridge, Realtime, and voice benchmark lifecycle.
anvil-serving mcp serve / controller Structured same-host or private control-plane access.

The reference split-host control plane runs the controller in the dedicated Linux controller image on Fakoli Dark and exposes it through host-owned Tailscale Serve. Fakoli Mini runs only the MCP stdio bridge used by OpenClaw. That bridge bundles the official TypeScript MCP SDK and accepts both the legacy initialize era through 2025-11-25 and the stateless 2026-07-28 era. Its authenticated downstream connection to Dark is pinned to 2026-07-28; the controller itself never exposes a legacy endpoint. Remote MCP proxy mode therefore requires Node.js 20+, while the Python router, controller, and ordinary CLI remain stdlib-only.

Documentation

Security and operating boundaries

  • Treat every tracked file as public. Keep real topology, active promotions, machine paths, and working evidence in a private operator repository selected through ANVIL_SERVING_HOME.
  • Use 127.0.0.1, never localhost, for same-host URLs.
  • Keep router authentication enabled before exposing it beyond loopback.
  • Store credentials only through environment-variable references.
  • Treat readiness and preflight as different checks: readiness says a serve can receive traffic; preflight and benchmark evidence establish whether it should.
  • Fakoli Mini is model-free in the reference topology. Its local audio proxy ports forward to Dark; they do not make Mini a serving host.

See SECURITY.md for the threat model and reporting policy.

Release files for anvil-serving 0.35.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for anvil-serving 0.35.0
File Size Uploaded
anvil_serving-0.35.0.tar.gz 1.5 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for anvil-serving 0.35.0
File Interpreter ABI Platform
anvil_serving-0.35.0-py3-none-any.whl Python 3 none any Details

Total release size: 2.8 MB

Release files / anvil_serving-0.35.0.tar.gz

Download URL anvil_serving-0.35.0.tar.gz
Size 1.5 MB
Tags Source
SHA-256 checksum
How to use checksums
81fa23e8ff8321060dfca974d568269d1ea9fd16dc121904d978a5364cd6a1ca
BLAKE2b-256 checksum
How to use checksums
55426e36338a09d0e8a560c0dc4a1f8e9edbc808b2d5d3725bf0b41cc8c35e84
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 22, 2026.

Transparency log

Release files / anvil_serving-0.35.0-py3-none-any.whl

Download URL anvil_serving-0.35.0-py3-none-any.whl
Size 1.3 MB
Tags Python 3
SHA-256 checksum
How to use checksums
13a331bf0eec6d2a356b1cbd1fea04385e30c8fb3ba0ce209f71c9a28e590b76
BLAKE2b-256 checksum
How to use checksums
0948a20344a5e51f00d5cbb0af4b30f518ccd30fd927119bb200f27712c3eda3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 22, 2026.

Transparency log
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page