Skip to main content

Anvil Serving

Operate, qualify, and expose local AI capabilities through explicit contracts.

License: MIT Source Version Docs

Get started · Choose a product journey · Browse benchmark evidence · Open the documentation

Anvil Serving is one local-first product for Model Serving, the Capability Gateway, Evaluation & Evidence, Anvil Voice, Anvil Media, Anvil Connect, and Control Plane & Fleet operations. The seven families share one package, CLI, topology, safety contract, evidence policy, and release line while retaining explicit authority boundaries. Anvil Voice and Anvil Media are first-class domains inside the umbrella, not separate products.

The result is a reproducible path from model artifact to qualified capability: operate a declared serve, collect independent evidence, make a human-gated exposure decision, and present callers with one exact endpoint. Anvil Serving does not hide fallback, cloud escalation, model substitution, or automatic promotion behind that path.

Product families

Family User outcome Primary commands
Model Serving Pin and operate reproducible model serves. init, models, serves
Capability Gateway Expose exact authenticated capability aliases. router
Evaluation & Evidence Prove compatibility and retain benchmark evidence. eval
Anvil Voice Operate qualified STT, TTS, and realtime voice paths. voice
Anvil Media Run bounded named image/video workflows with durable artifacts. media
Anvil Connect Publish local browser surfaces through one self-hosted, authenticated edge. connect
Control Plane & Fleet Resolve ownership and operate declared hosts and integrations. topology, controller, mcp, fleet, host

The installed product map is read-only and machine-readable:

anvil-serving product families
anvil-serving product journey anvil-media
anvil-serving product journey control-plane-fleet --json

Operating model

The Capability Gateway is an explicit capability meta-router: callers use a stable capability alias, operators map that alias to exactly one tier, and the selected inference service may report bounded facts about the model and context it currently serves. The request path remains deliberately thin. There is no request classifier, quality-profile router, semantic fallback, cloud escalation, or hidden substitute model.

The reference topology has two equivalent RTX PRO 6000 Blackwell Max-Q GPUs. In split mode, compatible LLM, Omni, voice, purpose-model, and ComfyUI workloads reserve Compute A or Compute B independently. In dual-gpu-exclusive mode, one explicitly declared TP=2 serve owns both cards and every other GPU inference workload is offline. Capability aliases remain independent of that placement. The gateway keeps authentication, dialect translation, streaming, readiness, admission, and decision evidence consistent across those capabilities. See Product families and user journeys for each authority boundary and its ordered path.

Anvil Connect exposure contract

Anvil Connect is the browser gateway for the surfaces you build with Anvil Serving. It answers one question: how do you open a dashboard, an agent UI, or a family chat page running on your own workstation from any browser — the same convenience hosted tools give you for their cloud workloads, without port-forwarding an unhardened application, without a VPN per device, and without handing identity for your own services to a hosted vendor. Applications stay bound to loopback and unchanged; exposure is one closed declaration that is validated, rendered, activated, and enrolled. Browsers authenticate once through the OpenID Connect provider you already operate, spoofable identity headers are stripped at the edge, and applications that need a user identity receive signed handoff instead of a raw header.

See Why Anvil Connect for the product story and Anvil Connect for the architecture and component boundaries.

Capability meta-router contract

[router.model_routes]
llm.primary = "primary-local"
llm.voice = "omni-local"
vision.ocr = "omni-local"
vision.general = "omni-local"

Send one of those aliases as the chat model. Matching is case-insensitive after trimming; compatibility prefixes are not accepted. /v1/models advertises the configured aliases plus each alias's effective context_window and max_output_tokens. A tier may keep its served model and context explicitly configured, or opt into bounded metadata reported by its inference service. The alias-to-tier route stays static in both modes. Unknown or missing chat aliases return 404. An unavailable selected tier returns an exhaustion error, not an alternate model. The authenticated /v1/models/capacity endpoint joins declared model/GPU capacity with bounded live engine telemetry; it does not operate a serve or grant the router GPU-device access. Related authenticated endpoints expose declared capabilities and fingerprints, router build/config identity, bounded-buffer statistics, request traces, and Prometheus gauges. See the router observability API.

Purpose models and audio are equally explicit: embeddings and reranking use their configured model names on dedicated endpoints, while STT/TTS use operator-configured audio routes. ComfyUI is lifecycle-managed rather than a chat capability. Its named media workflows may publish bounded, caller-selected quality profiles such as draft, standard, and high; each profile resolves to exact parameters inside the same workflow and never selects another model, host, backend, or provider. Durable media jobs report gateway-observed phase latency, and image artifacts within the six-MiB binary transport bound can return as native MCP image content.

The word meta describes the separation between a stable caller contract and the mutable configuration behind it. It does not mean that Anvil Serving chooses among models. The complete request path is:

  1. The caller chooses a declared capability alias such as llm.primary.
  2. Operator configuration maps that alias to exactly one tier: a direct endpoint or an explicit set of 2–16 qualified equivalent replicas on one host.
  3. The tier supplies configured metadata. A direct tier may instead opt into bounded live metadata from its one inference service; replica tiers may not.
  4. The router checks readiness and admission, selects one eligible member when replicas are configured, and relays once or fails closed. Round-robin and capacity-aware selection never choose a different model, tier, or host.

See qualified replica configuration and member lifecycle limits. The workload activity view adds bounded, read-only node and fleet visibility without acquiring scheduling or lifecycle authority.

See Capability meta-router for the authority model and the product decisions that keep dynamic metadata separate from dynamic route selection.

Quick start

Python 3.11+ is the only Anvil runtime prerequisite. Model serves additionally need their declared engine: Docker and compatible hardware for container recipes, or an installed MLX service on Apple Silicon macOS. A single-host installation can stay entirely on loopback; the reference multi-device deployment also installs Tailscale and follows Private networking with Tailscale.

pip install -e .
anvil-serving product families
anvil-serving init
anvil-serving serves groups
anvil-serving serves up SERVE_NAME --dry-run
anvil-serving serves up SERVE_NAME --confirm
anvil-serving serves mode status
anvil-serving router run

init writes the packaged operational manifests to ~/.anvil-serving. It detects NVIDIA GPU UUIDs with nvidia-smi, assigns stable Compute A and Compute B roles, and resolves the host's Tailscale IPv4 address. Capacity is sorted largest-first; equal-capacity cards use canonical UUID ordering so runtime-index changes cannot swap the roles. Use --compute-a-gpu-uuid and --compute-b-gpu-uuid to override discovery, --no-detect-host to keep placeholders, --out-dir to choose another location, or --single-model for a focused one-model scaffold. serves up is the canonical bring-up path for models and other manifest-owned resources. Rerunning init leaves content-identical files untouched; only changed files receive numbered backups before replacement. Preview the resolved operation before confirming it.

With the selected serves running, call the gateway:

curl -s http://127.0.0.1:8000/v1/models
curl -s 'http://127.0.0.1:8000/v1/models/capacity?model=llm.primary&images=1&image_tokens=2048'
curl -s http://127.0.0.1:8000/v1/chat/completions \
  -H 'content-type: application/json' \
  -d '{"model":"llm.primary","messages":[{"role":"user","content":"hello"}]}'

Use anvil-serving eval preflight before mapping a real model to an alias, and record capacity and quality evidence with anvil-serving eval benchmark. A mapping is an exposure decision, not a model promotion claim.

Portable host services

Use anvil-serving host services to discover supervised services, inspect process and engine state, read logs, and manage their lifecycle through the CLI or typed MCP tools. The supervisor and serving engine are separate contracts:

Platform Implemented serving path
Windows Docker
macOS Native MLX through launchd, or Docker
Linux Docker
NeoCloud: Vast.ai, Runpod Provider lifecycle TBD
Cloud: AWS, Azure Provider lifecycle TBD

The new lifecycle has passed an isolated live macOS LaunchAgent smoke. Docker adapters have simulated supervisor tests; live Docker lifecycle qualification on Windows, macOS, and Linux remains pending.

After declaring the service and its topology owner in the private operator home:

anvil-serving host services discover
anvil-serving host services status
anvil-serving host services logs SERVICE_NAME --tail 100
anvil-serving host services up SERVICE_NAME
anvil-serving host services up SERVICE_NAME --no-dry-run --confirm
anvil-serving host services down SERVICE_NAME --no-dry-run --confirm

Mutations preview by default; applying requires both flags shown above. restart, enable, and disable use the same contract. Existing macOS Parakeet.cpp and Kokoro LaunchAgents can be adopted as legacy services without restarting or migrating them. Native serves, recipes, and voice endpoints can delegate lifecycle to a declared service binding. See Host-supervised services for adoption, installation, ownership, and the distinction between running, enabled, and ready.

What it provides

Surface Purpose
anvil-serving product Read-only family, boundary, and ordered-journey discovery.
anvil-serving router run Authenticated Anthropic/OpenAI-compatible capability meta-router.
anvil-serving serves Docker Compose and declared native-service lifecycle; Docker GPU reservations and split/exclusive TP=2 mode transactions.
anvil-serving host services Supervisor discovery, adoption, state, bounded logs, and portable lifecycle through launchd or Docker.
anvil-serving eval preflight Functional qualification of a concrete endpoint.
anvil-serving eval benchmark Capacity and quality evidence collection.
anvil-serving models Model cache, source, and serve-recipe management.
anvil-serving voice Operator-owned STT/TTS, bridge, Realtime, and voice benchmark lifecycle.
anvil-serving media Named image/video workflows, qualification, durable jobs, cancellation, and opaque artifacts.
anvil-serving mcp serve / controller Structured same-host or private control-plane access.
anvil-serving topology / fleet / host Ownership resolution, fleet parity/drift, and supported host utilities.

The reference split-host control plane runs the controller in a dedicated Linux controller image on the resource-owning inference node and exposes it through host-owned Tailscale Serve. A separate model-free harness node runs only the MCP stdio bridge used by OpenClaw. That bridge bundles the official TypeScript MCP SDK and accepts both the legacy initialize era through 2025-11-25 and the stateless 2026-07-28 era. Its authenticated downstream connection to the resource owner is pinned to 2026-07-28; the controller itself never exposes a legacy endpoint. Remote MCP proxy mode therefore requires Node.js 20+, while the Python router, controller, and ordinary CLI remain stdlib-only.

Documentation

Security and operating boundaries

  • Treat every tracked file as public. Keep real topology, active promotions, machine paths, and working evidence in a private operator repository selected through ANVIL_SERVING_HOME.
  • Use 127.0.0.1, never localhost, for same-host URLs.
  • Prefer Tailscale Serve to project selected loopback services into the tailnet. Use least-privilege grants for network reachability and keep router or controller authentication enabled on every published path.
  • Store credentials only through environment-variable references.
  • Treat readiness and preflight as different checks: readiness says a serve can receive traffic; preflight and benchmark evidence establish whether it should.
  • The dedicated harness node is model-free in the reference topology. A local proxy port that forwards to another owner does not make the harness a model serving host.

See SECURITY.md for the threat model and reporting policy.

Release files for anvil-serving 1.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for anvil-serving 1.2.0
File Size Uploaded
anvil_serving-1.2.0.tar.gz 2.5 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for anvil-serving 1.2.0
File Interpreter ABI Platform
anvil_serving-1.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 4.8 MB

Release files / anvil_serving-1.2.0.tar.gz

Download URL anvil_serving-1.2.0.tar.gz
Size 2.5 MB
Tags Source
SHA-256 checksum
How to use checksums
04a42b84e237a2868407eef7b3a0d39e5a3cda3eeb8cd14ccd05a71e9e16ff3e
BLAKE2b-256 checksum
How to use checksums
1677b78f36e17200c94c78f404506f1066ad704787574f46f0380ec1c396ab72
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 12, 2026.

Transparency log

Release files / anvil_serving-1.2.0-py3-none-any.whl

Download URL anvil_serving-1.2.0-py3-none-any.whl
Size 2.3 MB
Tags Python 3
SHA-256 checksum
How to use checksums
640106504e72ba952f2899a1e9b3fd0401a2b885975359f32956d24490bc2ac9
BLAKE2b-256 checksum
How to use checksums
680cf1b1a1b71c0d6689eb8e08da8b5ec9df8f254f60670b88aa185623a1b060
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 12, 2026.

Transparency log
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page