Skip to main content

VibeSys: Generating Bespoke Systems with AI Agents

arXiv

An agentic framework that generates bespoke systems from application requirements, workload characteristics, and the underlying hardware.

One of VibeSys's first initiatives is VibeServe, which asks whether AI agents can generate a bespoke LLM serving system for each model, workload, and hardware target. The figures, blog post, and paper below document that initiative.

Generic serving today vs. VibeServe's per-target bespoke systems

Updates

Introduction

VibeSys explores a broader approach to systems development: use application requirements, workload characteristics, and hardware capabilities as the inputs to an agentic search process that creates a purpose-built system. Each target defines its own implementation contract, correctness checks, and performance benchmark, allowing VibeSys to work across domains rather than assuming a single runtime, programming language, or deployment shape.

The framework is organized as a multi-agent optimization loop. An outer loop plans the search over system designs using persistent state such as issues, memory, and git history, while an inner loop implements candidates, validates correctness against target-specific requirements, and measures performance on the target workload and hardware. VibeServe is the first substantial initiative built on this approach; its serving-focused results include predicted-output decoding, hybrid prompt caching, streaming ASR, constrained JSON decoding, multimodal inference, and Apple Silicon deployment.

Architecture

VibeServe architecture: outer loop dispatches per-round tasks to an inner loop of Implementer / Accuracy Judge / Performance Evaluator agents

The framework factors the work along two axes:

  • Outer loop — a fresh designer selects one falsifiable causal hypothesis from git history, profiling evidence, and durable roadmap/progress memory, then hands it off until it is proven, disproven, or otherwise terminated.
  • Inner loop — a hypothesis-scoped implementer session edits the candidate, chooses targeted experiments and parameter ranges, and reports whether to continue or nominate the result.
  • Independent judge — a fresh, read-only reviewer checks the implementation, activation evidence, invariants, and reward-hacking risks at a sparse cadence. After a PASS, the framework—not an agent—runs and records the canonical accuracy and benchmark commands.
  • Performance evaluator — profiles the implementation (Nsight Systems, PyTorch profiler) and feeds bottleneck hints into future design decisions.
  • Skills library — Agent Skills entries distilled from existing serving engines and research literature (continuous batching, paged-KV, FlashInfer/FlashAttention, MLX, hybrid-cache management, …). New model families, hardware platforms, and optimization techniques are added by writing a skill, not by modifying the framework.
  • Execution environment — an isolated workspace that mounts the user-provided artifacts read-only (so the Implementer cannot edit the checker or reference) and exposes the target hardware (local CUDA, Modal, Docker, or Apple Silicon) plus profilers.

Each round is recorded in git and a framework-owned audit. Provisional rounds remain explicitly unreviewed; only judge-approved candidates receive official accuracy and performance results.

Installation

  1. Install Python 3.12+, Git, and uv.
  2. For the default GitHub-synced runs, install the GitHub CLI and sign in with gh auth login. You do not need gh when every run uses --local.
  3. From the repository root, create the local configuration files:
cp .env.example .env       # provider keys for API-backed/deepagents runs
cp agent.toml.example agent.toml

Coding-agent setup

Choose one of these agent options in agent.toml or with command-line flags:

Agent Selection Authentication
Codex CLI --agent-backend cli --cli-provider codex Install Codex and run codex login.
Claude Code --agent-backend cli --cli-provider claude Install Claude Code and use its login flow.
Gemini CLI --agent-backend cli --cli-provider gemini Install Gemini CLI and use its login flow.
OpenCode --agent-backend cli --cli-provider opencode Install OpenCode and configure its provider.
DeepAgents --agent-backend deepagents Put the selected provider's API credentials in .env.

The default agent.toml.example selects the Codex CLI. CLI credentials stay with the CLI; API credentials are loaded from .env automatically.

  1. Check the installation:
uv run vibesys validate examples/data-structures/queue-spsc

uv run creates the Python environment automatically. You do not need to run uv sync first.

To use the interactive TUI, install Node.js 20+, Bun, and pnpm 11 (or enable Corepack). Run uv run vibesys --runs-dir "$PWD/exp_env"; it installs the frontend dependencies and builds the TUI when needed. npm is not required.

Installing from GitHub source

python -m pip install "git+https://github.com/uw-syfi/vibesys.git"

Source installs skip the repository's optional submodules and do not build the bundled native TUI. They run VibeSys headless; pass --headless explicitly to suppress the fallback notice. Use a supported PyPI wheel for the one-command install with the bundled Bun and OpenTUI runtime.

Installing from PyPI

uv tool install vibesys

The published wheel installs the vibesys command and bundles the native Bun and OpenTUI runtime for Linux x86-64, Linux ARM64, macOS Intel, or macOS Apple Silicon. It accepts the same run flags described in docs/cli-flags.md. No separate JavaScript runtime is needed for the installed tool. Every run requires --runs-dir PATH; VibeSys stores experiment directories, synthesized inputs, and shared caches there.

# Interactive TUI
vibesys --runs-dir ~/vibesys-runs --input <bundle> --local --agent-backend cli --cli-provider codex

# Same run, headless (no TUI, no Bun)
vibesys --headless --runs-dir ~/vibesys-runs --input <bundle> --local

vibesys --help                    # full flag list

vibesys also runs headless automatically for validate and when stdin/stdout is not a TTY (pipes, CI). The engine is additionally available directly as python -m vibesys ....

The wheel does not bundle external coding-agent CLIs or their credentials. Install the selected Codex, Claude Code, Gemini, or OpenCode CLI separately and complete its login flow. For API-backed agents, export the provider credentials (for example, OPENAI_API_KEY) or load them through your own configuration. The wheel does not include agent.toml. When --config is omitted, VibeSys loads exactly ./agent.toml from the working directory where it was launched and reports an error if that file is absent. Pass --config /path/to/agent.toml to override that path. Add --local to skip GitHub sync when gh is not installed.

Quickstart

# Issue-tracker outer loop, Codex CLI, Docker on local CUDA, 4 rounds
uv run vibesys \
  --runs-dir ~/vibesys-runs \
  --input examples/model-serving/moonshine-streaming \
  --exp-name my-experiment \
  --docker \
  --agent-backend cli --cli-provider codex \
  --max-rounds 4 \
  --modality speech_to_text

--outer-loop defaults to agent. Pass --outer-loop plain or --outer-loop evolve to switch. See uv run vibesys --outer-loop <kind> --help for loop-specific flags, and docs/cli-flags.md for the supported flag combinations.

Search strategies

VibeSys supports three outer-loop search strategies:

  • agent (default): an orchestrator plans each round and delegates to the implementer, judge, and profiler. It uses the multi-agent inner loop by default; pass --inner-loop single-agent to run the single-agent ablation.
  • evolve: an evolutionary search over candidate implementations.
  • plain: an issue-board loop that drains implementation issues and evaluates performance.

For contributor workflows and extension guides, see docs/development.md.

Per-target inputs

Each evaluation target lives under examples/<name>/:

examples/<name>/
├── OBJECTIVE.md          # free-form deployment goal (model + hardware + workload + interface)
├── vibesys.input.toml  # manifest-declared domain, checker, benchmark, and optional inputs
├── reference/            # optional reference or seed inputs
│   ├── reference.py
│   ├── config.json
│   └── meta.json         # model id + revision
├── accuracy_checker/     # optional checker.py + tests/data — the correctness gate
├── benchmark/            # benchmark.py + load levels — emits the metric to optimize
└── README.md             # human-readable description

Pass the bundle root with --input examples/<name>/. The manifest declares the agent domain, correctness command, benchmark command, optional starter workspace, optional trusted evaluator source, and optional benchmark result metric.

For external usage without a bundle on disk, supply the same pieces as separate --input-objective/--input-objective-file, --input-domain, --input-accuracy-command, and --input-benchmark-command flags (plus optional --input-reference, --input-evaluator-dir, and others). VibeSys synthesizes a bundle and runs it identically. See docs/cli-flags.md for the full flag list.

OBJECTIVE.md is read at the start of every run and must live next to the reference/ directory (sibling, not inside). See examples/model-serving/Llama-3-8B/, examples/model-serving/moonshine-streaming/, examples/model-serving/qwen3-32b-code-edit/, examples/model-serving/olmo-hybrid-prefix-caching/, examples/model-serving/Llama-3.1-8B-Instruct-MLX-8bit/, examples/model-serving/show-o2-1.5B-HQ-h100/, and examples/model-serving/show-o2-1.5B-HQ-macbook/ for the paper scenarios.

For multi-objective evolutionary runs, drop an objectives.toml next to OBJECTIVE.md (or pass --objective name:max|min flags) — see uv run vibesys --outer-loop evolve --help.

Configuration (agent.toml)

[model]
name = "gpt-5.4"             # auto-detected provider for claude-* / gpt-* / gemini-* / gemma-*
# provider = "openai"        # optional override

[backend]
name = "cuda"                 # or "metal", "trainium", "rocm", or "cpu"

[agent]
backend = "cli"               # "cli" (codex/claude/gemini/opencode) or "deepagents"
cli_provider = "codex"        # which coding-agent harness to drive
# cli_timeout = 1800          # per-invocation timeout (seconds)

# Optional role-specific CLI models. Other roles, including the independent
# judge, continue to use [model].name and [thinking].level.
[agent.outer]
model = "gpt-5.6-sol"         # orchestrator pre-round and planning calls
reasoning_effort = "xhigh"

[agent.inner]
model = "gpt-5.6-luna"        # implementer calls
reasoning_effort = "xhigh"

[repository]
# Optional GitHub user/org override. If omitted, use the account from `gh auth status`.
# owner = "your-github-user"
visibility = "private"        # private, public, or internal

[tui]
theme = "dark"                # interactive client theme; --theme overrides

# Optional: benchmark load levels handed to the perf evaluator.
# [[perf_eval.load_levels]]
# rate = 1
# duration = 20
# max_tokens = 128

Provider credentials live in .env — see .env.example. The CLI flags --agent-backend / --cli-provider / --backend override these.

The interactive client ships four light/dark theme pairs — dark (default) / light, solarized-dark / solarized-light, catppuccin-mocha / catppuccin-latte, and high-contrast-dark / high-contrast-light. Set one with [tui].theme or --theme, or switch mid-session with /theme <name>. See docs/cli-flags.md.

The config is validated against a typed schema on load (vibesys/config.py): unknown sections or keys, unknown providers/backends, and missing required fields are rejected with an error rather than silently ignored.

Fresh runs use GitHub-backed tracking by default. Pass --local for a local-only run in the collection selected by --runs-dir; see docs/cli-flags.md for repository and resume options.

Citation

If you use the VibeServe initiative in your research, please cite:

@misc{kamahori2026vibeserveaiagentsbuild,
      title={VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?},
      author={Keisuke Kamahori and Shihang Li and Simon Peter and Baris Kasikci},
      year={2026},
      eprint={2605.06068},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2605.06068},
}

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distributions

No source distribution files available for this release.See tutorial on generating distribution archives.

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

vibesys-0.1.0-py3-none-manylinux_2_28_x86_64.whl (50.6 MB view details)

Uploaded Python 3manylinux: glibc 2.28+ x86-64

vibesys-0.1.0-py3-none-manylinux_2_28_aarch64.whl (50.2 MB view details)

Uploaded Python 3manylinux: glibc 2.28+ ARM64

vibesys-0.1.0-py3-none-macosx_13_0_x86_64.whl (33.5 MB view details)

Uploaded Python 3macOS 13.0+ x86-64

vibesys-0.1.0-py3-none-macosx_13_0_arm64.whl (31.2 MB view details)

Uploaded Python 3macOS 13.0+ ARM64

File details

Details for the file vibesys-0.1.0-py3-none-manylinux_2_28_x86_64.whl.

File metadata

File hashes

Hashes for vibesys-0.1.0-py3-none-manylinux_2_28_x86_64.whl
Algorithm Hash digest
SHA256 678f4a8abd59771d5b384d2ac1bf07bc7f18d52fca35b560717f655e0c69ece8
MD5 9513dfd492e1c5ffbb647d9cac672ddf
BLAKE2b-256 ed9cb89c48049fe691ccb2c3fadf03cf10ef14fb1ea706f649c0f072d87f23e6

See more details on using hashes here.

Provenance

The following attestation bundles were made for vibesys-0.1.0-py3-none-manylinux_2_28_x86_64.whl:

Publisher: publish.yml on uw-syfi/vibesys

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file vibesys-0.1.0-py3-none-manylinux_2_28_aarch64.whl.

File metadata

File hashes

Hashes for vibesys-0.1.0-py3-none-manylinux_2_28_aarch64.whl
Algorithm Hash digest
SHA256 cd6d65a0aac8f7533163a6daac70f29ceb55e86c6db58477a748195a455f4136
MD5 9251bfeef8a2bb53923530811e1970f6
BLAKE2b-256 ef3e3e43a9061485bd8485c8e2efce31cf865246ff3566e7882b054f4b3ccbf1

See more details on using hashes here.

Provenance

The following attestation bundles were made for vibesys-0.1.0-py3-none-manylinux_2_28_aarch64.whl:

Publisher: publish.yml on uw-syfi/vibesys

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file vibesys-0.1.0-py3-none-macosx_13_0_x86_64.whl.

File metadata

File hashes

Hashes for vibesys-0.1.0-py3-none-macosx_13_0_x86_64.whl
Algorithm Hash digest
SHA256 00c110997c31eeb1fa7bf3b9be9c3ce0b5b1475a69018cc13fbf017a192b65f0
MD5 6e2db95bd96b35a5cd1f6a3a4dffd267
BLAKE2b-256 ba6010233c27a7dc3ae3af3df94441b1d1d15a04e0aaf03518c1ddabab2f9b85

See more details on using hashes here.

Provenance

The following attestation bundles were made for vibesys-0.1.0-py3-none-macosx_13_0_x86_64.whl:

Publisher: publish.yml on uw-syfi/vibesys

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file vibesys-0.1.0-py3-none-macosx_13_0_arm64.whl.

File metadata

File hashes

Hashes for vibesys-0.1.0-py3-none-macosx_13_0_arm64.whl
Algorithm Hash digest
SHA256 00ab30d85c79835a098d883229eb33879ec1d4dda7bf47596d7f40ea2c8cfbbe
MD5 58c7af4c253eac6d7ab3c059b5fce4b2
BLAKE2b-256 28c01e6a0fcb53f7134a4d6c9839f079fc8d7b96a83f7bd967be682d87113e45

See more details on using hashes here.

Provenance

The following attestation bundles were made for vibesys-0.1.0-py3-none-macosx_13_0_arm64.whl:

Publisher: publish.yml on uw-syfi/vibesys

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.1.1

4 files

This release

0.1.0 This release

4 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page