Skip to main content

🌊 Stream

Ruff Python 3.12+ Docs

Stream is a design space exploration (DSE) and constraint-optimization framework for heterogeneous dataflow accelerators: accelerator systems built by combining cores that each have their own dataflow and performance model (AIE and TPU-like are two example core types among others). Scheduling is layer-fused, and the TETRA constraint optimization uses MILP (Mixed-Integer Linear Programming) to decide tensor placement and transfer paths across the cores of such a system. Stream builds on top of ZigZag for per-core cost estimation.

📖 Explore the Documentation

🚀 Getting Started Guide


✨ Key Features

✔ Heterogeneous dataflow cores: compose an accelerator from cores that each carry their own dataflow and cost model (AIE, TPU-like, pooling, SIMD, and more).

✔ Layer-fused scheduling across the whole system of cores.

✔ TETRA constraint optimization: a MILP (TransferAndTensorAllocator) decides tensor placement and transfer-path routing.

✔ Pluggable solver backends: OR-Tools GSCIP (default, license-free), OR-Tools HiGHS, and Gurobi behind one unified SolverModel API.

✔ ONNX workloads with auto-generated or hand-written mappings.

✔ AMD AIE code generation: emit aie / aiex MLIR for the Ryzen AI NPU, ready for the mlir-aie / IRON toolchain.

✔ Built for AI agents: an MCP server and typed IR models expose the pipeline programmatically.

The pipeline runs as a chain of stages: parse → tile → cost → MILP allocation → memory estimation.


🚀 Installation

Python >=3.12 is required.

Full install with MCP server support (from the repo root):

pip install -e ".[mcp]"

Base install (no MCP server):

pip install -e .

The authoritative dependency source is pyproject.toml (package stream-dse). The base install pulls in zigzag-dse, ortools>=9.15 (the default, license-free MILP backend), pydantic, pydot, and xdsl. Optional extras: [mcp] adds fastmcp (required for the MCP server); [gurobi] adds gurobipy (commercial solver, opt-in).

AIE code generation

AIE-target MLIR codegen and tracing additionally need the AMD AIE toolchain (mlir_aie, llvm-aie, xdsl-aie, snax-mlir, aie-python-extras). These are git/URL installs that PyPI does not allow in package metadata, so a console script installs them after the base install rather than via an extra:

pip install -e .       # or, once published: pip install stream-dse
stream-setup-aie       # installs the AIE toolchain into the current environment

stream-setup-aie --dry-run prints exactly what it will install without making changes.

⚠️ Platform caveat: the AIE toolchain is Linux x86_64 only (manylinux wheels), CPython 3.12 or 3.13.

💡 Solver license note: OR-Tools (ortools_gscip, the default backend) is open-source and needs no license. Gurobi requires the [gurobi] extra (pip install -e ".[gurobi]") plus a separate commercial license; backend="gurobi" errors at solve time without a valid license.

Optional pre-commit setup:

pre-commit install

⚡ Quick Start

Price a small two-Conv workload (a committed test fixture) on a TPU-like quad-core system, with a mapping the generic mapping generator proposes, through the public API; just co-2conv runs the same pair from the test matrix (this repo uses just as a task runner).


🧩 Hardware and Core Types

An accelerator in Stream is described as a system of heterogeneous dataflow cores. Core roles include compute, memory, shim, and offchip; example dataflow core types include AIE, TPU-like, and pooling.

Hardware and mapping files are organized as follows:

  • stream/inputs/examples/hardware/ - system-level hardware YAMLs (e.g. tpu_like_quad_core.yaml, eyeriss_like_*.yaml, simba*.yaml, fusemax.yaml).
  • stream/inputs/examples/hardware/cores/ - per-core-type YAMLs (e.g. tpu_like.yaml, pooling.yaml, simd.yaml, offchip.yaml, eyeriss_like.yaml).
  • stream/inputs/aie/hardware/ and stream/inputs/aie/hardware/cores/ - AMD AIE example core types (e.g. aie_tile.yaml, mem_tile_256KB.yaml, shim_dma.yaml).
  • stream/inputs/examples/mapping/, stream/inputs/aie/mapping/, and stream/inputs/testing/mapping/ - mapping descriptions.

A mapping can be generated (as in Quick Start above) or hand-written and passed as mapping.


📊 Workload × Hardware Matrix

The generic CO pipeline runs any ONNX workload on any of the example hardware systems. The repo ships two small workloads and exercises them across all eight non-AIE example architectures, through the pytest suite (tests/test_hardware_combinations.py).

Workloads - committed test fixtures under stream/inputs/testing/workload/ (weight values are cleared, only tensor shapes matter for cost estimation, so the ONNX stay tiny; just gen-workloads regenerates them via the builders):

  • 2-conv - two chained Conv layers (make_2_conv.py).
  • swiglu - a 5-node SwiGLU block: two Gemms, SiLU, an elementwise Mul, and a down-projection Gemm (make_swiglu.py).
Hardware (stream/inputs/examples/hardware/) Description 2-conv swiglu
eyeriss_like_single_core one Eyeriss-like compute core (+ pooling, SIMD, DRAM) ✓ ✓
eyeriss_like_dual_core two Eyeriss-like compute cores ✓ ✓
eyeriss_like_quad_core four Eyeriss-like compute cores ✓ ✓
tpu_like_quad_core four TPU-like compute cores ✓ ✓
simba_small small Simba chiplet mesh ✓ ✓
simba 36-core Simba chiplet mesh ✓ ✓
fusemax FuseMax array + vector + DRAM ✓ ✓
meta_prototype_dual_core_simd_offchip two Meta-prototype compute cores (+ pooling, SIMD, DRAM) ✓ ✓

✓ = completes through the generic CO pipeline. All combinations run in the default fast suite; on these small single-fusion-group workloads even the 36-core simba mesh finishes in seconds.

Run one combination - hw is any hardware stem from the table (default tpu_like_quad_core):

just co-2conv fusemax           # 2-conv on an architecture
just co-swiglu simba_small      # swiglu on an architecture

Run the whole matrix - the justfile wraps pytest tests/test_hardware_combinations.py, which runs 2-conv + swiglu over all eight architectures plus a parse-only check confirming every hardware definition loads:

just matrix          # parse + 2-conv + swiglu over all 8 architectures (incl. simba)

🐍 Public API

stream/api.py has three calls. Each takes a hardware description (a YAML path or an Accelerator), a workload (anything a registered frontend loads, such as an ONNX path, or a Workload) and an output directory:

  • evaluate_mapping(hardware, workload, output_path, mapping=None, options=None) solves the allocation of each fused group and returns a MappingEstimate. Without a mapping, the mapping generator that claims the hardware proposes one.
  • select_mapping(hardware, workload, output_path, candidates, options=None) returns the estimate of the cheapest of the candidate mappings.
  • generate_code(hardware, workload, output_path, mapping=None, options=None) also writes each fused group's design under output_path/group_<index>, through the code generation backend that claims the hardware.
import tempfile
from stream.api import evaluate_mapping

with tempfile.TemporaryDirectory() as tmp:
    estimate = evaluate_mapping(
        "stream/inputs/examples/hardware/tpu_like_quad_core.yaml",
        "stream/inputs/testing/workload/2conv_1_8_32_32_16_32_3.onnx",
        tmp,
    )
    print("cycles:", estimate.cycles)

A MappingEstimate holds cycles, the fused groups' estimates plus the reconfiguration the hardware declares, the per-group group_cycles, and the solved context, whose useful keys are scheduler, workload, accelerator and group_latencies. SolveOptions sets the solver backend, the number of columns, the constraint selection, the kernel library, the tile search and instrumentation, and its stage_options carries what a plugin's stages read, such as fusion_cut_points and intra_core_tiling for the generic mapping generator or npu and trace_size for the AIE code generator.

Whatever depends on the hardware is found through entry-point groups, so a separate package extends Stream without a fork: stream.frontends (workload formats), stream.mapping_generators (a mapping when none is given), stream.constraints (namespace MILP constraints), stream.core_cost_backends (per-core cost) and stream.codegen_backends (code generation).


🤖 MCP Server (for AI agents)

Stream ships an MCP server (stream/mcp/server.py, server name stream) that lets an AI agent submit and inspect TETRA CO jobs. Requires the [mcp] extra (pip install -e ".[mcp]").

⚠️ Install caveat: [mcp] does not currently resolve against the pinned PyPI xdsl 0.29.1 - fastmcp's dependency tree needs newer typing-extensions/pydantic than xdsl 0.29.1 permits. For now it installs only in the dev environment that uses the git build of xdsl; a clean fix awaits the xdsl upgrade.

Launch command (from the repo root):

python3 -c "from stream.mcp.server import mcp; mcp.run(transport='stdio')"

The server runs on STDIO (JSON-RPC) transport and blocks until the client disconnects.

The 6 tools:

Tool Purpose
run_optimization(hardware, workload, mapping, output_path, backend, ...) Submit a TETRA CO job; returns a job_id immediately; solve runs in the background.
poll_optimization(job_id) Check job status (pending / running / complete / failed / not_found).
get_workload_ir(workload=None, experiment_id=None) Return the workload DAG as WorkloadIR JSON.
get_accelerator_ir(hardware=None, experiment_id=None) Return the hardware model as AcceleratorIR JSON.
get_allocation_ir(job_id) Return the TETRA allocation result as AllocationIR JSON (3 persona views).
get_solve_stats(job_id) Return MILP solve statistics (objective, time, gap, node count, backend).

Run / poll / inspect flow:

  1. run_optimization(...) returns {"job_id": "...", "status": "pending"}.
  2. Poll poll_optimization(job_id) until {"status": "complete"}.
  3. Inspect with get_allocation_ir(job_id) for the AllocationIR (algorithmic / hardware / compiler views) and get_solve_stats(job_id) for solve statistics.

🧠 Working in This Repo (AI agents)

Programmatic / IR API for structured JSON output:

from stream.ir import WorkloadIR, AcceleratorIR, AllocationIR

# ctx = evaluate_mapping(...).context
workload_ir = WorkloadIR.from_internal(ctx.get("workload"))
accelerator_ir = AcceleratorIR.from_internal(ctx.get("accelerator"))
allocation_ir = AllocationIR.from_internal(ctx.get("scheduler"))

workload_data = workload_ir.model_dump()      # JSON-compatible dict
hardware_data = accelerator_ir.model_dump()
allocation_data = allocation_ir.model_dump()

AllocationIR offers .algorithmic_view(), .hardware_view(), and .compiler_view() persona views.


📚 Further Documentation

Metadata

Release files for stream-dse 1.16.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for stream-dse 1.16.0
File Size Uploaded
stream_dse-1.16.0.tar.gz 405.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for stream-dse 1.16.0
File Interpreter ABI Platform
stream_dse-1.16.0-py3-none-any.whl Python 3 none any Details

Total release size: 873.6 kB

Release files / stream_dse-1.16.0.tar.gz

Download URL stream_dse-1.16.0.tar.gz
Size 405.9 kB
Tags Source
SHA-256 checksum
How to use checksums
0743a0ee9caeadb8c040e3e7b83ebff08572a6c6d2d8f1cd3a86bf7ecf8b0b43
BLAKE2b-256 checksum
How to use checksums
6588a2015fc3a6164ae7eb40395b3ea670e8db9b8a2852a670b4cb6388184412
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.3

Release files / stream_dse-1.16.0-py3-none-any.whl

Download URL stream_dse-1.16.0-py3-none-any.whl
Size 467.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
16d544a8fcb0195ed4133e1fd734584939019a3fe2bae0a296c41c82e06f14e2
BLAKE2b-256 checksum
How to use checksums
3d6df7e6d83d6e6ceb0d9d87b60e79fdefd74567473398e0dae67399a7c67cf3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.3
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page