🌊 Stream
Stream is a design space exploration (DSE) and constraint-optimization framework for heterogeneous dataflow accelerators: accelerator systems built by combining cores that each have their own dataflow and performance model (AIE and TPU-like are two example core types among others). Scheduling is layer-fused, and the TETRA constraint optimization uses MILP (Mixed-Integer Linear Programming) to decide tensor placement and transfer paths across the cores of such a system. Stream builds on top of ZigZag for per-core cost estimation.
📖 Explore the Documentation
🚀 Getting Started Guide
✨ Key Features
✔ Heterogeneous dataflow cores: compose an accelerator from cores that each carry their own dataflow and cost model (AIE, TPU-like, pooling, SIMD, and more).
✔ Layer-fused scheduling across the whole system of cores.
✔ TETRA constraint optimization: a MILP (TransferAndTensorAllocator) decides tensor placement and transfer-path routing.
✔ Pluggable solver backends: OR-Tools GSCIP (default, license-free), OR-Tools HiGHS, and Gurobi behind one unified SolverModel API.
✔ ONNX workloads with auto-generated or hand-written mappings.
✔ AMD AIE code generation: emit aie / aiex MLIR for the Ryzen AI NPU, ready for the mlir-aie / IRON toolchain.
✔ Built for AI agents: an MCP server and typed IR models expose the pipeline programmatically.
The pipeline runs as a chain of stages: parse → tile → cost → MILP allocation → memory estimation.
🚀 Installation
Python >=3.12 is required.
Full install with MCP server support (from the repo root):
pip install -e ".[mcp]"
Base install (no MCP server):
pip install -e .
The authoritative dependency source is pyproject.toml (package stream-dse). The base install pulls in zigzag-dse, ortools>=9.15 (the default, license-free MILP backend), pydantic, pydot, and xdsl. Optional extras: [mcp] adds fastmcp (required for the MCP server); [gurobi] adds gurobipy (commercial solver, opt-in).
AIE code generation
AIE-target MLIR codegen and tracing additionally need the AMD AIE toolchain (mlir_aie, llvm-aie, xdsl-aie, snax-mlir, aie-python-extras). These are git/URL installs that PyPI does not allow in package metadata, so a console script installs them after the base install rather than via an extra:
pip install -e . # or, once published: pip install stream-dse
stream-setup-aie # installs the AIE toolchain into the current environment
stream-setup-aie --dry-run prints exactly what it will install without making changes.
⚠️ Platform caveat: the AIE toolchain is Linux x86_64 only (manylinux wheels), CPython 3.12 or 3.13.
💡 Solver license note: OR-Tools (
ortools_gscip, the default backend) is open-source and needs no license. Gurobi requires the[gurobi]extra (pip install -e ".[gurobi]") plus a separate commercial license;backend="gurobi"errors at solve time without a valid license.
Optional pre-commit setup:
pre-commit install
⚡ Quick Start
Price a small two-Conv workload (a committed test fixture) on a TPU-like quad-core system, with a mapping the generic mapping generator proposes, through the public API; just co-2conv runs the same pair from the test matrix (this repo uses just as a task runner).
🧩 Hardware and Core Types
An accelerator in Stream is described as a system of heterogeneous dataflow cores. Core roles include compute, memory, shim, and offchip; example dataflow core types include AIE, TPU-like, and pooling.
Hardware and mapping files are organized as follows:
stream/inputs/examples/hardware/- system-level hardware YAMLs (e.g.tpu_like_quad_core.yaml,eyeriss_like_*.yaml,simba*.yaml,fusemax.yaml).stream/inputs/examples/hardware/cores/- per-core-type YAMLs (e.g.tpu_like.yaml,pooling.yaml,simd.yaml,offchip.yaml,eyeriss_like.yaml).stream/inputs/aie/hardware/andstream/inputs/aie/hardware/cores/- AMD AIE example core types (e.g.aie_tile.yaml,mem_tile_256KB.yaml,shim_dma.yaml).stream/inputs/examples/mapping/,stream/inputs/aie/mapping/, andstream/inputs/testing/mapping/- mapping descriptions.
A mapping can be generated (as in Quick Start above) or hand-written and passed as mapping.
📊 Workload × Hardware Matrix
The generic CO pipeline runs any ONNX workload on any of the example hardware systems. The repo ships two small workloads and exercises them across all eight non-AIE example architectures, through the pytest suite (tests/test_hardware_combinations.py).
Workloads - committed test fixtures under stream/inputs/testing/workload/ (weight values are cleared, only tensor shapes matter for cost estimation, so the ONNX stay tiny; just gen-workloads regenerates them via the builders):
- 2-conv - two chained Conv layers (
make_2_conv.py). - swiglu - a 5-node SwiGLU block: two Gemms, SiLU, an elementwise Mul, and a down-projection Gemm (
make_swiglu.py).
Hardware (stream/inputs/examples/hardware/) |
Description | 2-conv | swiglu |
|---|---|---|---|
eyeriss_like_single_core |
one Eyeriss-like compute core (+ pooling, SIMD, DRAM) | ✓ | ✓ |
eyeriss_like_dual_core |
two Eyeriss-like compute cores | ✓ | ✓ |
eyeriss_like_quad_core |
four Eyeriss-like compute cores | ✓ | ✓ |
tpu_like_quad_core |
four TPU-like compute cores | ✓ | ✓ |
simba_small |
small Simba chiplet mesh | ✓ | ✓ |
simba |
36-core Simba chiplet mesh | ✓ | ✓ |
fusemax |
FuseMax array + vector + DRAM | ✓ | ✓ |
meta_prototype_dual_core_simd_offchip |
two Meta-prototype compute cores (+ pooling, SIMD, DRAM) | ✓ | ✓ |
✓ = completes through the generic CO pipeline. All combinations run in the default fast suite; on these small single-fusion-group workloads even the 36-core simba mesh finishes in seconds.
Run one combination - hw is any hardware stem from the table (default tpu_like_quad_core):
just co-2conv fusemax # 2-conv on an architecture
just co-swiglu simba_small # swiglu on an architecture
Run the whole matrix - the justfile wraps pytest tests/test_hardware_combinations.py, which runs 2-conv + swiglu over all eight architectures plus a parse-only check confirming every hardware definition loads:
just matrix # parse + 2-conv + swiglu over all 8 architectures (incl. simba)
🐍 Public API
stream/api.py has three calls. Each takes a hardware description (a YAML path or an Accelerator), a workload (anything a registered frontend loads, such as an ONNX path, or a Workload) and an output directory:
evaluate_mapping(hardware, workload, output_path, mapping=None, options=None)solves the allocation of each fused group and returns aMappingEstimate. Without a mapping, the mapping generator that claims the hardware proposes one.select_mapping(hardware, workload, output_path, candidates, options=None)returns the estimate of the cheapest of the candidate mappings.generate_code(hardware, workload, output_path, mapping=None, options=None)also writes each fused group's design underoutput_path/group_<index>, through the code generation backend that claims the hardware.
import tempfile
from stream.api import evaluate_mapping
with tempfile.TemporaryDirectory() as tmp:
estimate = evaluate_mapping(
"stream/inputs/examples/hardware/tpu_like_quad_core.yaml",
"stream/inputs/testing/workload/2conv_1_8_32_32_16_32_3.onnx",
tmp,
)
print("cycles:", estimate.cycles)
A MappingEstimate holds cycles, the fused groups' estimates plus the reconfiguration the hardware declares, the per-group group_cycles, and the solved context, whose useful keys are scheduler, workload, accelerator and group_latencies. SolveOptions sets the solver backend, the number of columns, the constraint selection, the kernel library, the tile search and instrumentation, and its stage_options carries what a plugin's stages read, such as fusion_cut_points and intra_core_tiling for the generic mapping generator or npu and trace_size for the AIE code generator.
Whatever depends on the hardware is found through entry-point groups, so a separate package extends Stream without a fork: stream.frontends (workload formats), stream.mapping_generators (a mapping when none is given), stream.constraints (namespace MILP constraints), stream.core_cost_backends (per-core cost) and stream.codegen_backends (code generation).
🤖 MCP Server (for AI agents)
Stream ships an MCP server (stream/mcp/server.py, server name stream) that lets an AI agent submit and inspect TETRA CO jobs. Requires the [mcp] extra (pip install -e ".[mcp]").
⚠️ Install caveat:
[mcp]does not currently resolve against the pinned PyPIxdsl 0.29.1- fastmcp's dependency tree needs newertyping-extensions/pydanticthan xdsl 0.29.1 permits. For now it installs only in the dev environment that uses the git build of xdsl; a clean fix awaits the xdsl upgrade.
Launch command (from the repo root):
python3 -c "from stream.mcp.server import mcp; mcp.run(transport='stdio')"
The server runs on STDIO (JSON-RPC) transport and blocks until the client disconnects.
The 6 tools:
| Tool | Purpose |
|---|---|
run_optimization(hardware, workload, mapping, output_path, backend, ...) |
Submit a TETRA CO job; returns a job_id immediately; solve runs in the background. |
poll_optimization(job_id) |
Check job status (pending / running / complete / failed / not_found). |
get_workload_ir(workload=None, experiment_id=None) |
Return the workload DAG as WorkloadIR JSON. |
get_accelerator_ir(hardware=None, experiment_id=None) |
Return the hardware model as AcceleratorIR JSON. |
get_allocation_ir(job_id) |
Return the TETRA allocation result as AllocationIR JSON (3 persona views). |
get_solve_stats(job_id) |
Return MILP solve statistics (objective, time, gap, node count, backend). |
Run / poll / inspect flow:
run_optimization(...)returns{"job_id": "...", "status": "pending"}.- Poll
poll_optimization(job_id)until{"status": "complete"}. - Inspect with
get_allocation_ir(job_id)for theAllocationIR(algorithmic / hardware / compiler views) andget_solve_stats(job_id)for solve statistics.
🧠 Working in This Repo (AI agents)
Programmatic / IR API for structured JSON output:
from stream.ir import WorkloadIR, AcceleratorIR, AllocationIR
# ctx = evaluate_mapping(...).context
workload_ir = WorkloadIR.from_internal(ctx.get("workload"))
accelerator_ir = AcceleratorIR.from_internal(ctx.get("accelerator"))
allocation_ir = AllocationIR.from_internal(ctx.get("scheduler"))
workload_data = workload_ir.model_dump() # JSON-compatible dict
hardware_data = accelerator_ir.model_dump()
allocation_data = allocation_ir.model_dump()
AllocationIR offers .algorithmic_view(), .hardware_view(), and .compiler_view() persona views.
📚 Further Documentation
- Hosted documentation site: kuleuven-micas.github.io/stream, the human-facing docs (installation, getting started, the workload/hardware/mapping input formats, and driving Stream from an AI agent via the MCP server and IR models), rebuilt from
docs/on every push tomain. - Stream paper (IEEE): A. Symons, L. Mei, S. Colleman, P. Houshmand, S. Karl and M. Verhelst, "Stream: Design Space Exploration of Layer-Fused DNNs on Heterogeneous Dataflow Accelerators".
- ZigZag: zigzag-project.github.io/zigzag, the per-core cost-estimation framework Stream builds on.
Metadata
Release files for stream-dse 1.15.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| stream_dse-1.15.0.tar.gz | 399.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| stream_dse-1.15.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 858.7 kB
Release files / stream_dse-1.15.0.tar.gz
| Download URL | stream_dse-1.15.0.tar.gz |
|---|---|
| Size | 399.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
ee39522fed18a981e7dcaf017ebe7407c8ef99b769a626a01601306d77697e71
|
|
BLAKE2b-256 checksum How to use checksums |
587101e0dd23413acc43010d28d838ec9134cc1a34137e58a6e03a8bc4848a58
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.3
|
Release files / stream_dse-1.15.0-py3-none-any.whl
| Download URL | stream_dse-1.15.0-py3-none-any.whl |
|---|---|
| Size | 459.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
2d45d4a0c7a1676f184b88b9822f565bd66bb1cb5d12970c0fdbae2e1030f17f
|
|
BLAKE2b-256 checksum How to use checksums |
797d08f0b79497b8a3ce686ac9dfd13a1ada4c38b590ab46928d1dbff054f016
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.3
|