Config-driven data flow framework with pluggable ops and extensions (Excel, Postgres, FastAPI).
Project description
flowbook — a framework for flexible data flows.
Quickstart
pip install flowbook
flowbook --version
flowbook doctor
Core-only install has no heavy dependencies. For Excel (.xlsx, .xls), Postgres, and FastAPI extensions:
pip install "flowbook[full]"
Optional: add bundled configs for hands-on (no extra deps):
pip install "flowbook[full,demo]"
Dev CLI (Typer/Rich) for local development and demos:
pip install "flowbook[dev]"
flowbook --version
flowbook doctor
flowbook db init # Create schema (first-time only)
flowbook db reset # DB reset + seed (needs flowbook[dev])
flowbook db up # Start Postgres (Docker, optional)
flowbook api # Run API (uvicorn)
flowbook streamlit # Streamlit UI (venv)
flowbook hands-on # API hands-on flow
flowbook steps list # List available steps (ops)
flowbook steps show <op_name> # Show step spec (inputs, outputs)
flowbook doctor prints Python/OS/flowbook version and suggests pip install "flowbook[excel]", "flowbook[postgres]", "flowbook[fastapi]", or "flowbook[full]" for missing extensions.
Concept
- Config-driven: Which steps run, in what order, and how inputs are bound—all come from config (plan config, ConfigStore, plan templates). Change the flow without changing framework code.
- Steps (ops): Flowbook uses pluggable steps (ops). Each step has inputs and outputs; plans compose steps. List available steps:
flowbook steps list. Show step details (docstring, inputs, outputs):flowbook steps show <op_name>. API:GET /steps,GET /steps/{op_name}. Streamlit: Steps tab. - Extend via extensions: The behavior of each step is an op registered in a
Registry. Add new ops in your own package; the framework only resolvesop name → run op. No need to touch the core. - Single data rule: Data lives only in Artifacts; steps receive resolved values and return a dict. Contracts are explicit (e.g.
PortSpecfor inputs). - AI-friendly: Config (templates, rules, mappings) is easy for LLMs to generate or choose. New ops (including AI-backed ones) plug in the same way. You can call LLMs inside an op; the engine stays agnostic.
Usage (high-level)
- Engine = store (artifacts) + registry (ops) + optional config store. You build it once.
- Session =
with engine.create_run() as session:. Put inputs (logical name → value), then run a plan config (list of steps withname,op,inputs). - Optionally run a planner first (e.g.
plan_from_template); it produces a plan config that you then execute in the same session. - Steps read from the store (via resolved inputs) and write outputs back; later steps can depend on them. All orchestration is driven by config; new capabilities are new ops in your extensions.
To add your own steps: see Adding custom steps (minimal: one module + one line at startup; optional: package with entry points). To compose plans from steps: see Plan from steps. To add CLI commands: see Adding custom CLI.
Development
- CI before commit:
npm run ci(lint, typecheck, test) runs automatically via pre-commit. It runs only unit (and smoke) tests; integration and e2e are skipped so CI does not require Postgres. After clone, run:uv sync pre-commit install
(The dev group includes full extras so tests can run; for a minimal env usepip install flowbookonly.) - Full test suite (integration + e2e): Start Postgres (see below), then
npm run testoruv run pytest. To run only integration:uv run pytest -m integration. - Releasing: See Releasing. Publish = tag + twine upload. Pre-release (alpha/beta) = push to dev branch only.
License
Apache License 2.0
Dev commands
The commands below are for local development. The Docker subcommands (flowbook db up, flowbook api up, flowbook streamlit up) require an infra/ directory (compose files, env files). Clone this repo or copy infra/ to use them.
.env: uv does not auto-load .env. Use uv run --env-file .env flowbook ... or export UV_ENV_FILE=.env. Or use npm run api / npm run hands-on which pass --env-file .env.
Dev / Demo
API and Streamlit UI run from the repo for development and demos. Postgres is required for the API.
Postgres
Run Postgres (local install, cloud, or Docker):
# Option: Docker Compose (requires infra/)
flowbook db up # Start Postgres (uses infra/.env.postgres)
flowbook db down # Stop Postgres
Use --env-file PATH / --no-env-file to override. Or run Postgres yourself and set FLOWBOOK_DATABASE_URL.
First-time: Docker Compose runs init scripts automatically. For cloud or local Postgres, run FLOWBOOK_DB_RESET=1 flowbook db init to create the schema.
API
# With .env: cp .env.example .env, then:
npm run api
# Or: uv run --env-file .env flowbook api
# Or: UV_ENV_FILE=.env uv run flowbook api
API docs: http://localhost:8000/docs
Docker: flowbook api up / flowbook api down (uses infra/.env.api, network_mode: host). Cloud deploy: docker build -f infra/Dockerfile.api -t flowbook-api . — set FLOWBOOK_DATABASE_URL via env/secrets; image listens on $PORT.
Streamlit UI
flowbook streamlit
Runs in a separate venv (pandas version compatibility). Requires the API. If you see ModuleNotFoundError: altair.vegalite.v4, remove .venv-ui and run again.
Tabs: Steps, Inspect, Import, Artifacts, Export, Download, Configs.
Docker: flowbook streamlit up / flowbook streamlit down (uses infra/.env.streamlit). Cloud deploy: docker build -f infra/Dockerfile.streamlit -t flowbook-streamlit . — set FLOWBOOK_API_URL via env/secrets.
DB init and reset (dev only)
First-time setup (empty DB, no tables): Create schema before reset. Requires FLOWBOOK_DB_RESET=1 and localhost DSN.
Reset (truncate + seed): Same safety requirements. Run flowbook db init first if the DB has no schema.
FLOWBOOK_DATABASE_URL=... FLOWBOOK_DB_RESET=1 flowbook db init
FLOWBOOK_DATABASE_URL=... FLOWBOOK_DB_RESET=1 flowbook db reset
flowbook db init creates entities, runs, results, artifacts, configs. flowbook db reset truncates them and seeds from bundled configs (flowbook[demo]) + overlay from configs/ (default --config-dir configs). Use --config-dir bundled for bundled only.
Hands-on flow
flowbook hands-on
Runs Health -> Inspect -> Import -> Artifacts -> Export -> Download (interactive). Requires API up and a fixture. Excel import supports .xlsx and .xls (region-based import uses df-based detection). Generate fixture:
flowbook fixture generate -o tests/fixtures/excel
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file flowbook-0.1.0a6.tar.gz.
File metadata
- Download URL: flowbook-0.1.0a6.tar.gz
- Upload date:
- Size: 193.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.11
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
86848f200869250f6ff1ed3a376fec88b55e0a90ed201c7e179c09aaba9e3134
|
|
| MD5 |
4f5661e4344aa0cf1d120a8c45214c5f
|
|
| BLAKE2b-256 |
318ba7fbb906d2c74c9e5c9d0324a87796f1a90ffdbe4528f7e7fbd280c2cad9
|
File details
Details for the file flowbook-0.1.0a6-py3-none-any.whl.
File metadata
- Download URL: flowbook-0.1.0a6-py3-none-any.whl
- Upload date:
- Size: 109.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.11
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
cb8d397f1e590f3a4b2beadd8592b395928427244e639ed485539dd0a97d5c7d
|
|
| MD5 |
8cf6b7a32dc3bec62ac93cde277726fe
|
|
| BLAKE2b-256 |
0648de4536ddffb387347aa3a1a84d80d85d743d46ec5e02e1c8589f01febd8e
|