Skip to main content

Config-driven data flow framework with pluggable ops and extensions (Excel, Postgres, FastAPI).

Project description

flowbook — a framework for flexible data flows.

Quickstart

pip install flowbook
flowbook --version
flowbook doctor

Core-only install has no heavy dependencies. For Excel (.xlsx, .xls), Postgres, and FastAPI extensions:

pip install "flowbook[full]"

Optional: add bundled configs for hands-on (no extra deps):

pip install "flowbook[full,demo]"

Dev CLI (Typer/Rich) for local development and demos:

pip install "flowbook[dev]"
flowbook --version
flowbook doctor
flowbook db init    # Create schema (first-time only)
flowbook db reset   # DB reset + seed (needs flowbook[dev])
flowbook db up      # Start Postgres (Docker, optional)
flowbook api       # Run API (uvicorn)
flowbook streamlit # Streamlit UI (venv)
flowbook hands-on   # API hands-on flow
flowbook steps list # List available steps (ops)
flowbook steps show <op_name>  # Show step spec (inputs, outputs)

flowbook doctor prints Python/OS/flowbook version and suggests pip install "flowbook[excel]", "flowbook[postgres]", "flowbook[fastapi]", or "flowbook[full]" for missing extensions.

Concept

  • Config-driven: Which steps run, in what order, and how inputs are bound—all come from config (plan config, ConfigStore, plans). Change the flow without changing framework code.
  • Steps (ops): Flowbook uses pluggable steps (ops). Each step has inputs and outputs; plans compose steps. List available steps: flowbook steps list. Show step details (docstring, inputs, outputs): flowbook steps show <op_name>. API: GET /steps, GET /steps/{op_name}. Streamlit: Steps tab.
  • Extend via extensions: The behavior of each step is an op registered in a Registry. Add new ops in your own package; the framework only resolves op name → run op. No need to touch the core.
  • Single data rule: Data lives only in Artifacts; steps receive resolved values and return a dict. Contracts are explicit (e.g. PortSpec for inputs).
  • AI-friendly: Config (plans, rules, mappings) is easy for LLMs to generate or choose. New ops (including AI-backed ones) plug in the same way. You can call LLMs inside an op; the engine stays agnostic.

Usage (high-level)

  1. Engine = store (artifacts) + registry (ops) + optional config store. You build it once.
  2. Session = with engine.create_run() as session:. Put inputs (logical name → value), then run a plan config (list of steps with name, op, inputs).
  3. Optionally run a planner first (e.g. load_plan); it produces a plan config that you then execute in the same session.
  4. Steps read from the store (via resolved inputs) and write outputs back; later steps can depend on them. All orchestration is driven by config; new capabilities are new ops in your extensions.

To add your own steps: see Adding custom steps (minimal: one module + one line at startup; optional: package with entry points). To compose plans from steps: see Plan from steps. To add CLI commands: see Adding custom CLI.

Development

  • CI before commit: npm run ci (lint, typecheck, test) runs automatically via pre-commit. It runs only unit (and smoke) tests; integration and e2e are skipped so CI does not require Postgres. After clone, run:
    uv sync
    pre-commit install
    
    (The dev group includes full extras so tests can run; for a minimal env use pip install flowbook only.)
  • Full test suite (integration + e2e): Start Postgres (see below), then npm run test or uv run pytest. To run only integration: uv run pytest -m integration.
  • Releasing: See Releasing. Publish = tag + twine upload. Pre-release (alpha/beta) = push to dev branch only.

License

Apache License 2.0

Dev commands

The commands below are for local development. The Docker subcommands (flowbook db up, flowbook api up, flowbook streamlit up) require an infra/ directory (compose files, env files). Clone this repo or copy infra/ to use them.

.env: uv does not auto-load .env. Use uv run --env-file .env flowbook ... or export UV_ENV_FILE=.env. Or use npm run api / npm run hands-on which pass --env-file .env.

Dev / Demo

API and Streamlit UI run from the repo for development and demos. Postgres is required for the API.

Postgres

Run Postgres (local install, cloud, or Docker):

# Option: Docker Compose (requires infra/)
flowbook db up    # Start Postgres (uses infra/.env.postgres)
flowbook db down  # Stop Postgres

Use --env-file PATH / --no-env-file to override. Or run Postgres yourself and set FLOWBOOK_DATABASE_URL.

First-time: Docker Compose runs init scripts automatically. For cloud or local Postgres, run FLOWBOOK_DB_RESET=1 flowbook db init to create the schema.

API

# With .env: cp .env.example .env, then:
npm run api
# Or: uv run --env-file .env flowbook api
# Or: UV_ENV_FILE=.env uv run flowbook api

API docs: http://localhost:8000/docs

Docker: flowbook api up / flowbook api down (uses infra/.env.api, network_mode: host). Cloud deploy: docker build -f infra/Dockerfile.api -t flowbook-api . — set FLOWBOOK_DATABASE_URL via env/secrets; image listens on $PORT.

Streamlit UI

flowbook streamlit

Runs in a separate venv (pandas version compatibility). Requires the API. If you see ModuleNotFoundError: altair.vegalite.v4, remove .venv-ui and run again.

Tabs: Steps, Inspect, Import, Artifacts, Export, Download, Configs.

Docker: flowbook streamlit up / flowbook streamlit down (uses infra/.env.streamlit). Cloud deploy: docker build -f infra/Dockerfile.streamlit -t flowbook-streamlit . — set FLOWBOOK_API_URL via env/secrets.

DB init and reset (dev only)

First-time setup (empty DB, no tables): Create schema before reset. Requires FLOWBOOK_DB_RESET=1 and localhost DSN.

Reset (truncate + seed): Same safety requirements. Run flowbook db init first if the DB has no schema.

FLOWBOOK_DATABASE_URL=... FLOWBOOK_DB_RESET=1 flowbook db init
FLOWBOOK_DATABASE_URL=... FLOWBOOK_DB_RESET=1 flowbook db reset

flowbook db init creates entities, runs, results, artifacts, configs. flowbook db reset truncates them and seeds from bundled configs (flowbook[demo]) + overlay from configs/ (default --config-dir configs). Use --config-dir bundled for bundled only.

Hands-on flow

flowbook hands-on

Runs Health -> Inspect -> Import -> Artifacts -> Export -> Download (interactive). Requires API up and a fixture. Excel import supports .xlsx and .xls (region-based import uses df-based detection). Generate fixture:

flowbook fixture generate -o tests/fixtures/excel

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

flowbook-0.1.0a10.tar.gz (231.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

flowbook-0.1.0a10-py3-none-any.whl (138.7 kB view details)

Uploaded Python 3

File details

Details for the file flowbook-0.1.0a10.tar.gz.

File metadata

  • Download URL: flowbook-0.1.0a10.tar.gz
  • Upload date:
  • Size: 231.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.11

File hashes

Hashes for flowbook-0.1.0a10.tar.gz
Algorithm Hash digest
SHA256 53a0afe4d1fcc6bb713a26c1ed5aa3888ec84a7d7190533545c1b87bd992ff38
MD5 b2feb74f6895cde2cfac8b9e93a1bbcd
BLAKE2b-256 3e7aea529d0705a2c3e58f9de3e185a30842a7b7613110ae26e7d06b8038b331

See more details on using hashes here.

File details

Details for the file flowbook-0.1.0a10-py3-none-any.whl.

File metadata

  • Download URL: flowbook-0.1.0a10-py3-none-any.whl
  • Upload date:
  • Size: 138.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.11

File hashes

Hashes for flowbook-0.1.0a10-py3-none-any.whl
Algorithm Hash digest
SHA256 9eb9d8dee5beb50ee091ce0b85bc4722ee99e3fad77d15b3e4b101f18cde5218
MD5 185a0ef18a86b5192cf6c761e7f4573f
BLAKE2b-256 15330e5713a64a6564622f412cd86bad163f9a391cb7e4489acc201ee5089f24

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page