Skip to main content

Config-driven data flow framework with pluggable ops and extensions (Excel, Postgres, FastAPI).

Project description

flowbook — a framework for flexible data flows.

Quickstart

pip install flowbook
flowbook --version
flowbook doctor

Core-only install has no heavy dependencies. For Excel (.xlsx, .xls), Postgres, and FastAPI extensions:

pip install "flowbook[full]"

Optional: add bundled configs for hands-on (no extra deps):

pip install "flowbook[full,demo]"

Dev CLI (Typer/Rich) for local development and demos:

pip install "flowbook[dev]"
flowbook --version
flowbook doctor
flowbook db init    # Create schema (first-time only)
flowbook db reset   # DB reset + seed (needs flowbook[dev])
flowbook db up      # Start Postgres (Docker, optional)
flowbook api       # Run API (uvicorn)
flowbook streamlit # Streamlit UI (venv)
flowbook hands-on   # API hands-on flow
flowbook steps list # List available steps (ops)
flowbook steps show <op_name>  # Show step spec (inputs, outputs)

flowbook doctor prints Python/OS/flowbook version and suggests pip install "flowbook[excel]", "flowbook[postgres]", "flowbook[fastapi]", or "flowbook[full]" for missing extensions.

Concept

  • Config-driven: Which steps run, in what order, and how inputs are bound—all come from config (plan config, ConfigStore, plan templates). Change the flow without changing framework code.
  • Steps (ops): Flowbook uses pluggable steps (ops). Each step has inputs and outputs; plans compose steps. List available steps: flowbook steps list. Show step details (docstring, inputs, outputs): flowbook steps show <op_name>. API: GET /steps, GET /steps/{op_name}. Streamlit: Steps tab.
  • Extend via extensions: The behavior of each step is an op registered in a Registry. Add new ops in your own package; the framework only resolves op name → run op. No need to touch the core.
  • Single data rule: Data lives only in Artifacts; steps receive resolved values and return a dict. Contracts are explicit (e.g. PortSpec for inputs).
  • AI-friendly: Config (templates, rules, mappings) is easy for LLMs to generate or choose. New ops (including AI-backed ones) plug in the same way. You can call LLMs inside an op; the engine stays agnostic.

Usage (high-level)

  1. Engine = store (artifacts) + registry (ops) + optional config store. You build it once.
  2. Session = with engine.create_run() as session:. Put inputs (logical name → value), then run a plan config (list of steps with name, op, inputs).
  3. Optionally run a planner first (e.g. plan_from_template); it produces a plan config that you then execute in the same session.
  4. Steps read from the store (via resolved inputs) and write outputs back; later steps can depend on them. All orchestration is driven by config; new capabilities are new ops in your extensions.

To add your own steps: see Adding custom steps (minimal: one module + one line at startup; optional: package with entry points). To compose plans from steps: see Plan from steps. To add CLI commands: see Adding custom CLI.

Development

  • CI before commit: npm run ci (lint, typecheck, test) runs automatically via pre-commit. It runs only unit (and smoke) tests; integration and e2e are skipped so CI does not require Postgres. After clone, run:
    poetry install
    pre-commit install
    
    (The dev group includes full extras so tests can run; for a minimal env use pip install flowbook only.)
  • Full test suite (integration + e2e): Start Postgres (see below), then npm run test or poetry run pytest. To run only integration: poetry run pytest -m integration.
  • Releasing: See Releasing. Publish = tag + twine upload. Pre-release (alpha/beta) = push to dev branch only.

License

Apache License 2.0

Dev commands

The commands below are for local development. The Docker subcommands (flowbook db up, flowbook api up, flowbook streamlit up) require an infra/ directory (compose files, env files). Clone this repo or copy infra/ to use them.

Dev / Demo

API and Streamlit UI run from the repo for development and demos. Postgres is required for the API.

Postgres

Run Postgres (local install, cloud, or Docker):

# Option: Docker Compose (requires infra/)
flowbook db up    # Start Postgres (uses infra/.env.postgres)
flowbook db down  # Stop Postgres

Use --env-file PATH / --no-env-file to override. Or run Postgres yourself and set FLOWBOOK_DATABASE_URL.

First-time: Docker Compose runs init scripts automatically. For cloud or local Postgres, run FLOWBOOK_DB_RESET=1 flowbook db init to create the schema.

API

FLOWBOOK_DATABASE_URL=postgresql+psycopg://flowbook:flowbook@localhost:5432/flowbook poetry run flowbook api

API docs: http://localhost:8000/docs

Docker: flowbook api up / flowbook api down (uses infra/.env.api, network_mode: host). Cloud deploy: docker build -f infra/Dockerfile.api -t flowbook-api . — set FLOWBOOK_DATABASE_URL via env/secrets; image listens on $PORT.

Streamlit UI

flowbook streamlit

Runs in a separate venv (pandas version compatibility). Requires the API. If you see ModuleNotFoundError: altair.vegalite.v4, remove .venv-ui and run again.

Tabs: Steps, Inspect, Import, Artifacts, Export, Download, Configs.

Docker: flowbook streamlit up / flowbook streamlit down (uses infra/.env.streamlit). Cloud deploy: docker build -f infra/Dockerfile.streamlit -t flowbook-streamlit . — set FLOWBOOK_API_URL via env/secrets.

DB init and reset (dev only)

First-time setup (empty DB, no tables): Create schema before reset. Requires FLOWBOOK_DB_RESET=1 and localhost DSN.

Reset (truncate + seed): Same safety requirements. Run flowbook db init first if the DB has no schema.

FLOWBOOK_DATABASE_URL=... FLOWBOOK_DB_RESET=1 flowbook db init
FLOWBOOK_DATABASE_URL=... FLOWBOOK_DB_RESET=1 flowbook db reset

flowbook db init creates entities, runs, entity_runs, artifacts, configs. flowbook db reset truncates them and seeds from bundled configs (flowbook[demo]) + overlay from configs/ (default --config-dir configs). Use --config-dir bundled for bundled only.

Hands-on flow

flowbook hands-on

Runs Health -> Inspect -> Import -> Artifacts -> Export -> Download (interactive). Requires API up and a fixture. Excel import supports .xlsx and .xls (region-based import uses df-based detection). Generate fixture:

flowbook fixture generate -o tests/fixtures/excel

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

flowbook-0.1.0a4.tar.gz (67.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

flowbook-0.1.0a4-py3-none-any.whl (103.5 kB view details)

Uploaded Python 3

File details

Details for the file flowbook-0.1.0a4.tar.gz.

File metadata

  • Download URL: flowbook-0.1.0a4.tar.gz
  • Upload date:
  • Size: 67.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.11

File hashes

Hashes for flowbook-0.1.0a4.tar.gz
Algorithm Hash digest
SHA256 0e190ecc23d55fee1f194ebf5a07539b83c22e547d48aab1f6ce2d88e62f9614
MD5 2abbc61a7221ddf1db79b1fd794fcbc9
BLAKE2b-256 a99f5aacd4796dc7cacae2db19d4c6ad52a619fcd428cec33dd3654d0a2aaaab

See more details on using hashes here.

File details

Details for the file flowbook-0.1.0a4-py3-none-any.whl.

File metadata

  • Download URL: flowbook-0.1.0a4-py3-none-any.whl
  • Upload date:
  • Size: 103.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.11

File hashes

Hashes for flowbook-0.1.0a4-py3-none-any.whl
Algorithm Hash digest
SHA256 ebcc1220be1d3e637c1604fd24ecd1140e0573465c1d4eb0c0ac18d05901cd64
MD5 ea26616ab0c741f44c4f7e2f8c119e7e
BLAKE2b-256 e406946e0408b2cf06ef36d7eb068fffb5cef2c857c239a9837c2cb1e48edde0

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page