Skip to main content

Filedge

CI codecov Python License Ruff PRs Welcome

Files are the universal building block of data engineering. Whether data starts in Kafka, Stripe's API, a partner SFTP, or a CDC stream, every reliable pipeline eventually crystallizes it into a file before it touches the warehouse. Filedge is the load boundary built around that fact: atomic per-file ingestion, content-hash idempotency, and a full audit trail — into SQLite, PostgreSQL, BigQuery, Databricks, Snowflake, or DuckDB.

Why files?

Streams are continuous; files are discrete. That discreteness is what makes ingestion auditable: a file has a SHA-256, a row count, a state in the audit DB, and a row-level provenance trail in the destination. Every downstream question — did we load this?, replay this, where did this row come from? — has a deterministic anchor.

Filedge starts where the file lands and ends when its rows are committed. Upstream is your choice: dlt or vendor exporters for APIs, Kafka Connect or Vector for queues, rclone for SFTP. Downstream is your warehouse. The hard part in between — retry-safe commits, dedupe, retries, lineage — is all Filedge does.

What it gives you that a hand-rolled DAG doesn't

Failure mode Typical pipeline Filedge
Half-written tables after a crash Manual cleanup Per-file atomic commit, retry-safe by content hash
"Did we already load this file?" Filename heuristics SHA-256 dedupe at the entry point
"Where did this row come from?" Grep logs _source_file_hash + _ingested_at on every row
Stale lock from a killed worker Page someone Reclaimed automatically on next run
One bad file blocks the pipeline Skip and forget Bounded retry → terminal FAILED with audit
Schema drift in destination Silent corruption Loud failure with a clear diff

How it differs from neighbors

  • vs Airbyte / Fivetran / dlt — those fetch (paginate APIs, manage cursors). Filedge lands — it takes whatever they produce as files and makes the write to the warehouse audit-grade. Use them as Fetchers in front of Filedge.
  • vs Kafka Connect / Flink / Spark Structured Streaming — streaming systems own continuous offsets and incremental state. Filedge owns the file as the unit of work — simpler to reason about, replay, and audit. Materialize queues to files, then ingest.
  • vs Airflow + custom Python loaders — same DAG shape, but partial-load corruption, lock reclaim, retry caps, idempotent CDC apply, and row provenance are already wired in.
  • vs Iceberg / Delta tables — those are table formats. Filedge is what writes to them (or to plain BigQuery / Postgres / Databricks tables) with the per-file commit guarantee.

Quick start

Requires uv.

uv sync --extra dev                          # core (SQLite)
uv sync --extra dev --extra postgres         # + PostgreSQL
uv sync --extra dev --extra bigquery         # + BigQuery
uv sync --extra dev --extra databricks       # + Databricks
uv sync --extra dev --extra snowflake        # + Snowflake
uv sync --extra dev --extra duckdb           # + DuckDB
uv sync --extra dev --extra authoring        # + Authoring UI
uv sync --extra dev --extra excel            # + Excel (.xlsx)
uv sync --extra dev --extra kafka            # + Reference Queue Materializer

Declare a pipeline:

# pipeline.yaml
format: csv
dest_table: orders
write_mode: append          # append | truncate | cdc
retry_cap: 3
batch_size: 1000

connector:
  type: sqlite
  url: sqlite:///orders.db

columns:
  - { source: order_id,   dest: order_id,   type: string,  required: true }
  - { source: amount,     dest: amount,     type: float,   required: true }
  - { source: order_date, dest: order_date, type: date }

Run it:

filedge run --dir ./incoming --config pipeline.yaml --audit-db-url sqlite:///filedge.db
# Committed: 3  Failed: 0  Skipped: 0  New: 3  Reclaimed: 0  Retried: 0

filedge status --audit-db-url sqlite:///filedge.db
# PENDING: 0  PROCESSING: 0  COMMITTED: 3  FAILED: 0

Don't know the schema yet? filedge inspect data.csv samples the file and prints a columns: block with confidence tiers ready to paste.

Prefer to author interactively? filedge author data.csv launches a local terminal UI that runs schema inference, lets you review columns, write modes, connectors, and field encryption, validates the result, and writes a ready-to-run pipeline folder. To revise a pipeline later, filedge author --pipeline pipelines/<id> re-opens it in place — or run filedge author with no arguments to browse and pick from the registry. See the author guide.

Pulling from APIs or queues? Use an upstream Fetcher or Queue Materializer to land complete Files, then run Filedge. The first-party companions filedge-fetch and filedge-materialize demonstrate the audited materialize-to-files contract for API Sources, EDGAR companyConcept, and Kafka Queue Sources. See the API sources and queue sources guides.

Connectors

The destination is configured via a connector: block in pipeline.yaml. Built-ins:

Destination Extra Notes
SQLite (core) Default for local dev; configure with type: sqlite and a url
PostgreSQL postgres COPY bulk load; idempotent via per-hash DELETE
BigQuery bigquery NDJSON staging + load job; job-ID-keyed idempotency (7-day window)
Databricks databricks Unity Catalog volume staging
Snowflake snowflake Idempotent via per-hash DELETE + INSERT in one transaction
DuckDB duckdb File-based; single-writer, fails fast if locked

See docs/guides/run.md for full connector config, credentials, and live-integration test setup.

How a run works

filedge run
├── Reset FAILED below retry_cap → PENDING
├── Reclaim stale PROCESSING locks → PENDING
├── Connector: ensure destination table exists
├── Hash files in watched dir; enqueue new hashes as PENDING
└── For each PENDING file:
    ├── Audit DB: mark PROCESSING        (distributed lock)
    ├── Connector: stream rows → commit  (idempotent per file_hash)
    └── Audit DB: mark COMMITTED / FAILED

The audit DB and the destination are separate systems. A crash between connector commit and audit mark leaves the file PROCESSING — the next run reclaims it, and the connector's per-hash idempotency guarantees no duplicate rows.

More

License

Apache 2.0 — see LICENSE.

Release files for filedge 0.6.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for filedge 0.6.0
File Size Uploaded
filedge-0.6.0.tar.gz 541.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for filedge 0.6.0
File Interpreter ABI Platform
filedge-0.6.0-py3-none-any.whl Python 3 none any Details

Total release size: 695.8 kB

Release files / filedge-0.6.0.tar.gz

Download URL filedge-0.6.0.tar.gz
Size 541.5 kB
Tags Source
SHA-256 checksum
How to use checksums
1f871c4f1f85b23aabda86fad1919a188d8066e89d9f8e2a247ccd31b2dd2854
BLAKE2b-256 checksum
How to use checksums
bf73feffba697eec43c7b9fd080e8644de9bd38d9133840f396a6ccf659c2235
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.11.19 {"installer":{"name":"uv","version":"0.11.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release files / filedge-0.6.0-py3-none-any.whl

Download URL filedge-0.6.0-py3-none-any.whl
Size 154.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
4f0cc4ea0b1868defd560b34e12411384ee0264e200388a3a02ac2cbd3ac6c3c
BLAKE2b-256 checksum
How to use checksums
b3ed2751e0af7b61bb2b9ed919e2ead03b594a9a307dddfc13001adbc6c08243
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via uv/0.11.19 {"installer":{"name":"uv","version":"0.11.19","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

Release history Release notifications | RSS feed

This release

0.6.0 This release

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page