Skip to main content

minilake — Local Databricks API Emulator

MiniLake

Free, open-source local Databricks emulator for offline development and testing.

Real SQL & Spark execution · Unity Catalog hierarchy · Databricks SDK compatible · Terraform compatible · MIT licensed

GitHub release CI Release GHCR image License Python

Website · Documentation · GitHub · Container Image (GHCR)


MiniLake is a free, local Databricks API emulator — a single-developer tool for testing databricks-sdk/Terraform code against real SQL, real Delta Lake, and real Job execution, without paying for cloud compute.

Quick Start

# Option 1: PyPI
pip install minilake
minilake --port 8000

# Option 2: GitHub Container Registry
docker run -p 8000:8000 ghcr.io/dmux/minilake:latest

# Option 3: Clone and build
git clone https://github.com/dmux/minilake && cd minilake
docker compose up -d

# Verify (any option)
curl http://localhost:8000/_minilake/health

Then open http://localhost:8000/ui/ for the built-in SQL workspace — an Athena-style query editor with a data catalog, saved queries and query history (details).

No account, no API key, no sign-up — and the container image downloads nothing at runtime: DuckDB's delta extension and the Delta / Unity Catalog Spark jars are baked in at build time, so it works air-gapped (details).

Then point any Databricks client at it:

from databricks.sdk import WorkspaceClient

w = WorkspaceClient(host="http://localhost:8000", token="dev")
w.catalogs.create(name="vendas")

More in Getting Started.

Documentation

Getting Started Install, first catalog and query, internal endpoints
Web UI The built-in SQL workspace at /ui
Configuration Every environment variable, persistence, HTTPS/TLS
Databricks SDK Unity Catalog, warehouses, SQL and jobs from Python
Terraform & Asset Bundles The provider, and bundle deploy / bundle run
Spark & Delta Lake Real Delta files, real Spark jobs, spark.table() by name
MCP Server 67 tools for LLM agents — examples, tool reference, troubleshooting
Testing & development Running the suite, adding an API group
Releases & CI/CD How a tag becomes a published image
Feature status Endpoint-by-endpoint status and design rationale

Supported Services

Service Status Notes
Unity Catalog (catalogs, schemas, tables, volumes) ✅ Real Each catalog = its own DuckDB database (ATTACH), native catalog.schema.table addressing
EXTERNAL Delta Tables ✅ Real Real Delta files; INSERT/UPDATE/DELETE via a generated Spark job, reads via delta_scan()
SQL Statement Execution ✅ Real Real DuckDB; JSON_ARRAY/ARROW_STREAM/CSV, INLINE/EXTERNAL_LINKS; result manifest carries column types
SQL Warehouses ✅ Real Full CRUD + lifecycle
Query History ✅ Real w.query_history.list() over everything executed, failures included
Saved Queries ✅ Real w.queries.* CRUD with update_mask
Web UI ✅ Real Athena-style SQL workspace at /ui — see Web UI
Jobs ✅ Real Sibling Docker container execution (Spark) or subprocess fallback; real DAG scheduling (depends_on/run_if); sql_task.file
Workspace ✅ Real File-backed notebook/script storage; raw-bytes workspace-files sync powers databricks bundle deploy / bundle run
DBFS & Files API ✅ Real File-backed storage, chunked upload
Secrets ✅ Real Real CRUD; values only resolvable inside job env vars, never via direct API (matches real Databricks)
Clusters ✅ Real state machine CRUD + timed lifecycle transitions; no real Spark compute (by design)
Permissions ✅ Real CRUD Single-user "allow-all" default (by design — see Gaps)
Identity (SCIM) ✅ Static Fake current-user endpoint
Persistence (MINILAKE_PERSIST=1) ✅ Real JSON snapshot on shutdown, restored on startup
Unity Catalog protocol for Spark ✅ Real spark.table("cat.sch.tbl") resolves against minilake — see Spark & Delta Lake
JupyterLab + PySpark + Delta (optional) ✅ Real docker compose --profile notebook up
MCP Server (optional, MINILAKE_MCP=1) ✅ Real 67 tools + resources + prompts at /mcp — see MCP Server
Secrets ACLs, Repos/Git, multi-language notebooks, DBT/pipeline tasks, Model Registry, Vector Search, Dashboards 🚫 Not implemented Returns 501 NOT_IMPLEMENTED

Known Gaps

These are deliberate, not oversights — minilake targets one developer running it locally, not a shared or multi-tenant server:

  • No real authentication — any token is accepted; there's only ever one real user.
  • No access-control enforcement — the Permissions API is real CRUD but always allow-all, so a test that passes here says nothing about grants in a real workspace.
  • No real Spark compute for Clusters — state machine only; real compute happens through Jobs' sibling containers instead.
  • Single process, no HA — and DuckDB's single-writer model means concurrent load contends on locks.
  • Uneven test coveragejobs.py, sql_statements.py and unity_catalog.py are covered mostly on happy paths, not edge cases.
  • Secrets ACLs not implemented — scope/secret CRUD is real, ACL endpoints aren't.

Contributing

See CONTRIBUTING.md for the project structure, how to add a new API group, and the PR checklist.

License

MIT — see LICENSE.

Release files for minilake 1.7.6

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for minilake 1.7.6
File Size Uploaded
minilake-1.7.6.tar.gz 7.6 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for minilake 1.7.6
File Interpreter ABI Platform
minilake-1.7.6-py3-none-any.whl Python 3 none any Details

Total release size: 14.5 MB

Release files / minilake-1.7.6.tar.gz

Download URL minilake-1.7.6.tar.gz
Size 7.6 MB
Tags Source
SHA-256 checksum
How to use checksums
e8c415fee25355dfe918ebd8a2dfb8686b85bf0ee08cabcbb2714c82d82f4192
BLAKE2b-256 checksum
How to use checksums
9fee78692dc7525320d7c89fcaafef5fcb1b4ed8e188fff06ee129608c4e7482
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 9, 2026.

Transparency log

Release files / minilake-1.7.6-py3-none-any.whl

Download URL minilake-1.7.6-py3-none-any.whl
Size 6.9 MB
Tags Python 3
SHA-256 checksum
How to use checksums
9ca6636d84be5afe6ba9984541c7a30739e9862724b11024dd800545ed255c90
BLAKE2b-256 checksum
How to use checksums
c0251d406f2303150309a4ff7f9204760b0cf50bf67552b3950d6f5c6e9f83b4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 9, 2026.

Transparency log

Release history Release notifications | RSS feed

1.7.8

2 release files

1.7.7

2 release files

This release

1.7.6 This release

2 release files

1.7.4

2 release files

1.7.3

2 release files

1.7.2

2 release files

1.7.1

2 release files

1.7.0

1 release file

1.6.0

2 release files

1.5.1

2 release files

1.5.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page