Skip to main content

jupyter-databricks-kernel

PyPI version CI License Python

Ask DeepWiki

A Jupyter kernel for complete remote execution on Databricks clusters.

1. Features

  • Execute Python code entirely on Databricks clusters
    • Works with VS Code, JupyterLab, and other Jupyter frontends
    • CLI execution support with jupyter execute command
  • Automatic file synchronization to the Databricks cluster driver node
    • Syncs your local project files to the cluster driver node before each execution
    • Respects .gitignore patterns and configurable exclude rules
    • Configurable size limits to prevent syncing large files

2. Requirements

  • Python 3.11 or later
  • Databricks workspace with authentication configured (supports Personal Access Token, OAuth M2M with Service Principal, etc.)
  • Classic all-purpose cluster

3. Quick Start

  1. Install the kernel:

    uv add jupyter-databricks-kernel
    uv run python -m jupyter_databricks_kernel.install
    

    Install options:

    Option Description
    (default) Install to current venv (sys.prefix)
    --user Install to user site (~/.local/share/jupyter/kernels/)
    --prefix PATH Install to custom path
  2. Configure authentication and cluster:

    # Recommended: Use Databricks CLI to set up everything
    databricks auth login --configure-cluster
    

    This creates ~/.databrickscfg with authentication credentials and cluster ID.

    Alternatively, use environment variables:

    # Override cluster ID (optional, takes priority over ~/.databrickscfg)
    export DATABRICKS_CLUSTER_ID=your-cluster-id
    
    # Authentication (if not using ~/.databrickscfg)
    export DATABRICKS_HOST=https://your-workspace.cloud.databricks.com
    export DATABRICKS_TOKEN=your-personal-access-token
    
    # Service Principal authentication (alternative to PAT)
    export DATABRICKS_CLIENT_ID=your-client-id
    export DATABRICKS_CLIENT_SECRET=your-client-secret
    
    # Use specific profile from ~/.databrickscfg (optional)
    export DATABRICKS_CONFIG_PROFILE=your-profile-name
    

    For authentication options, see Databricks SDK Authentication.

  3. Open a notebook and select "Databricks" kernel:

    VS Code:

    1. Install the Jupyter extension
    2. Open a .ipynb file
    3. Click "Select Kernel" and choose "Databricks"

    JupyterLab:

    jupyter-lab
    

    Select "Databricks" from the kernel list.

  4. Run a simple test:

    spark.version
    

If the cluster is stopped, the first execution may take 5-6 minutes while the cluster starts.

Examples

See examples/ for sample projects:

  • table-exporter — Skinny notebook wrapper with pure Python business logic for exporting an existing Databricks table.

4. Configuration

4.1. Cluster ID

Cluster ID is read from (in order of priority):

  1. DATABRICKS_CLUSTER_ID environment variable
  2. ~/.databrickscfg (from active profile)
  3. .databricks/jupyter-databricks-kernel.json project routing config

Active profile is determined by DATABRICKS_CONFIG_PROFILE environment variable, or DEFAULT if not set.

Example ~/.databrickscfg:

[DEFAULT]
host = https://your-workspace.cloud.databricks.com
token = dapi...
cluster_id = 0123-456789-abcdef12

4.2. Sync Settings

You can configure file synchronization in pyproject.toml:

[tool.jupyter-databricks-kernel.sync]
enabled = true
source = "."
exclude = ["*.log", "data/"]
max_size_mb = 100.0
max_file_size_mb = 10.0
compression_level = 1
use_gitignore = true
Option Description Default
sync.enabled Enable file synchronization true
sync.source Source directory to sync "."
sync.exclude Additional exclude patterns []
sync.max_size_mb Maximum total project size in MB No limit
sync.max_file_size_mb Maximum individual file size in MB No limit
sync.use_gitignore Respect .gitignore patterns true
sync.workspace_extract_dir Custom driver-local extraction directory null (auto)

The extraction directory can also be set via the JUPYTER_DATABRICKS_KERNEL_EXTRACT_DIR environment variable, which takes priority over pyproject.toml.

Extraction runs on the cluster driver under /tmp by default. Legacy Databricks Workspace mount paths such as /Workspace/... are not used as a fallback and are rejected when configured explicitly.

By default, files are extracted to /tmp/jupyter_databricks_kernel/<project>-<hash>/ on the cluster driver node, where <project> is derived from the local project root directory name and <hash> is derived from the project root path. This single path works uniformly for user accounts and service principals while avoiding collisions between different projects with the same directory name.

Set sync.compression_level to choose the ZIP compression level from 0 (fast) through 9 (small). If unset, Python's zipfile default is used.

5. CLI Execution

You can execute notebooks from the command line using jupyter execute:

jupyter execute notebook.ipynb --kernel_name=databricks --inplace

To save the output to a different file:

jupyter execute notebook.ipynb --kernel_name=databricks --output=output.ipynb

5.1. Options

Option Description
--kernel_name Kernel name (use databricks)
--output Output file name
--inplace Overwrite input file with results
--timeout Cell execution timeout in seconds
--startup_timeout Kernel startup timeout in seconds (default: 60)
--allow-errors Continue execution even if a cell raises an error

5.2. Notes

If the cluster is stopped, kernel startup may take 5-6 minutes. Increase --startup_timeout to avoid timeout errors:

jupyter execute notebook.ipynb --kernel_name=databricks --startup_timeout=600

5.3. Runner CLI (databricks-run)

Execute scripts and notebooks directly without launching Jupyter:

uv run databricks-run path/to/script.py
uv run databricks-run path/to/notebook.db.py
uv run databricks-run path/to/notebook.ipynb

The unified databricks-run command dispatches by file extension. Use --format to override detection:

uv run databricks-run --format db-py path/to/notebook.py

The previous entry points remain available as compatibility aliases:

uv run run path/to/script.py
uv run run-py path/to/script.py
uv run run-db-py path/to/notebook.py
uv run run-ipynb path/to/notebook.ipynb

Output is written to .cache/outputs/<stem>.<YYYYMMDDTHHMMSS>.output.md relative to the current working directory. Use --output-dir to override the directory. The --serverless flag is recognized but currently fails fast because the implemented runner backend uses the Databricks Command Execution API on classic all-purpose clusters. See serverless execution backend investigation for the required separate backend design.

Behavior notes

Behavior Details
Default output dir .cache/outputs/ (override with --output-dir DIR)
Output filename <stem>.<YYYYMMDDTHHMMSS>.output.md — timestamped, never overwritten
Default timeout 10 minutes per command/cell
Timeout handling Cluster command is cancelled; error written to output file
Exit code Exits with code 1 on error or timeout; code 0 on success
run-ipynb --inplace Writes cell outputs back into the notebook; backup at <path>.bak
databricks-run --serverless Recognized, but exits with an unsupported-backend error

6. Papermill Integration

papermill supports parameter injection for notebook pipelines. Use it with this kernel for parameterized remote execution on Databricks clusters.

Install papermill:

uv add papermill

Run a notebook with parameter injection:

papermill input.ipynb output.ipynb --kernel databricks \
  -p param1 value1 -p param2 value2

Do NOT use the --inplace flag with papermill. Papermill is designed to produce a new output notebook with injected parameters and captured cell outputs; --inplace overwrites the source notebook and defeats this purpose.

If the cluster is stopped, increase the startup timeout:

papermill input.ipynb output.ipynb --kernel databricks \
  --start_timeout 600 -p param1 value1

7. MCP Server Usage Pattern

jupyter-databricks-kernel can be used by an external MCP (Model Context Protocol) server as its Databricks execution dependency. This repository does not ship an MCP server, Databricks App, HTTP adapter, or MCP tool definition.

7.1. External Server Responsibilities

A companion MCP server can be deployed once per Databricks workspace. That server owns:

  • Workspace authentication and Service Principal credentials
  • MCP transport and tool definitions
  • HTTP routing, if the server exposes an HTTP adapter
  • Session storage and timeout policy
  • Any output file persistence outside the command result returned by this package

The server can import DatabricksExecutor from this package to run code on a configured all-purpose cluster. It should keep one executor per active client session when isolated command contexts are required.

7.2. Project Routing

If a companion server supports multiple workspaces, keep workspace routing in the companion server configuration. This package can also read a project-local .databricks/jupyter-databricks-kernel.json file as local routing metadata for mcp_profile and cluster_id.

The .databricks/ directory is also used by the Databricks CLI for local sync snapshots, bundle state, variable overrides, and generated artifacts. Keep this package-owned file limited to the top-level routing shape below, do not store companion-server state under CLI-managed subdirectories such as .databricks/bundle/, and do not rely on .databricks/ for files that must be synchronized to the cluster. This package's file synchronization excludes .databricks/ to match Databricks CLI behavior.

Generic .databricks/config.json is not read. The package-owned filename avoids treating generic config.json as this package's claim inside the Databricks CLI local namespace.

Databricks CLI authentication remains in the normal Databricks configuration locations, such as ~/.databrickscfg or the CLI token cache, not in this project routing file.

Example external routing file at .databricks/jupyter-databricks-kernel.json:

{
  "mcp_profile": "databricks-prod",
  "cluster_id": "0123-456789-abcdef12"
}

Only two routing fields are read:

Field Effect
cluster_id Sets Config.cluster_id; the Jupyter kernel, runner CLI, or companion server uses it as the target all-purpose cluster for DatabricksExecutor.
mcp_profile Sets Config.mcp_profile; the AI agent or companion adapter can use it to select the named MCP server/workspace profile.

Configuration precedence is field-by-field:

  1. Environment variables: DATABRICKS_CLUSTER_ID and DATABRICKS_MCP_PROFILE
  2. ~/.databrickscfg from the active Databricks profile
  3. .databricks/jupyter-databricks-kernel.json

If the JSON above is present and no higher-priority value is set, Config.load() selects cluster 0123-456789-abcdef12 and profile databricks-prod. The runner commands and Jupyter kernel execute on that cluster, and an MCP-aware agent can call the companion server identified by databricks-prod.

If the JSON is absent, the same fields fall back to environment variables and then ~/.databrickscfg. If no source provides cluster_id, execution cannot choose a Databricks cluster until cluster_id is configured. If DATABRICKS_CLUSTER_ID=dev-cluster is set while the JSON contains 0123-456789-abcdef12, the environment value wins and dev-cluster is used.

This routing file does not configure authentication, store tokens, set workspace_url, create an MCP server, configure the Databricks CLI, configure Asset Bundles, or add files to cluster sync. Authentication remains in the normal Databricks SDK locations, such as environment variables, ~/.databrickscfg, or the CLI token cache. Bundle state remains under CLI-managed paths such as .databricks/bundle/, and this package excludes .databricks/ from synchronization.

In this pattern, mcp_profile maps to a named companion server entry in the AI agent's global config. This makes a project directory self-routing for the intended MCP/Jupyter workflow: switching projects can switch the companion MCP profile and cluster without changing local credentials or editing notebook code. Avoid duplicating workspace identity in both mcp_profile and workspace_url; choose one source of truth in the companion server.

7.3. Execution Flow

  1. The AI agent reads project routing from .databricks/jupyter-databricks-kernel.json.
  2. The AI agent calls the companion MCP server identified by mcp_profile.
  3. The companion server maps the request to a Databricks cluster.
  4. The companion server calls DatabricksExecutor from this package.
  5. DatabricksExecutor executes code through the Databricks Command Execution API and returns the result.
  6. The companion server or AI agent decides whether and where to persist the returned output.

8. Known Limitations

  • Serverless compute is not supported for the current kernel backend (Command Execution API limitation). See the serverless execution backend investigation for the current recommendation and possible complementary backend paths.
  • input() and interactive prompts do not work
  • Interactive widgets (ipywidgets) are not supported

9. Troubleshooting

9.1. Kernel feels slow

File sync may be uploading unnecessary files. Check your sync settings:

  1. Ensure .gitignore includes large/unnecessary files:

    .venv/
    __pycache__/
    *.pyc
    data/
    *.parquet
    node_modules/
    
  2. Add exclude patterns in pyproject.toml:

    [tool.jupyter-databricks-kernel.sync]
    exclude = ["data/", "models/", "*.csv"]
    
  3. Set size limits to catch unexpected large files:

    [tool.jupyter-databricks-kernel.sync]
    max_size_mb = 50.0
    max_file_size_mb = 10.0
    
  4. Disable sync entirely if not needed:

    [tool.jupyter-databricks-kernel.sync]
    enabled = false
    

10. Development

See CONTRIBUTING.md for development setup and guidelines.

11. License

Apache License 2.0

Metadata

Release files for jupyter-databricks-kernel 1.4.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for jupyter-databricks-kernel 1.4.1
File Size Uploaded
jupyter_databricks_kernel-1.4.1.tar.gz 282.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for jupyter-databricks-kernel 1.4.1
File Interpreter ABI Platform
jupyter_databricks_kernel-1.4.1-py3-none-any.whl Python 3 none any Details

Total release size: 323.6 kB

Release files / jupyter_databricks_kernel-1.4.1.tar.gz

Download URL jupyter_databricks_kernel-1.4.1.tar.gz
Size 282.0 kB
Tags Source
SHA-256 checksum
How to use checksums
f50c7137ed0dbfe0078a84ccc5565e67cfd757a83a70879e04de4deef76b1dce
BLAKE2b-256 checksum
How to use checksums
70f62f86585331f90eff32d6bae57b89a497b891fb28e7f388c77b699c1e8d5c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.13

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jun 3, 2026.

Transparency log

Release files / jupyter_databricks_kernel-1.4.1-py3-none-any.whl

Download URL jupyter_databricks_kernel-1.4.1-py3-none-any.whl
Size 41.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
54047da8ac89933f9446567bd1dde3701d1578893b30b17a6f430d56e82500e5
BLAKE2b-256 checksum
How to use checksums
8499bd38721833899f419627fc8271088e6ff60f06b01e9320e3239340541dfd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.13

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jun 3, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

1.4.1 This release

2 release files

1.4.0

2 release files

1.3.1

2 release files

1.3.0

2 release files

1.2.7

2 release files

1.2.6

2 release files

1.2.5

2 release files

1.2.4

2 release files

1.2.1

2 release files

1.2.0

2 release files

1.1.5

2 release files

1.1.4

2 release files

1.1.2

2 release files

1.1.1

2 release files

1.1.0

2 release files

1.0.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page