Skip to main content

dbt-prune

dbt-prune is a pip-installable dbt-core companion CLI for finding dbt models, snapshots, and sources that are no longer referenced by anything in your current project, configured sibling projects, or dbt exposures.

Installation

pip install dbt-prune

Optional integrations stay lightweight by default:

pip install 'dbt-prune[gcs]'
pip install 'dbt-prune[datahub]'

Checking the installed version

dbt-prune --version

Configuration

By default dbt-prune looks for dbt_prune.config.yml in the dbt project root passed to --project-dir.

manifests:
  - name: other_project
    type: local
    config:
      path: ../other_project/target/manifest.json

datahub:
  enabled: false
  gms_server_env_var: DATAHUB_GMS_SERVER
  token_env_var: DATAHUB_TOKEN
  platform: bigquery
  env: PROD

whitelist:
  - my_model
  - important_snapshot

GCS manifest

Install the GCS extra and declare a gcs manifest entry:

pip install 'dbt-prune[gcs]'
manifests:
  - name: analytics
    type: gcs
    config:
      project_id: data-mdm-prod
      bucket_name: tron-dbt-artifacts-prod
      object_name: orchestration/analytics/manifest.json
    excluded_packages:
      - customers
      - customer_platform

A gs:// URI is accepted as shorthand, and project_id is optional (it then falls back to your default credentials' project):

manifests:
  - name: analytics
    type: gcs
    config:
      uri: gs://tron-dbt-artifacts-prod/orchestration/analytics/manifest.json

Unknown or misspelled config keys raise an error rather than being silently ignored.

Optional GCS config fields:

Field Description
credentials Path to a service-account JSON key file (uses ADC when omitted)
impersonate_service_account Service account email to impersonate

Package filtering

Each manifest entry supports excluded_packages and included_packages (mutually exclusive). Nodes whose package_name is in excluded_packages are dropped from the manifest before orphan evaluation; included_packages keeps only matching nodes.

Whitelisting models

Use the top-level whitelist to keep specific current-project models or snapshots from ever being reported as orphaned. Whitelisted assets never appear in dbt-prune ls output and are skipped by dbt-prune run:

whitelist:
  - my_model
  - important_snapshot

Entries are matched against the node name. When you inspect a whitelisted asset directly with -s/--select, dbt-prune reports it as not orphaned with the reason Whitelisted in dbt_prune.config.yml.

Optional manifests

Set optional: true on a manifest entry to allow it to fail without aborting the run:

manifests:
  - name: staging
    type: gcs
    optional: true
    config:
      project_id: my-project
      bucket_name: my-bucket
      object_name: staging/manifest.json

The config file stores the names of the environment variables only. The actual DataHub GMS server URL and token must come from the environment at runtime:

export DATAHUB_GMS_SERVER=https://datahub.example.com
export DATAHUB_TOKEN=...

How orphan detection works

dbt-prune only parses manifest.json files. It does not run dbt, compile models, or introspect a live warehouse.

It loads:

  • the current project's target/manifest.json
  • any configured external manifests from local paths or gs:// URIs
  • exposure dependencies from every loaded manifest

A current-project model, snapshot, or source is considered orphaned only when it has zero incoming references from any loaded manifest and is not referenced by any exposure. An unused source is one that no model in any loaded manifest selects via source(). If DataHub checking is enabled, downstream lineage consumers also keep a node off the orphan list.

Usage

Output

dbt-prune uses rich terminal output when stdout is a TTY, including colours, status spinners, and concise success/warning/error lines. Piped or redirected output is plain text with no ANSI escapes or emoji, and --format json remains machine-readable. Use group-level flags before the command: -v/--verbose for per-item detail, -q/--quiet for errors only, and --no-color (or NO_COLOR=1) to force plain output.

List all orphaned assets

dbt-prune ls --project-dir .

By default ls reports orphaned models, snapshots, and unused sources.

Filter by resource type

Restrict the listing to a single resource type with --resource-type (model, snapshot, or source):

dbt-prune ls --resource-type source

Inspect one specific model

dbt-prune ls --project-dir . -s my_model

When -s/--select is used, dbt-prune reports whether the named model or snapshot is orphaned and explains why.

Filter by package

Scope orphan detection to a specific dbt package within your project using -p/--package:

dbt-prune ls -p marketing_core

Combine with -s/--select to inspect a specific model within a package:

dbt-prune ls -p marketing_core -s some_model_name

The same flag works for run:

dbt-prune run -p marketing_core --dry-run

JSON output

dbt-prune ls --format json

Start a Copilot agent session to delete every orphan

dbt-prune run --project-dir .

Behavior:

  • scans all current-project models and snapshots by default
  • builds a single prompt listing every orphaned asset
  • starts one agent session via gh agent-task create "PROMPT"
  • the prompt instructs the agent to delete the model files and every properties.yml / schema.yml reference to them, then open a pull request
  • requires the GitHub CLI (gh) to be installed and authenticated

Dry-run the cleanup plan

dbt-prune run --dry-run

Restrict cleanup to one model

dbt-prune run -s my_model

Limit the number of assets in the prompt

dbt-prune run --limit 5

DataHub integration

When datahub.enabled: true, dbt-prune queries DataHub for downstream lineage before marking an otherwise unreferenced asset as orphaned.

Dataset URNs are built from each manifest node's materialized location: database.schema.alias (falling back to name when alias is omitted). Ensure the manifest's database / schema / alias values match how your warehouse assets are ingested into DataHub.

Use datahub.platform and datahub.env to match your DataHub dataset URN convention:

datahub:
  enabled: true
  gms_server_env_var: DATAHUB_GMS_SERVER
  token_env_var: DATAHUB_TOKEN
  platform: bigquery
  env: PROD
  extra_platforms: [dbt]   # also probed when the warehouse URN has no lineage

Building URNs from a production manifest

If your local manifest.json is compiled against a dev target, its database / schema / alias describe your dev tables, and DataHub will correctly report that those have no downstreams. Point location_manifest at another configured manifest (for example the prod artifact loaded from GCS) to build URNs from production locations instead:

manifests:
  - name: prod
    type: gcs
    config:
      project_id: data-mdm-prod
      bucket_name: tron-dbt-artifacts-prod
      object_name: orchestration/revshare/manifest.json

datahub:
  enabled: true
  platform: bigquery
  env: PROD
  location_manifest: prod

Nodes are matched by unique_id, falling back to name. Anything not found in the location manifest falls back to the current-project location.

Lineage is resolved with a single-hop DownstreamOf relationships query. One hop is all that is needed to know whether anything consumes the asset, and it avoids DataHub's maxRelations limit that a full multi-hop lineage walk hits on highly connected datasets. Deployments without that field fall back to searchAcrossLineage with maxHops: 1. Because dbt-managed assets are often ingested under both the warehouse platform and the dbt platform (as siblings), each URN in platform + extra_platforms is probed and the asset is kept if any of them reports downstreams. Set extra_platforms: [] to probe only the warehouse platform.

GraphQL errors returned by DataHub are surfaced as warnings (or raised with --strict) rather than being treated as "no downstreams".

If the required environment variables are missing or DataHub is unreachable, dbt-prune warns and continues unless you pass --strict.

GitHub Copilot coding-agent PR automation

dbt-prune run builds a single cleanup prompt covering every orphaned asset in dbt_prune.pr_agent and starts a Copilot coding-agent session with gh agent-task create. The module also still exposes a REST-based issue creator (GitHubCopilotPRAgent) for teams that prefer issue-driven workflows.

Development

python -m pip install -e '.[dev]'
ruff check .
pytest
mypy dbt_prune

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

dbt_prune-1.0.0.tar.gz (28.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

dbt_prune-1.0.0-py3-none-any.whl (24.1 kB view details)

Uploaded Python 3

File details

Details for the file dbt_prune-1.0.0.tar.gz.

File metadata

  • Download URL: dbt_prune-1.0.0.tar.gz
  • Upload date:
  • Size: 28.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.14.6

File hashes

Hashes for dbt_prune-1.0.0.tar.gz
Algorithm Hash digest
SHA256 e169ace33c40a2d0c2a3e8df474e431925f3f1bd71f09ec16f782d83a05238d9
MD5 31378fc9ce4b1112dc726c7dbf0915e6
BLAKE2b-256 a5d60ea6e7a82fcabc6074a12efa0fe88d155c48f8b78d7e4e5b0ac38b8ec213

See more details on using hashes here.

File details

Details for the file dbt_prune-1.0.0-py3-none-any.whl.

File metadata

  • Download URL: dbt_prune-1.0.0-py3-none-any.whl
  • Upload date:
  • Size: 24.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.14.6

File hashes

Hashes for dbt_prune-1.0.0-py3-none-any.whl
Algorithm Hash digest
SHA256 01181319fcd540ef9ee394b84978168e0ca61a7000e513060fe44f3453482579
MD5 abba63d52c6414a7c60a576ffb6e6246
BLAKE2b-256 89a76d93b1943ca018621c9323bf086e91bf5b64d4d9a165335bddaa0c17af2f

See more details on using hashes here.

Release history Release notifications | RSS feed

1.3.2

2 files

1.3.1

2 files

1.3.0

2 files

1.2.0

2 files

1.1.0

2 files

This release

1.0.0 This release

2 files

0.5.1

2 files

0.5.0

2 files

0.4.2

2 files

0.4.1

2 files

0.4.0

2 files

0.3.1

2 files

0.3.0

2 files

0.2.2

2 files

0.2.1

2 files

0.2.0

2 files

0.1.4

2 files

0.1.3

2 files

0.1.2

2 files

0.1.1

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page