Skip to main content

VDK Lineage plugin collects lineage (input -> job -> output) information and send it to a pre-configured destination.

Project description

VDK Lineage

monthly download count for vdk-lineage

VDK Lineage plugin provides lineage data (input data -> job -> output data) information and send it to a pre-configured destination. The lineage data is send using OpenLineage standard

At POC level currently.

Currently, lineage data is collected

  • For each data job run/execution both start and end events including the status of the job (failed/succeeded)
  • For each execute query we collect input and output tables.

TODOs:

  • Collect status of the SQL query (failed, succeeded)
  • Create parent /child relationship between sql event and job run event to track them better (single job can have multiple queries)
  • Non-SQL lineage (ingest, load data,etc)
  • Extend support for all queries
  • provide more information using facets – op id, job version,
  • figure out how to visualize parent/child relationships in Marquez
  • Explore openlineage.sqlparser instead of sqllineage library as alternative

Usage

pip install vdk-lineage

And it will start collecting lineage from job and sql queries.

To send data using openlineage specify VDK_OPENLINEAGE_URL. For example:

export VDK_OPENLINEAGE_URL=http://localhost:5002
vdk marquez-server --start
vdk run some-job
# check UI for lineage
# stopping the server will delete any lineage data.
vdk marquez-server --stop

Build and testing

In order to build and test a plugin go to the plugin directory and use ../build-plugin.sh script to build it

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

vdk-lineage-0.3.1190994517.tar.gz (12.7 kB view details)

Uploaded Source

File details

Details for the file vdk-lineage-0.3.1190994517.tar.gz.

File metadata

  • Download URL: vdk-lineage-0.3.1190994517.tar.gz
  • Upload date:
  • Size: 12.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/5.0.0 CPython/3.10.13

File hashes

Hashes for vdk-lineage-0.3.1190994517.tar.gz
Algorithm Hash digest
SHA256 c65c2a564344a25193ee33f5e02364f1e09d9c60d8f04f49b7117262c328780b
MD5 70937c62d9c7b586005c5b95e26678c2
BLAKE2b-256 f03b529eb80fa4e256f37dc797d4b210cbfa7a76171e0286d1da5c58f5bd1e82

See more details on using hashes here.

Supported by

AWS AWS Cloud computing and Security Sponsor Datadog Datadog Monitoring Fastly Fastly CDN Google Google Download Analytics Microsoft Microsoft PSF Sponsor Pingdom Pingdom Monitoring Sentry Sentry Error logging StatusPage StatusPage Status page