OpenLineage integration with Airflow

Project description

OpenLineage Airflow Integration

A library that integrates Airflow DAGs with OpenLineage for automatic metadata collection.

Features

Metadata

Task lifecycle
Task parameters
Task runs linked to versioned code
Task inputs / outputs

Lineage

Track inter-DAG dependencies

Built-in

SQL parser
Link to code builder (ex: GitHub)
Metadata extractors

Requirements

Installation

$ pip3 install openlineage-airflow

Note: You can also add openlineage-airflow to your requirements.txt for Airflow.

To install from source, run:

$ python3 setup.py install

Setup

Airflow 2.3+

The integration automatically registers itself for Airflow 2.3 if it's installed on the Airflow worker's Python. This means you don't have to do anything besides configuring it, which is described in the Configuration section.

Airflow 2.1 - 2.2

This method has limited support: it does not support tracking failed jobs, and job starts are registered only when a job ends.

Set your LineageBackend in your airflow.cfg or via environmental variable AIRFLOW__LINEAGE__BACKEND to

openlineage.lineage_backend.OpenLineageBackend

In contrast to integration via subclassing a DAG, a LineageBackend-based approach collects all metadata for a task on each task's completion.

The OpenLineageBackend does not take into account manually configured inlets and outlets.

When enabled, the library will:

On DAG start, collect metadata for each task using an Extractor if it exists for a given operator.
Collect task input / output metadata (source, schema, etc.)
Collect task run-level metadata (execution time, state, parameters, etc.)
On DAG complete, also mark the task as complete in OpenLineage

Configuration

`HTTP` Backend Environment Variables

openlineage-airflow uses the OpenLineage client to push data to OpenLineage backend.

The OpenLineage client depends on environment variables:

OPENLINEAGE_URL - point to the service that will consume OpenLineage events.
OPENLINEAGE_API_KEY - set if the consumer of OpenLineage events requires a Bearer authentication key.
OPENLINEAGE_NAMESPACE - set if you are using something other than the default namespace for the job namespace.
OPENLINEAGE_AIRFLOW_DISABLE_SOURCE_CODE - set to False if you want the source code of callables provided in the PythonOperator to be sent in OpenLineage events.

For backwards compatibility, openlineage-airflow also supports configuration via MARQUEZ_URL, MARQUEZ_NAMESPACE and MARQUEZ_API_KEY variables.

MARQUEZ_URL=http://my_hosted_marquez.example.com:5000
MARQUEZ_NAMESPACE=my_special_ns

Extractors : Sending the correct data from your DAGs

If you do nothing, the OpenLineage backend will receive the Job and the Run from your DAGs, but, unless you use one of the few operators for which this integration provides an extractor, input and output metadata will not be sent.

openlineage-airflow allows you to do more than that by building "Extractors." An extractor is an object suited to extract metadata from a particular operator (or operators).

Name : The name of the task
Inputs : A list of input datasets
Outputs : A list of output datasets
Context : The Airflow context for the task

Bundled Extractors

openlineage-airflow provides extractors for:

PostgresOperator
MySqlOperator
AthenaOperator
BigQueryOperator
SnowflakeOperator
TrinoOperator
GreatExpectationsOperator
SFTPOperator
FTPFileTransmitOperator
PythonOperator
RedshiftDataOperator, RedshiftSQLOperator
SageMakerProcessingOperator, SageMakerProcessingOperatorAsync
SageMakerTrainingOperator, SageMakerTrainingOperatorAsync
SageMakerTransformOperator, SageMakerTransformOperatorAsync
S3CopyObjectExtractor, S3FileTransformExtractor
GCSToGCSOperator

SQL Operators utilize the SQL parser. There is an experimental SQL parser activated if you install openlineage-sql on your Airflow worker.

Custom Extractors

If your DAGs contain additional operators from which you want to extract lineage data, fear not - you can always provide custom extractors. They should derive from BaseExtractor.

There are two ways to register them for use in openlineage-airflow.

One way is to add them to the OPENLINEAGE_EXTRACTORS environment variable, separated by a semi-colon (;).

OPENLINEAGE_EXTRACTORS=full.path.to.ExtractorClass;full.path.to.AnotherExtractorClass

To ensure OpenLineage logging propagation to custom extractors you should use self.log instead of creating a logger yourself.

Default Extractor

When you own operators' code this is not neccessary to provide custom extractors. You can also use Default Extractor's capability.

In order to do that you should define at least one of two methods in operator:

get_openlineage_facets_on_start()

Extracts metadata on start of task.

get_openlineage_facets_on_complete(task_instance: TaskInstance)

Extracts metadata on complete of task. This should accept task_instance argument, similar to extract_on_complete method in base extractors.

If you don't define get_openlineage_facets_on_complete method it would fall back to get_openlineage_facets_on_start.

Great Expectations

The Great Expectations integration works by providing an OpenLineageValidationAction. You need to include it into your action_list in great_expectations.yml.

The following example illustrates a way to change the default configuration:

validation_operators:
  action_list_operator:
    # To learn how to configure sending Slack notifications during evaluation
    # (and other customizations), read: https://docs.greatexpectations.io/en/latest/autoapi/great_expectations/validation_operators/index.html#great_expectations.validation_operators.ActionListValidationOperator
    class_name: ActionListValidationOperator
    action_list:
      - name: store_validation_result
        action:
          class_name: StoreValidationResultAction
      - name: store_evaluation_params
        action:
          class_name: StoreEvaluationParametersAction
      - name: update_data_docs
        action:
          class_name: UpdateDataDocsAction
+     - name: openlineage
+       action:
+         class_name: OpenLineageValidationAction
+         module_name: openlineage.common.provider.great_expectations.action
      # - name: send_slack_notification_on_validation_result
      #   action:
      #     class_name: SlackNotificationAction
      #     # put the actual webhook URL in the uncommitted/config_variables.yml file
      #     slack_webhook: ${validation_notification_slack_webhook}
      #     notify_on: all # possible values: "all", "failure", "success"
      #     renderer:
      #       module_name: great_expectations.render.renderer.slack_renderer
      #       class_name: SlackRenderer

If you're using GreatExpectationsOperator, you need to set validation_operator_name to an operator that includes OpenLineageValidationAction. Setting it in great_expectations.yml files isn't enough - the operator overrides it with the default name if a different one is not provided.

To see an example of a working configuration, see DAG and Great Expectations configuration in the integration tests.

Triggering Child Jobs

Commonly, Airflow DAGs will trigger processes on remote systems, such as an Apache Spark or Apache Beam job. Those systems may have their own OpenLineage integrations and report their own job runs and dataset inputs/outputs. To propagate the job hierarchy, tasks must send their own run ids so that the downstream process can report the ParentRunFacet with the proper run id.

The lineage_run_id and lineage_parent_id macros exists to inject the run id or whole parent run information of a given task into the arguments sent to a remote processing job's Airflow operator. The macro requires the DAG run_id and the task to access the generated run_id for that task. For example, a Spark job can be triggered using the DataProcPySparkOperator with the correct parent run_id using the following configuration:

t1 = DataProcPySparkOperator(
    task_id=job_name,
    #required pyspark configuration,
    job_name=job_name,
    dataproc_pyspark_properties={
        'spark.driver.extraJavaOptions':
            f"-javaagent:{jar}={os.environ.get('OPENLINEAGE_URL')}/api/v1/namespaces/{os.getenv('OPENLINEAGE_NAMESPACE', 'default')}/jobs/{job_name}/runs/{{{{macros.OpenLineagePlugin.lineage_run_id(task, task_instance)}}}}?api_key={os.environ.get('OPENLINEAGE_API_KEY')}"
        dag=dag)

Secrets redaction

The integration uses Airflow SecretsMasker to hide secrets from produced metadata events. As not all fields in the metadata should be redacted, RedactMixin is used to pass information about which fields should be ignored by the process.

Typically, you should subclass RedactMixin and use the _skip_redact attribute as a list of names of fields to be skipped.

However, all facets inheriting from BaseFacet should use the _additional_skip_redact attribute as an addition to the regular list (['_producer', '_schemaURL']).

Development

To install all dependencies for local development:

The Airflow integration depends on openlineage.sql, openlineage.common and openlineage.client.python. You should install them first independently or try to install them with following command:

$ pip install -r dev-requirements.txt

There is also a bash script that can run an arbitrary Airflow image with an OpenLineage integration built from the current branch. Additionally, it mounts OpenLineage Python packages as Docker volumes. This enables you to change your code without the need to constantly rebuild Docker images to run tests. Run it as:

$ AIRFLOW_IMAGE=<airflow_image_with_tag> ./scripts/run-dev-airflow.sh [--help]

Unit tests

To run the entire unit test suite, use the below command:

$ tox

or choose one of the environments, e.g.:

$ tox -e py-airflow214

You can also skip using tox and run pytest on your own dev environment.

Integration tests

The integration tests require the use of docker compose. There are scripts prepared to make build images and run tests easier.

$ AIRFLOW_IMAGE=<name-of-airflow-image> ./tests/integration/docker/up.sh

$ AIRFLOW_IMAGE=apache/airflow:2.3.1-python3.7 ./tests/integration/docker/up.sh

When using run-dev-airflow.sh, you can add the -i flag or --attach-integration flag to run integration tests in a dev environment. This can be helpful when you need to run arbitrary integration tests during development. For example, the following command run in the integration container...

python -m pytest test_integration.py::test_integration[great_expectations_validation-requests/great_expectations.json]

...runs a single test which you can repeat after changes in code.

SPDX-License-Identifier: Apache-2.0
Copyright 2018-2023 contributors to the OpenLineage project

Project details

Release history Release notifications | RSS feed

1.13.1

Apr 25, 2024

1.12.0

Apr 9, 2024

1.11.3

Apr 4, 2024

1.11.2

Apr 4, 2024

1.11.1

Apr 4, 2024

1.10.2

Mar 15, 2024

1.10.1

Mar 14, 2024

1.10.0

Mar 14, 2024

1.9.1

Feb 26, 2024

1.9.0

Feb 23, 2024

1.8.0

Jan 22, 2024

1.7.0

Dec 21, 2023

1.6.2

Dec 7, 2023

1.6.1

Dec 7, 2023

1.6.0

Dec 5, 2023

1.5.0

Nov 2, 2023

1.4.1

Oct 9, 2023

1.3.1

Oct 3, 2023

1.3.0

Oct 3, 2023

1.2.2

Sep 20, 2023

1.2.1

Sep 19, 2023

1.2.0

Sep 14, 2023

1.1.0

Aug 23, 2023

1.0.0

Aug 1, 2023

0.30.1

Jul 25, 2023

0.30.0

Jul 25, 2023

0.29.2

Jun 30, 2023

0.28.0

Jun 12, 2023

0.27.2

Jun 6, 2023

0.27.1

Jun 5, 2023

0.26.0

May 18, 2023

0.25.0

May 15, 2023

0.24.0

May 3, 2023

0.23.0

Apr 20, 2023

0.22.0

Apr 3, 2023

0.21.1

Mar 2, 2023

This version

0.21.0

Mar 2, 2023

0.20.6

Feb 10, 2023

0.20.4

Feb 7, 2023

0.19.2

Jan 5, 2023

0.19.1

Jan 5, 2023

0.18.0

Dec 8, 2022

0.17.0

Nov 16, 2022

0.16.1

Nov 3, 2022

0.15.1

Oct 4, 2022

0.14.1

Sep 7, 2022

0.14.0

Sep 6, 2022

0.13.1

Aug 25, 2022

0.13.0

Aug 22, 2022

0.12.0

Aug 1, 2022

0.11.0

Jul 7, 2022

0.10.0

Jun 24, 2022

0.9.0

Jun 3, 2022

0.8.2

May 19, 2022

0.8.1

Apr 29, 2022

0.7.1

Apr 19, 2022

0.7.0

Mar 10, 2022

0.6.2

Mar 16, 2022

0.6.1

Mar 7, 2022

0.6.0

Mar 4, 2022

0.5.2

Feb 11, 2022

0.5.1

Jan 19, 2022

0.4.0

Dec 13, 2021

0.3.1

Oct 22, 2021

0.3.0

Oct 21, 2021

0.2.3

Oct 7, 2021

0.2.2

Sep 8, 2021

0.2.1

Aug 27, 2021

0.2.0

Aug 23, 2021

0.1.0

Aug 13, 2021

0.1.0rc5 pre-release

Aug 13, 2021

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

openlineage-airflow-0.21.0.tar.gz (40.5 kB view hashes)

Uploaded Mar 2, 2023 Source

Built Distribution

openlineage_airflow-0.21.0-py3-none-any.whl (53.3 kB view hashes)

Uploaded Mar 2, 2023 Python 3

Hashes for openlineage-airflow-0.21.0.tar.gz

Hashes for openlineage-airflow-0.21.0.tar.gz
Algorithm	Hash digest
SHA256	`51c7b8d15801c1fe5a813e80ab2f4a845111ed801f2eb4f01cbc11f31b8430d3`
MD5	`40b947c3f2a7bed15ab74f6b98ee0c0b`
BLAKE2b-256	`4969ed79bb6b69361dceb05b1d5ec67751a927278029c25712b7ca7bb4eac372`

Hashes for openlineage_airflow-0.21.0-py3-none-any.whl

Hashes for openlineage_airflow-0.21.0-py3-none-any.whl
Algorithm	Hash digest
SHA256	`199cc59d94360449def009663c76a6d5369400f9b9ae75036859d245255a5e97`
MD5	`a37ec47d343652e7a6c45bffd103fe6d`
BLAKE2b-256	`14c5868660c7c74a48b9dd1d679ea57999ec93a12c1c6b75506c89fae134f21a`

openlineage-airflow 0.21.0

Navigation

Verified details

Maintainers

Unverified details

GitHub Statistics

Meta

Project description

OpenLineage Airflow Integration

Features

Requirements

Installation

Setup

Airflow 2.3+

Airflow 2.1 - 2.2

Configuration

`HTTP` Backend Environment Variables

Extractors : Sending the correct data from your DAGs

Bundled Extractors

Custom Extractors

Default Extractor

Great Expectations

Triggering Child Jobs

Secrets redaction

Development

Unit tests

Integration tests

Project details

Verified details

Maintainers

Unverified details

GitHub Statistics

Meta

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

openlineage-airflow 0.21.0

Navigation

Verified details

Maintainers

Unverified details

GitHub Statistics

Meta

Project description

OpenLineage Airflow Integration

Features

Requirements

Installation

Setup

Airflow 2.3+

Airflow 2.1 - 2.2

Configuration

HTTP Backend Environment Variables

Extractors : Sending the correct data from your DAGs

Bundled Extractors

Custom Extractors

Default Extractor

Great Expectations

Triggering Child Jobs

Secrets redaction

Development

Unit tests

Integration tests

Project details

Verified details

Maintainers

Unverified details

GitHub Statistics

Meta

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

`HTTP` Backend Environment Variables