Apache Airflow connector for Ocean for Apache Spark

These details have been verified by PyPI

Project links

Repository

GitHub Statistics

Maintainers

ImpSy sigmarkarl spotinst z4ck404

These details have not been verified by PyPI

Project links

Homepage

Project description

Airflow connector to Ocean for Apache Spark

An Airflow plugin and provider to launch and monitor Spark applications on Ocean for Apache Spark.

Installation

pip install ocean-spark-airflow-provider

Usage

For general usage of Ocean for Apache Spark, refer to the official documentation.

Setting up the connection

In the connection menu, register a new connection of type Ocean for Apache Spark. The default connection name is ocean_spark_default. You will need to have:

The Ocean Spark cluster ID of the cluster you just created (of the format osc-e4089a00). You can find it in the Spot console in the list of clusters, or by using the Cluster List API.
A Spot token to interact with the Spot API.

connection setup dialog

The Ocean for Apache Spark connection type is not available for Airflow 1, instead create an HTTP connection and fill your cluster id as host, and your API token as password.

You will need to create a separate connection for each Ocean Spark cluster that you want to use with Airflow. In the OceanSparkOperator, you can select which Ocean Spark connection to use with the connection_name argument (defaults to ocean_spark_default). For example, you may choose to have one Ocean Spark cluster per environment (dev, staging, prod), and you can easily target an environment by picking the correct Airflow connection.

Using the Spark operator

from ocean_spark.operators import OceanSparkOperator

# DAG creation

spark_pi_task = OceanSparkOperator(
    job_id="spark-pi",
    task_id="compute-pi",
    dag=dag,
    config_overrides={
        "type": "Scala",
        "sparkVersion": "3.2.0",
        "image": "gcr.io/datamechanics/spark:platform-3.2-latest",
        "imagePullPolicy": "IfNotPresent",
        "mainClass": "org.apache.spark.examples.SparkPi",
        "mainApplicationFile": "local:///opt/spark/examples/jars/examples.jar",
        "arguments": ["10000"],
        "driver": {
            "cores": 1,
            "spot": false
        },
        "executor": {
            "cores": 4,
            "instances": 1,
            "spot": true,
            "instanceSelector": "r5"
        },
    },
)

Using the Spark Connect operator (available since airflow 2.6.2)

from airflow import DAG, utils
from ocean_spark.operators import (
    OceanSparkConnectOperator,
)

args = {
    "owner": "airflow",
    "depends_on_past": False,
    "start_date": utils.dates.days_ago(0, second=1),
}


dag = DAG(dag_id="spark-connect-task", default_args=args, schedule_interval=None)

spark_pi_task = OceanSparkConnectOperator(
    task_id="spark-connect",
    dag=dag,
)

Trigger the DAG with config, such as

{
  "sql": "select random()"
}

more examples are available for Airflow 2.

Test locally

You can test the plugin locally using the docker compose setup in this repository. Run make serve_airflow at the root of the repository to launch an instance of Airflow 2 with the provider already installed.

Project details

These details have been verified by PyPI

Project links

Repository

GitHub Statistics

Maintainers

ImpSy sigmarkarl spotinst z4ck404

These details have not been verified by PyPI

Project links

Homepage

Release history Release notifications | RSS feed

This version

1.1.4

Sep 30, 2024

1.1.3

Sep 9, 2024

1.1.2

Sep 4, 2024

1.1.1

May 6, 2024

1.1.0

Jan 29, 2024

1.0.0

Sep 18, 2023

0.1.6

Aug 29, 2023

0.1.5

Mar 8, 2023

0.1.4 yanked

Mar 7, 2023

Reason this release was yanked:

Incorrect python version pinned

0.1.3

Jul 8, 2022

0.1.2

Feb 11, 2022

0.1.1

Feb 2, 2022

0.1.0

Jan 31, 2022

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

ocean_spark_airflow_provider-1.1.4.tar.gz (12.4 kB view details)

Uploaded Sep 30, 2024 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

ocean_spark_airflow_provider-1.1.4-py3-none-any.whl (16.5 kB view details)

Uploaded Sep 30, 2024 Python 3

File details

Details for the file ocean_spark_airflow_provider-1.1.4.tar.gz.

File metadata

Download URL: ocean_spark_airflow_provider-1.1.4.tar.gz
Upload date: Sep 30, 2024
Size: 12.4 kB
Tags: Source
Uploaded using Trusted Publishing? Yes
Uploaded via: twine/5.1.1 CPython/3.12.6

File hashes

Hashes for ocean_spark_airflow_provider-1.1.4.tar.gz
Algorithm	Hash digest
SHA256	`6db73eb59f77a4582e7a6933db444596420140596015b391ed16d6d147e4c863`
MD5	`d2eb865df1fea7db61c9fc2e81afdbb5`
BLAKE2b-256	`79d276989946167a053307dcd061fdff682aabec3182354be3df9d1df17a5ae8`

See more details on using hashes here.

File details

Details for the file ocean_spark_airflow_provider-1.1.4-py3-none-any.whl.

File metadata

Download URL: ocean_spark_airflow_provider-1.1.4-py3-none-any.whl
Upload date: Sep 30, 2024
Size: 16.5 kB
Tags: Python 3
Uploaded using Trusted Publishing? Yes
Uploaded via: twine/5.1.1 CPython/3.12.6

File hashes

Hashes for ocean_spark_airflow_provider-1.1.4-py3-none-any.whl
Algorithm	Hash digest
SHA256	`9dcd05da2cbd988c5c084e53da1fb9b201a71123167183d2c4c33f9bb48741d3`
MD5	`f93498c199400ee149d98890870f767d`
BLAKE2b-256	`079a2196110e185db76e3ece2a74875dfd37d2e9dbb96c919dd24edbb7978ce7`

See more details on using hashes here.

ocean-spark-airflow-provider 1.1.4

Navigation

Verified details

Project links

GitHub Statistics

Maintainers

Unverified details

Project links

Meta

Classifiers

Project description

Airflow connector to Ocean for Apache Spark

Installation

Usage

Setting up the connection

Using the Spark operator

Using the Spark Connect operator (available since airflow 2.6.2)

Trigger the DAG with config, such as

Test locally

Project details

Verified details

Project links

GitHub Statistics

Maintainers

Unverified details

Project links

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

File details

File metadata

File hashes