Skip to main content

SageMakerStudioDataEngineeringSessions

SageMaker Unified Studio Data Engineering Sessions

This pacakge depends on SageMaker Unified Studio environment, if you are using SageMaker Unified Studio, see AWS Doc for guidance.

This package contains functionality to support SageMaker Unified Studio connecting to various AWS Compute including EMR/EMR Serverless/Glue/Redshift etc.

It is utilizing ipython magics and AWS DataZone Connections to achieve the following features.

Features

  • Connect to remote compute
  • Execute Spark code in remote compute in Python/Scala
  • Execute SQL queries in remote compute
  • Send local variables to remote compute

How to setup

If you are using SageMaker Unifed Studio, you can skip this part, SageMaker Unifed Studio already set up the package.

This package contains various Jupyter Magics to achieve its functionality.

To load these magics, make sure you have iPython config file generated. If not, you could run ipython profile create, then a file with path ~/.ipython/profile_default/ipython_config.py should be generated

Then you will need to add the following line in the end of that config file

c.InteractiveShellApp.extensions.extend(['sagemaker_studio_dataengineering_sessions.sagemaker_connection_magic'])

Once that is finished, you could restart the ipython kernel and run %help to see a list of supported magics

Interactive vs background session

This packages uses SM_INPUT_NOTEBOOK_NAME environment variable to determine if the execution is through interactive or background session. See sagemaker_studio_dataengineering_sessions/sagemaker_database_session_manager/redshift/redshift_session.py file for usage.

Examples

To connect to remote compute, a DataZone Connection is required, you could create it via CreateConnection API, Let's say there's an existing connection called project.spark.

Supported Connection Type:

  • IAM
  • SPARK
  • REDSHIFT
  • ATHENA

Connect to remote compute and Execute Spark Code in Python

The following example will connect to AWS Glue Interactive session and run the spark code in Glue.

%%pyspark project.spark

import sys
import boto3
from awsglue.utils import getResolvedOptions
from pyspark.context import SparkContext
from pyspark.sql import SparkSession
from pyspark.sql.functions import col

args = getResolvedOptions(sys.argv, ["redshift_url", "redshift_iam_role", "redshift_tempdir","redshift_jdbc_iam_url"])
print(f"{args}")

sc = SparkContext.getOrCreate()
spark = SparkSession(sc)

df = spark.read.csv(f"s3://sagemaker-example-files-prod-{boto3.session.Session().region_name}/datasets/tabular/dirty-titanic/", header=True)
df.show(5, truncate=False)
df.printSchema()

df.createOrReplaceTempView("df_sql_tempview")

Execute Spark Code in Scala

The following example will connect to AWS Glue Interactive session and run the spark code in Scala.

%%scalaspark project.spark
val dfScala = spark.sql("SELECT count(0) FROM df_sql_tempview")
dfScala.show()

Execute SQL query in remote compute

The following example will connect to AWS Glue Interactive session and run the spark code in Scala.

%%sql project.redshift
select current_user()

Some other helpful magics

%help - list available magics and related information

%send_to_remote - send local variable to remote compute

%%configure - configure spark application config in remote compute

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

File details

Details for the file sagemaker_studio_dataengineering_sessions-1.3.25.tar.gz.

File metadata

File hashes

Hashes for sagemaker_studio_dataengineering_sessions-1.3.25.tar.gz
Algorithm Hash digest
SHA256 3bcfc28735bca1d1b192dc00a4dd5f8378191d9eb6a5bd1c7ab6f2d3f6e9c36c
MD5 26b2840c312deab606b292c33af15d04
BLAKE2b-256 da5b47ce35c02b77c3d80f648a460e6922b85112041707b6e15da1c2ad6b2a1c

See more details on using hashes here.

File details

Details for the file sagemaker_studio_dataengineering_sessions-1.3.25-py3-none-any.whl.

File metadata

File hashes

Hashes for sagemaker_studio_dataengineering_sessions-1.3.25-py3-none-any.whl
Algorithm Hash digest
SHA256 5694931b70ccae840a432ce9b0b9ad3c934e29cd735fcd124a7d03e2472856cf
MD5 d32e034fa3c9fdc8e29b529f6641a68f
BLAKE2b-256 5be7aa2641b5da03737bfa360a5dea7230a736cbd5ceb3ce5b639390caa85df3

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

1.3.25 This release

2 files

1.3.24

2 files

1.3.23

2 files

1.3.22

2 files

1.3.21

2 files

1.3.20

2 files

1.3.19

2 files

1.3.18

2 files

1.3.17

2 files

1.3.16

2 files

1.3.14

2 files

1.3.13

2 files

1.3.12

2 files

1.3.11

2 files

1.3.9

2 files

1.3.7

2 files

1.3.6

2 files

1.3.5

2 files

1.3.4

2 files

1.2.6

2 files

1.2.5

2 files

1.2.4

2 files

1.2.3

2 files

1.2.2

2 files

1.2.1

2 files

1.2.0

2 files

1.1.8

2 files

1.1.7

2 files

1.1.5

2 files

1.1.4

2 files

1.1.2

2 files

1.1.1

2 files

1.1.0

2 files

1.0.13

2 files

1.0.12

2 files

1.0.11

2 files

1.0.10

2 files

1.0.9

2 files

1.0.7

2 files

1.0.6

2 files

1.0.4

2 files

1.0.3

2 files

1.0.2

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page