A utility library for working with data pipelines on GCP
Project description
mlslib
A lightweight utility library to simplify working with Google Cloud Storage and BigQuery on Google Cloud Platform (GCP). This library provides a set of high-level functions to streamline common data engineering and data science workflows.
🚀 Key Features
- Google Cloud Storage Integration: Upload pandas or Spark DataFrames to GCS
- File Management: Upload any local file (CSV, Parquet, Pickle, etc.) to GCS
- Public Access: Make GCS files public and get downloadable links
- BigQuery Integration: Query BigQuery tables directly into Spark DataFrames
- Notebook Display: Beautifully display PySpark DataFrames in Jupyter notebooks
- Data Sampling: Perform session-based sampling on pandas and Spark DataFrames
📦 Installation
Install mlslib directly from PyPI:
pip install mlslib
Dependencies
The library requires the following Python packages:
ipython>=7.0.0- For notebook display functionalitypyarrow>=6.0.0- For efficient data serialization
Note: The library's functions assume that google-cloud-storage and pyspark are installed and configured in your environment.
🔧 Setup
Before using mlslib, ensure you have:
- Google Cloud SDK installed and configured
- Authentication set up (service account key or gcloud auth)
- Required packages installed:
pip install google-cloud-storage pyspark
📖 Usage
Google Cloud Storage Utilities (gcs_utils)
Upload Local Files to GCS
from mlslib.gcs_utils import upload_file_to_gcs
# Upload any local file to GCS
gcs_uri = upload_file_to_gcs(
file_path="/path/to/your/local/model.pkl",
bucket_name="my-gcp-bucket",
gcs_path="models/model.pkl"
)
print(f"File uploaded to: {gcs_uri}")
Upload DataFrames to GCS
from mlslib.gcs_utils import upload_df_to_gcs
# Upload pandas DataFrame
gcs_path_pandas = upload_df_to_gcs(
my_pandas_df,
bucket_name="my-gcp-bucket",
gcs_path="data/pandas_export.parquet",
format="parquet"
)
# Upload Spark DataFrame
gcs_path_spark = upload_df_to_gcs(
my_spark_df,
bucket_name="my-gcp-bucket",
gcs_path="data/spark_export.csv",
format="csv"
)
Make Files Public and Get Download Links
from mlslib.gcs_utils import download_csv
# Make a GCS file public and get HTTPS download link
public_url = download_csv(
bucket_name="my-gcp-bucket",
file_path="data/public_file.csv"
)
print(f"Public URL: {public_url}")
BigQuery Utilities (bigquery_utils)
Load BigQuery Data into Spark DataFrames
from mlslib.bigquery_utils import load_bigquery_table_spark
# Load data from BigQuery table into Spark DataFrame
sql = "SELECT user_id, event_name FROM my_table WHERE event_date = '2025-06-22'"
df = load_bigquery_table_spark(
spark=spark, # Your SparkSession object
sql_query=sql,
table_name="my_table",
project_id="my-gcp-project",
dataset_id="my_analytics_dataset"
)
df.show()
Display Utilities (display_utils)
Beautiful DataFrame Display in Notebooks
from mlslib.display_utils import display_df
# Display Spark DataFrame as styled HTML table
display_df(df, limit_rows=50, title="User Events Preview")
Sampling Utilities (sampling_utils)
Session-Based Data Sampling
from mlslib.sampling_utils import sample_by_session
# Sample 1% of unique sessions from a DataFrame
sampled_df = sample_by_session(
df=my_dataframe,
session_column="user_session_id",
fraction=0.01,
seed=42
)
# Works with both pandas and Spark DataFrames
sampled_pandas = sample_by_session(pandas_df, "session_id", 0.05)
sampled_spark = sample_by_session(spark_df, "session_id", 0.05)
Date Utilities (date_utils)
Generate Periodic Date Ranges
from mlslib.date_utils import generate_periodic_date_ranges
# Generate 4 weekly (7-day) date ranges
weekly_batches = generate_periodic_date_ranges(
start_date_str="2025-07-01",
num_periods=4,
period_days=7
)
# Returns: [('2025-07-01', '2025-07-07'), ('2025-07-08', '2025-07-14'), ...]
print(weekly_batches)
Get Relative Date Ranges
from mlslib.date_utils import get_relative_day_range
# Get the date range for the last 30 days, ending yesterday
# Assuming today is 2025-06-29
last_30_days = get_relative_day_range(days=30, offset_days=-1)
# Returns: ('2025-05-30', '2025-06-28')
print(last_30_days)
📁 Project Structure
mlslib/
├── __init__.py # Package initialization and exports
├── gcs_utils.py # Google Cloud Storage utilities
├── bigquery_utils.py # BigQuery integration utilities
├── display_utils.py # Notebook display utilities
└── sampling_utils.py # Data sampling utilities
🔍 API Reference
gcs_utils Module
upload_file_to_gcs(file_path, bucket_name, gcs_path)- Upload local file to GCSupload_df_to_gcs(df, bucket_name, gcs_path, format='parquet')- Upload DataFrame to GCSdownload_csv(bucket_name, file_path)- Make GCS file public and get download URL
bigquery_utils Module
load_bigquery_table_spark(spark, sql_query, table_name, project_id, dataset_id)- Load BigQuery data into Spark DataFrame
display_utils Module
display_df(df, limit_rows=100, title=None)- Display Spark DataFrame in notebook
sampling_utils Module
sample_by_session(df, session_column, fraction, seed=42)- Perform session-based sampling on pandas or Spark DataFrames
date_utils Module
-
generate_periodic_date_ranges(start_date_str, num_periods, period_days)- Generate sequential date ranges of N days each. -
get_relative_day_range(days, offset_days, base_date_str)- Get date range for N days relative to a base date.
🤝 Contributing
Contributions are welcome! Please feel free to submit a Pull Request. For major changes, please open an issue first to discuss what you would like to change.
📄 License
This project is licensed under the MIT License - see the LICENSE file for details.
👨💻 Author
Raj Jha - rjha4@wayfair.com
🔗 Links
- PyPI: https://pypi.org/project/mlslib/
- Repository: https://github.com/wayfair-sandbox/mlslib
Note: This library is designed to work with Google Cloud Platform services. Make sure you have proper authentication and permissions set up before using these utilities.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file mlslib-0.1.10.tar.gz.
File metadata
- Download URL: mlslib-0.1.10.tar.gz
- Upload date:
- Size: 11.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.10.17
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
6900ba5fe97fe2acbb3e86b48e2fdff644f6c4f99483f37685bddd95bb88891b
|
|
| MD5 |
0373babceb52fc7892ae48924f83e99a
|
|
| BLAKE2b-256 |
93b272d25315d3e7a9dbd21a427af1cd986f11db6257ca600985aacef47e0de7
|
File details
Details for the file mlslib-0.1.10-py3-none-any.whl.
File metadata
- Download URL: mlslib-0.1.10-py3-none-any.whl
- Upload date:
- Size: 10.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.10.17
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
6737ed40515bf1cb72fe8e3b32b1a13beed95ea6a2aa06e2b40e9cf52ae66e4a
|
|
| MD5 |
534c5966e8001f385fde7994395ea44b
|
|
| BLAKE2b-256 |
eaa07b238788efbc7f426b9a7f25564da2af179b40d5290a80147ba4e0c90fd2
|