Skip to main content

Spark AI

Toolbox for building Generative AI applications on top of Apache Spark.

Many developers are companies are trying to leverage LLMs to enhance their existing applications or build completely new ones. Thanks to LLMs most of them no longer have to train new ML models. However, still the major challenge is data and infrastructure. This includes data ingestion, transformation, vectorization, lookup, and model serving.

Over the last few months, the industry has seen a spur of new tools and frameworks to help with these challenges. However, none of them are easy to use, deploy to production, nor can deal with the scale of data.

This project aims to provide a toolbox of Spark extensions, data sources, and utilities to make building robust data infrastructure on Spark for Generative AI applications easy.

PyPI version Maven Central

Example Applications

Complete examples that anyone can start from to build their own Generative AI applications.

Read about our thoughts on Prompt engineering, LLMs, and Low-code here.

Quickstart

Installation

Currently, the project is aimed mainly at PySpark users, however, because it also features high-performance connectors, both the PySpark and Scala dependencies have to be present on the Spark cluster.

Ingestion

from spark_ai.webapps.slack import SlackUtilities

# Batch version
slack = SlackUtilities(token='xoxb-...', spark=spark)
df_channels = slack.read_channels()
df_conversations = slack.read_conversations(df_channels)

# Live streaming version
df_messages = (spark.readStream
    .format('io.prophecy.spark_ai.webapps.slack.SlackSourceProvider')
    .option('token', 'xapp-...')
    .load())

Pre-processing & Vectorization

from spark_ai.llms.openai import OpenAiLLM
from spark_ai.dbs.pinecone import PineconeDB

OpenAiLLM(api_key='sk-...').register_udfs(spark=spark)
PineconeDB('8045...', 'us-east-1-aws').register_udfs(self.spark)

(df_conversations
    # Embed the text from every conversation into a vector
    .withColumn('embeddings', expr('openai_embed_texts(text)'))
    # Do some more pre-processing
    ... 
    # Upsert the embeddings into Pinecone
    .withColumn('status', expr('pinecone_upsert(\'index-name\', embeddings)'))
    # Save the status of the upsertion to a standard table
    .saveAsTable('pinecone_status'))

Inference

df_messages = spark.readStream \
    .format("io_prophecy.spark_ai.SlackStreamingSourceProvider") \
    .option("token", token) \
    .load()

# Handle a live stream of messages from Slack here

Roadmap

Data sources supported:

  • 🚧 Slack
  • 🗺️ PDFs
  • 🗺️ Asana
  • 🗺️ Notion
  • 🗺️ Google Drive
  • 🗺 Web-scrape

Vector databases supported:

  • 🚧 Pinecone
  • 🚧 Spark-ML (table store & cos sim)
  • 🗺 ElasticSearch

LLMs supported:

  • 🚧 OpenAI
  • 🚧 Spark-ML
  • 🗺️ Databrick's Dolly
  • 🗺️ HuggingFace's Models

Application interfaces supported:

  • 🚧 Slack
  • 🗺️ Microsoft Teams

And many more are coming soon (feel free to request as issues)! 🚀

✅: General Availability; 🚧: Beta availability; 🗺️: Roadmap;

Metadata

Release files for prophecy-spark-ai 0.1.14

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distribution (wheel)

Table of built distributions (wheels) for prophecy-spark-ai 0.1.14
File Interpreter ABI Platform
prophecy_spark_ai-0.1.14-py3-none-any.whl Python 3 none any Details

Release files / prophecy_spark_ai-0.1.14-py3-none-any.whl

Download URL prophecy_spark_ai-0.1.14-py3-none-any.whl
Size 15.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
420273d5aace1edb1623783efe101d2584e860d1efeb18b10c785df625fbe2c0
BLAKE2b-256 checksum
How to use checksums
1881d98854ad435d4fd9cd46c155115673395d14fe6a52d414c9929875629786
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.1.0 CPython/3.12.7

Release history Release notifications | RSS feed

This release

0.1.14 This release

1 release file

0.1.13

1 release file

0.1.12

1 release file

0.1.11

1 release file

0.1.10

1 release file

0.1.9

1 release file

0.1.8

1 release file

0.1.7

1 release file

0.1.6

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page