Skip to main content

supernova

supernova

Generate massive pre-embedded datasets, then load them into vector databases.

Overview

supernova has two pipelines:

  1. Embedding -- stream data from HuggingFace, embed with dense and/or sparse models, write parquet to S3
  2. Loading -- stream pre-embedded parquet from S3/HuggingFace into vector stores (Qdrant)

Both pipelines are streaming (never loads the full dataset into memory), pluggable (add new sources/embedders/stores by subclassing), and parallelizable (SkyPilot for distributed embedding and loading).

Quickstart

uv sync

# 1. Embed a dataset locally
nova embed configs/embedder/nick007x_arxiv_papers.yaml

# 2. Embed distributed across SkyPilot GPU pool
nova embed-dist configs/embedder/nick007x_arxiv_papers.yaml

# 3. Load into Qdrant
nova load configs/loader/ccnews_bge_large.yaml

# 4. Distributed loading (SkyPilot)
nova load-dist configs/loader/ccnews_bge_large.yaml

Project structure

supernova/
  sources/            # Data sources (HuggingFace)
  embedders/
    dense/            # Dense embedding backends (OpenAI, sentence-transformers)
    sparse/           # Sparse embedding backends (sentence-transformers SparseEncoder)
    engine.py         # EmbeddingEngine -- orchestrates dense/sparse/hybrid
    hybrid.py         # HybridEmbedder -- single forward pass for both
  storage/            # Output backends (S3, HuggingFace Hub, local)
  pipeline/           # Embedding orchestration (runner, worker, buffer)
  loader/
    datasource/       # Parquet readers (S3, HuggingFace)
    vectorstore/      # Vector store backends (Qdrant)
    runner.py         # Loading orchestration

configs/
  embedder/           # Embedding pipeline configs (single + distributed)
  loader/             # Loading pipeline configs (single + distributed)

scripts/
  run_embedder.py           # supernova CLI
  run_embed_distributed.py  # nova embed-dist CLI
  run_loader.py             # nova load CLI
  run_load_distributed.py   # nova load-dist CLI

Pipeline 1: Embedding

Configuration

source:
  type: huggingface
  dataset_name: nick007x/arxiv-papers
  split: train
  text_field: abstract

dense_embedder:
  type: sentence_transformer    # or openai
  model: Alibaba-NLP/gte-multilingual-base
  trust_remote_code: true
  batch_size: 64
  dtype: bfloat16

pipeline:
  chunk_size: 100000
  num_workers: 2

storage:
  type: s3                      # or hf, local
  bucket: qdrant--vectorforge
  prefix: arxiv-papers/gte-multilingual-base
  output_dir: /tmp/supernova

Sparse embeddings

Add a sparse_embedder section to produce sparse vectors alongside dense:

dense_embedder:
  type: sentence_transformer
  model: Alibaba-NLP/gte-multilingual-base
  trust_remote_code: true
  batch_size: 64
  dtype: bfloat16

sparse_embedder:
  type: sentence_transformer
  model: Alibaba-NLP/gte-multilingual-base
  batch_size: 64
  dtype: bfloat16

When both point to the same model, supernova automatically uses a hybrid encoder to minimize forward passes. You must specify at least one of dense_embedder or sparse_embedder.

Dense embedders

Type Config key Notes
OpenAI openai model, dimensions, batch_size, max_concurrent, base_url, api_key
Sentence Transformers sentence_transformer model, batch_size, dtype. Auto-detects CUDA/MPS/CPU

The OpenAI embedder supports any OpenAI-compatible API via base_url (llama.cpp, vLLM, Ollama, etc). Set api_key: none for local servers that don't require auth.

Storage backends

Type Config key Notes
S3 s3 bucket, prefix
HuggingFace Storage Buckets hf bucket_id, optional prefix, private. Writes to hf://buckets/{bucket_id}/...
Local local output_dir

Running locally

nova embed configs/embedder/nick007x_arxiv_papers.yaml

Running at scale with SkyPilot

SkyPilot pools create GPU workers and distribute embedding jobs across them. Workers are reused -- setup happens once, not per-slice.

# Preview the plan
nova embed-dist configs/embedder/nick007x_arxiv_papers.yaml --dry-run

# Run (default: A10G spot, autoscaling)
nova embed-dist configs/embedder/nick007x_arxiv_papers.yaml

# Custom parallelism
nova embed-dist configs/embedder/nick007x_arxiv_papers.yaml --num-jobs 20

Override resources in your config:

resources:
  accelerators: A10G:1
  cloud: aws
  use_spot: true

Output format

Parquet files with this schema:

Column Type Description
row_id int64 Auto-incrementing record ID
source_row_id int64 Original row in the source dataset
chunk_id int32 Pipeline batch / slice ID
chunk_index int32 Position within a text split (0 if not split)
text string The embedded text
dense_embedding list<float32> Dense embedding vector (when configured)
sparse_embedding struct{indices, values} Sparse embedding (when configured)

Query with DuckDB:

SELECT row_id, text[:80] AS preview, length(dense_embedding) AS dim
FROM 's3://qdrant--vectorforge/dataset/model/**/*.parquet'
LIMIT 10;

Pipeline 2: Loading

Configuration

vectors:                            # one entry per Qdrant vector name
  dense:
    type: dense                     # dense | sparse | multivector
    column: dense_embedding         # parquet column to read
    distance: cosine                # cosine | dot | euclid | manhattan

datasource:
  type: s3                          # s3 or huggingface
  bucket: qdrant--vectorforge
  prefix: stanford-oval--ccnews/baai_bge_large_en_v1.5
  id_column: row_id                 # default
  payload_fields:                   # what goes into the vector store payload
    text: text                      # payload key: parquet column name
    title: title

vectorstore:
  type: qdrant
  collection_name: ccnews-bge-large
  url: ${QDRANT_URL}                # env var substitution
  api_key: ${QDRANT_API_KEY}

loader:
  batch_size: 1000                  # points per upsert
  prefetch_size: 100000             # rows per DuckDB fetch (default: batch_size * 10)
  concurrency: 8                    # parallel upsert tasks

Running

nova load configs/loader/ccnews_bge_large.yaml

Datasources

Type Config key Notes
S3 s3 bucket, prefix. Streams via DuckDB httpfs
HuggingFace huggingface repo_id, optional subdir. Streams via DuckDB hf:// protocol

Vector stores

Type Config key Notes
Qdrant qdrant url, api_key, collection_name. Retry with backoff on timeouts

How it works

  1. DuckDB streams parquet data in large prefetch chunks (minimizes S3 round trips)
  2. Chunks are sliced into upsert-sized batches and written concurrently via asyncio
  3. Deferred indexing -- HNSW construction is disabled during load, then enabled for one efficient batch build
  4. Failed upserts are retried with exponential backoff

Distributed loading with SkyPilot

For terabyte-scale datasets, fan out across SkyPilot spot instances:

nova load-dist configs/loader/ccnews_bge_large.yaml
nova load-dist configs/loader/ccnews_bge_large.yaml --dry-run
nova load-dist configs/loader/ccnews_bge_large.yaml --num-shards 20

Environment variables

Variable Required for
OPENAI_API_KEY OpenAI embedder
HF_TOKEN HuggingFace Hub storage / datasource
AWS_ACCESS_KEY_ID S3 storage / datasource
AWS_SECRET_ACCESS_KEY S3 storage / datasource
AWS_SESSION_TOKEN S3 with AWS SSO
QDRANT_URL Qdrant vector store
QDRANT_API_KEY Qdrant vector store

Tests

uv run pytest tests/ -v

Documentation

Metadata

Release files for supernova 0.1.5

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for supernova 0.1.5
File Size Uploaded
supernova-0.1.5.tar.gz 113.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for supernova 0.1.5
File Interpreter ABI Platform
supernova-0.1.5-py3-none-any.whl Python 3 none any Details

Total release size: 252.8 kB

Release files / supernova-0.1.5.tar.gz

Download URL supernova-0.1.5.tar.gz
Size 113.7 kB
Tags Source
SHA-256 checksum
How to use checksums
f4a484b1e1699eedc4eef094b71029ae7b8c6f05b571403f5ae9100aca2f8a7c
BLAKE2b-256 checksum
How to use checksums
d05538769ddfb53c05cd8598f35632fc8e8d50e99a745a05eb12be887dd8c1bc
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jun 17, 2026.

Transparency log

Release files / supernova-0.1.5-py3-none-any.whl

Download URL supernova-0.1.5-py3-none-any.whl
Size 139.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
1543b84bc4be750ee8829e2a424da7fe04d66609802bd64dcfdef30ca3b6826d
BLAKE2b-256 checksum
How to use checksums
c094c1ff47fd65a574e335d595adad715afaaeca44de797e5c8d436bc068ec80
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jun 17, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.5 This release

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page