Unified serving layer for non-text foundation models

These details have not been verified by PyPI

Project description

Sheaf

Unified serving layer for non-text foundation models.

vLLM solved inference for text LLMs by defining a standard compute contract and optimizing behind it. The same problem exists for every other class of foundation model — time series, tabular, molecular, geospatial, diffusion, audio — and nobody has solved it. Sheaf is that solution.

Each model type gets a typed request/response contract. Batching, caching, and scheduling are optimized per model type. Ray Serve is the substrate. Feast is a first-class input primitive.

In mathematics, a sheaf tracks locally-defined data that glues consistently across a space. Each model type defines its own local contract; Sheaf ensures they cohere into a unified serving layer.

Install

pip install sheaf-serve                           # core only
pip install "sheaf-serve[time-series]"            # + Chronos2 / TimesFM / Moirai
pip install "sheaf-serve[tabular]"                # + TabPFN
pip install "sheaf-serve[molecular]"              # + ESM-3  (Python 3.12+)
pip install "sheaf-serve[genomics]"               # + Nucleotide Transformer
pip install "sheaf-serve[small-molecule]"         # + MolFormer
pip install "sheaf-serve[materials]"              # + MACE-MP
pip install "sheaf-serve[audio]"                  # + Whisper / faster-whisper
pip install "sheaf-serve[audio-generation]"       # + MusicGen
pip install "sheaf-serve[tts]"                    # + Bark
pip install "sheaf-serve[vision]"                 # + DINOv2 / OpenCLIP / SAM2 / Depth Anything / DETR
pip install "sheaf-serve[earth-observation]"      # + Prithvi
pip install "sheaf-serve[weather]"                # + GraphCast
pip install "sheaf-serve[feast]"                  # + Feast feature store integration
pip install "sheaf-serve[modal]"                  # + Modal serverless deployment
pip install "sheaf-serve[batch]"                  # + offline batch inference (Ray Data)
pip install "sheaf-serve[all]"                    # everything

Quickstart

Direct backend inference:

from sheaf.api.time_series import Frequency, OutputMode, TimeSeriesRequest
from sheaf.backends.chronos import Chronos2Backend

backend = Chronos2Backend(model_id="amazon/chronos-bolt-tiny", device_map="cpu")
backend.load()

req = TimeSeriesRequest(
    model_name="chronos-bolt-tiny",
    history=[312, 298, 275, 260, 255, 263, 285, 320,
             368, 402, 421, 435, 442, 438, 430, 425],
    horizon=12,
    frequency=Frequency.HOURLY,
    output_mode=OutputMode.QUANTILES,
    quantile_levels=[0.1, 0.5, 0.9],
)

response = backend.predict(req)
# response.mean, response.quantiles

Ray Serve (production, autoscaling):

from sheaf import ModelServer
from sheaf.spec import ModelSpec, ResourceConfig
from sheaf.api.base import ModelType

server = ModelServer(models=[
    ModelSpec(
        name="chronos",
        model_type=ModelType.TIME_SERIES,
        backend="chronos2",
        backend_kwargs={"model_id": "amazon/chronos-bolt-small"},
        resources=ResourceConfig(num_gpus=1),
    ),
])
server.run()  # POST /chronos/predict, GET /chronos/health

Feast feature store (resolve features at request time):

# ModelSpec wires Feast — no history needed in the request
spec = ModelSpec(
    name="chronos",
    model_type=ModelType.TIME_SERIES,
    backend="chronos2",
    feast_repo_path="/feast/feature_repo",
)

# Client sends feature_ref instead of raw history
{
    "model_type": "time_series",
    "model_name": "chronos",
    "feature_ref": {
        "feature_view": "asset_prices",
        "feature_name": "close_history_30d",
        "entity_key": "ticker",
        "entity_value": "AAPL"
    },
    "horizon": 7,
    "frequency": "1d"
}

Modal (serverless, zero-infra):

from sheaf import ModalServer

server = ModalServer(models=[spec], app_name="my-sheaf", gpu="A10G")
app = server.app  # modal deploy my_server.py

Typed Python client:

from sheaf.client import SheafClient
from sheaf.api.time_series import Frequency, TimeSeriesRequest

with SheafClient(base_url="http://localhost:8000") as client:
    resp = client.predict(
        "chronos",
        TimeSeriesRequest(
            model_name="chronos",
            history=[1.0, 2.0, 3.0, 4.0, 5.0],
            horizon=3,
            frequency=Frequency.HOURLY,
        ),
    )
# resp is a typed TimeSeriesResponse — same Pydantic class the server returned
print(resp.mean)

AsyncSheafClient is the async-mirror; client.stream(deployment, request) yields SSE events for streaming backends like FLUX.

See examples/ for time series comparison, tabular, audio, vision, and the Feast feature store quickstart.

Supported model types

Type	Status	Backends
Time series	✅ v0.1	Chronos2, Chronos-Bolt, TimesFM, Moirai
Tabular	✅ v0.1	TabPFN v2
Audio transcription	✅ v0.3	Whisper, faster-whisper
Audio generation	✅ v0.3	MusicGen
Text-to-speech	✅ v0.3	Bark
Vision embeddings	✅ v0.3	OpenCLIP, DINOv2
Segmentation	✅ v0.3	SAM2
Depth estimation	✅ v0.3	Depth Anything v2
Object detection	✅ v0.3	DETR / RT-DETR
Protein / molecular	✅ v0.3	ESM-3 (Python 3.12+)
Genomics	✅ v0.3	Nucleotide Transformer
Small molecule	✅ v0.3	MolFormer-XL
Materials science	✅ v0.3	MACE-MP-0
Earth observation	✅ v0.3	Prithvi (IBM/NASA)
Weather forecasting	✅ v0.3	GraphCast
Cross-modal embeddings	✅ v0.3	ImageBind (text, vision, audio, depth, thermal)
Feast feature store	✅ v0.3	Any Feast online store (SQLite, Redis, DynamoDB, …)
Modal serverless	✅ v0.3	`ModalServer` — zero-infra GPU deployment
Diffusion / image gen	✅ v0.4	FLUX (schnell, dev)
Video understanding	✅ v0.4	VideoMAE, TimeSformer
LiDAR / 3D point cloud	✅ v0.5	PointNet (pure PyTorch; embed + ModelNet40 classify)
Pose estimation	✅ v0.5	ViTPose (COCO 17-keypoint, optional person bboxes)
Optical flow	✅ v0.5	RAFT (raft_large / raft_small via torchvision)
Multimodal generation	✅ v0.5	SDXL img2img + inpainting
Speech synthesis	✅ v0.5	Kokoro (voice + speed per request)
Offline batch inference	✅ v0.6	`BatchRunner` (Ray Data; tasks + actor-pool modes)
Async-job worker	✅ v0.7	`SheafWorker` (Redis Streams; pluggable queue/result ABCs)
LoRA adapter multiplexing	✅ v0.8	FLUX, SDXL via `ModelSpec.lora` (local paths + HF Hub sources)

Roadmap to production

v0.2 — serving layer (complete)

Ray Serve integration tested end-to-end
Async predict() handlers
HTTP API with proper request validation (422 on bad input)
Health check and readiness probe endpoints
Batching scheduler (BatchPolicy wired into @serve.batch per deployment)
Error handling at the service boundary (backend exceptions → structured HTTP 500)
Model hot-swap without restart (ModelServer.update())
Container-friendly auth for TabPFN v2 (TABPFN_TOKEN env var)

v0.3 — model types + integrations (complete)

ESM-3 protein embeddings
Nucleotide Transformer genomics embeddings
MolFormer-XL small molecule embeddings
MACE-MP-0 materials (energy, forces, stress)
Whisper / faster-whisper audio transcription
MusicGen audio generation
Bark text-to-speech
OpenCLIP image/text embeddings
DINOv2 image embeddings
SAM2 segmentation
Depth Anything v2 depth estimation
DETR / RT-DETR object detection
Prithvi earth observation embeddings
GraphCast weather forecasting
ImageBind cross-modal embeddings (text, vision, audio, depth, thermal)
Feast feature store integration (feature_ref in requests, FeastResolver, feast_repo_path on ModelSpec)
Modal serverless deployment (ModalServer — zero-infra alternative to Ray Serve)

v0.4 — generation + video (complete)

FLUX diffusion / image generation
VideoMAE / TimeSformer video understanding

v0.5 — observability + new modalities

Ops / DX:

PyPI publish (v0.4.0)
Prometheus metrics endpoint per deployment
Structured logging with request IDs end-to-end
OpenTelemetry traces through the request path

Serving / infra:

Streaming responses (POST /{name}/stream → SSE; FLUX emits per-step progress events)
Request caching (CacheConfig on ModelSpec — in-process LRU, optional TTL)
bucket_by batching — group requests by field value before @serve.batch

New model types:

LiDAR / 3D point cloud (PointNet — pure-PyTorch, no torch-geometric; embed + ModelNet40 classify; install with pip install 'sheaf-serve[lidar]')
Pose estimation (ViTPose — COCO 17-keypoint skeleton, optional person bboxes; install with pip install 'sheaf-serve[pose]')
Optical flow (RAFT — raft_large/raft_small via torchvision; (H, W, 2) float32 flow field; install with pip install 'sheaf-serve[optical-flow]')
Multimodal generation — text+image-conditioned (SDXL img2img + inpainting; install with pip install 'sheaf-serve[multimodal-generation]')
Speech synthesis with fine-grained control (Kokoro — voice + speed per request; install with pip install 'sheaf-serve[kokoro]')

v0.6 — offline batch inference (complete)

BatchRunner — same backend, same typed contract, offline batch mode; Ray Data map_batches substrate, stateless tasks with a worker-local backend cache so load() fires once per worker (not once per batch); install with pip install 'sheaf-serve[batch]'
BatchSpec — mirrors ModelSpec for backend selection; JsonlSource/JsonlSink in v1; new sources/sinks (S3, Parquet, Delta) slot in as additional BatchSource/BatchSink subclasses without changing the runner API
Actor-pool execution mode for warm loads on expensive backends (FLUX, GraphCast, SDXL) — opt-in via BatchSpec.compute="actors" + num_actors=N; load() runs once per actor at __init__ and persists for the actor's lifetime (#13)
Resumable checkpointing across process restarts (#12)

v0.7 — async-job queue (complete)

SheafWorker — queue-consumer pattern for long-running inference; v1 ships Redis Streams + consumer groups (horizontal scaling), pluggable JobQueue / ResultStore ABCs for SQS / Kafka follow-ups; install with pip install 'sheaf-serve[worker]'
Job lifecycle: enqueue → processing → result / dead-letter; at-least-once delivery via XACK-after-persist; per-job webhook on completion (best-effort POST)
Priority lanes + per-tenant fair queuing

v0.8 — LoRA adapter multiplexing (complete)

ModelSpec.lora = LoRAConfig(adapters={...}, default="...") — declare per-deployment adapter registry; one GPU deployment serves many fine-tunes
Per-request adapter selection via DiffusionRequest.adapters / MultimodalGenerationRequest.adapters (with optional adapter_weights for fusion)
First targets: FLUX (FLUX.1-schnell + FLUX.1-dev), SDXL (img2img + inpaint)
Local paths and HF Hub sources both supported (hf:org/repo[:weight_file] convention)
Bucket-by-resolved-adapter inside Ray Serve batch windows: set_active_adapters is called exactly once per homogeneous sub-batch
Hot-add adapters at runtime without ModelServer.update(spec) (deferred — adds VRAM-eviction / index-sync surface area)

v0.9 — typed Python client (complete)

Ships as sheaf.client inside sheaf-serve (not a separate sheaf-client PyPI package — schemas stay in one tree, no codegen, no drift). Splittable into its own package later if external client contributors arrive or install footprint becomes a real cost.

SheafClient (sync) + AsyncSheafClient (async, httpx-backed); typed predict(deployment, request) -> response against the discriminated AnyResponse union
health() / ready() helpers; structured exceptions (ValidationError for 422, ServerError for 5xx, ClientError for transport / decode failures)
SSE streaming via client.stream(deployment, request) async generator
RetryConfig with exponential backoff: configurable status codes, connection-error retry toggle, and max_attempts cap. Streams bypass retry by design (re-running yields interleaved progress events).
Server-side request_id (the UUID minted on the request) is attached to every raised SheafError subclass so callers can log-correlate without holding the original request object.
OpenAPI export via python -m sheaf.openapi --specs my_module:specs > openapi.json (or sheaf.openapi.generate(specs) programmatically) — backends are not loaded during generation, so it runs without GPU.

Architecture

┌─────────────────────────────────────────┐
│           API Layer                      │  typed contracts per model type
│  TimeSeriesRequest  TabularRequest  ...  │
├─────────────────────────────────────────┤
│         Scheduling Layer                 │  model-type-aware batching
│  BatchPolicy  RequestQueue               │
├─────────────────────────────────────────┤
│          Backend Layer                   │  pluggable execution + Ray Serve
│  ModelBackend  CacheManager  Feast       │
└─────────────────────────────────────────┘

Adding a new backend takes one class:

from sheaf.backends.base import ModelBackend
from sheaf.registry import register_backend

@register_backend("my-model")
class MyModelBackend(ModelBackend):
    def load(self) -> None:
        self._model = load_my_model()

    def predict(self, request):
        ...

    @property
    def model_type(self):
        return "time_series"

Contributing

Issues and PRs welcome. See CONTRIBUTING.md for development setup.

License

Apache 2.0

Project details

These details have not been verified by PyPI

Release history Release notifications | RSS feed

0.11.0

May 28, 2026

0.10.0

May 8, 2026

This version

0.9.0

May 7, 2026

0.8.0

May 7, 2026

0.7.0

Apr 30, 2026

0.6.0

Apr 20, 2026

0.5.1

Apr 19, 2026

0.5.0

Apr 17, 2026

0.4.0

Apr 16, 2026

0.3.0

Apr 16, 2026

0.1.0 yanked

Apr 15, 2026

Reason this release was yanked:

first release; missing the Python 3.11 floor that all subsequent versions require — caused silent broken installs on Python 3.10

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

sheaf_serve-0.9.0.tar.gz (842.7 kB view details)

Uploaded May 7, 2026 Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

The dropdown lists show the available interpreters, ABIs, and platforms. Enable javascript to be able to filter the list of wheel files.

sheaf_serve-0.9.0-py3-none-any.whl (161.5 kB view details)

Uploaded May 7, 2026 Python 3

File details

Details for the file sheaf_serve-0.9.0.tar.gz.

File metadata

Download URL: sheaf_serve-0.9.0.tar.gz
Upload date: May 7, 2026
Size: 842.7 kB
Tags: Source
Uploaded using Trusted Publishing? Yes
Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for sheaf_serve-0.9.0.tar.gz
Algorithm	Hash digest
SHA256	`b746d4e0317e6f363524d973f8c3c90a595d8da7bbedbce4b7ed48dc4385da5b`
MD5	`dfb97eb01364ca76a30e65a3611b7e55`
BLAKE2b-256	`8401cd57f8daa02932c9543007f19529738c643f8b1059a7044066583be9f600`

See more details on using hashes here.

Provenance

The following attestation bundles were made for sheaf_serve-0.9.0.tar.gz:

Publisher: publish.yml on korbonits/sheaf

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Statement:
- Statement type: https://in-toto.io/Statement/v1
- Predicate type: https://docs.pypi.org/attestations/publish/v1
- Subject name: sheaf_serve-0.9.0.tar.gz
- Subject digest: b746d4e0317e6f363524d973f8c3c90a595d8da7bbedbce4b7ed48dc4385da5b
- Sigstore transparency entry: 1458019527
- Sigstore integration time: May 7, 2026
Source repository:
- Permalink: korbonits/sheaf@4a3e3144d0219390d426e485842e5f08fe9c5647
- Branch / Tag: refs/tags/v0.9.0
- Owner: https://github.com/korbonits
- Access: public
Publication detail:
- Token Issuer: https://token.actions.githubusercontent.com
- Runner Environment: github-hosted
- Publication workflow: publish.yml@4a3e3144d0219390d426e485842e5f08fe9c5647
- Trigger Event: push

File details

Details for the file sheaf_serve-0.9.0-py3-none-any.whl.

File metadata

Download URL: sheaf_serve-0.9.0-py3-none-any.whl
Upload date: May 7, 2026
Size: 161.5 kB
Tags: Python 3
Uploaded using Trusted Publishing? Yes
Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for sheaf_serve-0.9.0-py3-none-any.whl
Algorithm	Hash digest
SHA256	`397f6d8fd69a56465d6db3e5472a103962b0d0038aac71d01a1217b2979baa5e`
MD5	`1865f69c05a7f7d957801f6b9223480d`
BLAKE2b-256	`ec8be8ef51caa0fb12366bcfe515e37c3a4d02a0ffc831e838d665172357b337`

See more details on using hashes here.

Provenance

The following attestation bundles were made for sheaf_serve-0.9.0-py3-none-any.whl:

Publisher: publish.yml on korbonits/sheaf

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Statement:
- Statement type: https://in-toto.io/Statement/v1
- Predicate type: https://docs.pypi.org/attestations/publish/v1
- Subject name: sheaf_serve-0.9.0-py3-none-any.whl
- Subject digest: 397f6d8fd69a56465d6db3e5472a103962b0d0038aac71d01a1217b2979baa5e
- Sigstore transparency entry: 1458019695
- Sigstore integration time: May 7, 2026
Source repository:
- Permalink: korbonits/sheaf@4a3e3144d0219390d426e485842e5f08fe9c5647
- Branch / Tag: refs/tags/v0.9.0
- Owner: https://github.com/korbonits
- Access: public
Publication detail:
- Token Issuer: https://token.actions.githubusercontent.com
- Runner Environment: github-hosted
- Publication workflow: publish.yml@4a3e3144d0219390d426e485842e5f08fe9c5647
- Trigger Event: push

sheaf-serve 0.9.0

Navigation

Verified details

Maintainers

Unverified details

Meta

Classifiers

Project description

Sheaf

Install

Quickstart

Supported model types

Roadmap to production

Architecture

Contributing

License

Project details

Verified details

Maintainers

Unverified details

Meta

Classifiers

Release history Release notifications | RSS feed

Download files

Source Distribution

Built Distribution

File details

File metadata

File hashes

Provenance

File details

File metadata

File hashes

Provenance