Skip to main content

sference Python SDK

Installable package: sference-sdk (import: sference_sdk). Used by the sference CLI and your own automation.

Install

uv add sference-sdk

Fallback:

pip install sference-sdk

From a clone of this repo:

uv sync --package sference-sdk

Usage

Set SFERENCE_API_KEY, or pass api_key= to the client.

Completion window: async workloads use "24h" (the only supported value) on background responses (metadata.completion_window), streams (window=), and batches (window=). Sync realtime endpoints (/v1/chat/completions, /v1/messages, blocking /v1/responses) do not take a window.

./workload.jsonl

Batch APIs take a JSONL file: one JSON object per line. OpenAI-compatible lines include custom_id, method, url, and body (only custom_id + inner body are sent to POST /v1/batches; method/url are ignored). Content-only lines are {"content": "..."} (then pass model= on submit).

Inner body accepts chat completions (messages) or Responses (input, max_output_tokens, …). Responses fields are normalized to chat format at create and validated before enqueue. Invalid rows return HTTP 400 with requests[i] and optional custom_id.

Example workload.jsonl:

{"custom_id":"example-1","method":"POST","url":"/v1/chat/completions","body":{"model":"Qwen/Qwen3.6-35B-A3B","messages":[{"role":"user","content":"Say hello in exactly one word."}]}}
{"custom_id":"example-2","method":"POST","url":"/v1/chat/completions","body":{"model":"Qwen/Qwen3.6-35B-A3B","messages":[{"role":"system","content":"You reply with one short sentence only."},{"role":"user","content":"What is 2+2?"}]}}
{"custom_id":"example-3","method":"POST","url":"/v1/responses","body":{"model":"Qwen/Qwen3.6-35B-A3B","input":[{"role":"user","content":"Reply with one word."}],"max_output_tokens":32}}

Batches (sync)

Best for a fixed JSONL workload: one submit, poll until terminal, then fetch structured results or download JSONL via the API.

from sference_sdk import SferenceClient

client = SferenceClient(api_key="sk_...")

batch = client.submit_batch(
    input_file="./workload.jsonl",
    model="Qwen/Qwen3.6-35B-A3B",
    window="24h",
)
done = client.wait_for_completion(batch.id, poll_interval=2.0, timeout=3600.0)
results = client.get_results(done.id)
by_id = results.index_by_custom_id()
print(by_id["row-a"].completion_text)

Or in one call:

by_id = client.get_results_indexed(done.id)

Build chat rows without hand-assembling body.messages:

from sference_sdk.models import InferenceRequest

req = InferenceRequest.chat(
    custom_id="row-a",
    user_content="Summarize this.",
    system_content="One sentence only.",
    model="Qwen/Qwen3.6-35B-A3B",
    temperature=0,
)
batch = client.submit_batch(requests=[req], window="24h")

Use a model supported by your sference deployment.

OpenAI-compatible responses (sync)

Standalone or stream-associated jobs via POST /v1/responses. Keys need responses:read and responses:write (default on newly issued keys).

from sference_sdk import SferenceClient

client = SferenceClient(api_key="sk_...")

created = client.create_response(
    model="Qwen/Qwen3.6-35B-A3B",
    input=[{"role": "user", "content": "Hello"}],
    metadata={"completion_window": "24h"},
)
row = client.get_response(created.id)

For a stream, add stream_id inside metadata next to completion_window.

OpenAI Python SDK (openai package)

If you already use the official OpenAI client, point it at sference’s /v1 endpoint and the same API key (with responses:read and responses:write).

pip install openai
import asyncio
import os

from openai import AsyncOpenAI


async def main() -> None:
    client = AsyncOpenAI(
        base_url="https://api.sference.com/v1",
        api_key=os.environ["SFERENCE_API_KEY"],
    )

    response = await client.responses.create(
        model="Qwen/Qwen3.6-35B-A3B",
        input=[{"role": "user", "content": "Hello, world!"}],
        background=True,
    )
    # Poll GET /v1/responses/{id} until terminal; your openai version may expose
    # something like await client.responses.retrieve(response.id), or use
    # AsyncSferenceClient.get_response(response.id) with the same API key.


asyncio.run(main())

Metadata: to set completion_window or stream_id like the native SDK, pass them in the request body your openai version supports (for example metadata= on create, or extra_body={"metadata": {...}} if the helper does not list those fields yet).

Async client — batches

AsyncSferenceClient uses httpx.AsyncClient so batch polling can run alongside other async I/O without blocking threads.

Use case: You already know the full set of prompts (for example a JSONL file) and want one scheduled unit of work with a clear terminal state and bulk results.

Benefits: Simple lifecycle (submit → wait → fetch results), fits large static workloads and JSONL-heavy pipelines.

import asyncio

from sference_sdk import AsyncSferenceClient


async def main() -> None:
    async with AsyncSferenceClient(api_key="sk_...") as client:
        batch = await client.submit_batch(
            input_file="./workload.jsonl",
            model="Qwen/Qwen3.6-35B-A3B",
            window="24h",
        )
        done = await client.wait_for_completion(batch.id, poll_interval=2.0, timeout=3600.0)
        results = await client.get_results(done.id)
        print(results.status, results.output_url)


asyncio.run(main())

Async client — streams

Stream-associated jobs use create_response(..., metadata={"stream_id": ..., "completion_window": "24h"}). Consume completions with list_responses_events / iter_responses_events (optional stream_id, wait_ms long-poll; optional checkpoints align with CLI sference responses tail).

Use case: Work arrives over time, or you want one id to group many responses and observe completions as they land.

Benefits: Independent submits with aggregated progress, stream-level status in the API/UI, and efficient event tailing.

import asyncio

from sference_sdk import AsyncSferenceClient


async def main() -> None:
    async with AsyncSferenceClient(api_key="sk_...") as client:
        stream = await client.create_stream(name="sdk-demo", window="24h")
        await client.create_response(
            model="Qwen/Qwen3.6-35B-A3B",
            input=[{"role": "user", "content": "Hello"}],
            metadata={"stream_id": stream.id, "completion_window": "24h"},
        )
        async for ev in client.iter_responses_events(stream_id=stream.id, checkpoint=False):
            print(ev.completion_id, ev.status)


asyncio.run(main())

CLI

For sference batch … and sference stream … commands, see the CLI README.

Metadata

Release files for sference-sdk 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for sference-sdk 0.2.0
File Size Uploaded
sference_sdk-0.2.0.tar.gz 15.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for sference-sdk 0.2.0
File Interpreter ABI Platform
sference_sdk-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 35.9 kB

Release files / sference_sdk-0.2.0.tar.gz

Download URL sference_sdk-0.2.0.tar.gz
Size 15.5 kB
Tags Source
SHA-256 checksum
How to use checksums
ed2d7ce292817c649304aada8f16751c0395282e8dea1e2e47929d53f004e11c
BLAKE2b-256 checksum
How to use checksums
4ee1f97d3a9189da68e864439de2b5523695148c3113426a15b777e8cf7b95b8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 28, 2026.

Transparency log

Release files / sference_sdk-0.2.0-py3-none-any.whl

Download URL sference_sdk-0.2.0-py3-none-any.whl
Size 20.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
3f128acaae6ca4e1b810e6b69fd638d76db65dc9caa578ba3815969740dddd0e
BLAKE2b-256 checksum
How to use checksums
77af469f792feebe65a4e598b47ec35afdd8e9c27e5124da6221e16ef1bb8d91
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 28, 2026.

Transparency log

Release history Release notifications | RSS feed

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

This release

0.2.0 This release

2 release files

0.1.9

2 release files

0.1.8

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

0.0.9

2 release files

0.0.8

2 release files

0.0.7

2 release files

0.0.6

2 release files

0.0.5

2 release files

0.0.4

2 release files

0.0.3

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page