Skip to main content

sference Python SDK

Installable package: sference-sdk (import: sference_sdk). Used by the sference CLI and your own automation.

Install

uv add sference-sdk

Fallback:

pip install sference-sdk

From a clone of this repo:

uv sync --package sference-sdk

Usage

Set SFERENCE_API_KEY, or pass api_key= to the client.

Completion window: async workloads use "24h" (the only supported value) on background responses (metadata.completion_window), streams (window=), and batches (window=). Sync realtime endpoints (/v1/chat/completions, /v1/messages, blocking /v1/responses) do not take a window.

./workload.jsonl

Batch APIs take a JSONL file: one JSON object per line. OpenAI-compatible lines include custom_id, method, url, and body (only custom_id + inner body are sent to POST /v1/batches; method/url are ignored). Content-only lines are {"content": "..."} (then pass model= on submit).

Inner body accepts chat completions (messages) or Responses (input, max_output_tokens, …). Responses fields are normalized to chat format at create and validated before enqueue. Invalid rows return HTTP 400 with requests[i] and optional custom_id.

Example workload.jsonl:

{"custom_id":"example-1","method":"POST","url":"/v1/chat/completions","body":{"model":"Qwen/Qwen3.6-35B-A3B","messages":[{"role":"user","content":"Say hello in exactly one word."}]}}
{"custom_id":"example-2","method":"POST","url":"/v1/chat/completions","body":{"model":"Qwen/Qwen3.6-35B-A3B","messages":[{"role":"system","content":"You reply with one short sentence only."},{"role":"user","content":"What is 2+2?"}]}}
{"custom_id":"example-3","method":"POST","url":"/v1/responses","body":{"model":"Qwen/Qwen3.6-35B-A3B","input":[{"role":"user","content":"Reply with one word."}],"max_output_tokens":32}}

Batches (sync)

Best for a fixed JSONL workload: one submit, poll until terminal, then fetch structured results or download JSONL via the API.

from sference_sdk import SferenceClient

client = SferenceClient(api_key="sk_...")

batch = client.submit_batch(
    input_file="./workload.jsonl",
    model="Qwen/Qwen3.6-35B-A3B",
    window="24h",
)
done = client.wait_for_completion(batch.id, poll_interval=2.0, timeout=3600.0)
results = client.get_results(done.id)
by_id = results.index_by_custom_id()
print(by_id["row-a"].completion_text)

Or in one call:

by_id = client.get_results_indexed(done.id)

Build chat rows without hand-assembling body.messages:

from sference_sdk.models import InferenceRequest

req = InferenceRequest.chat(
    custom_id="row-a",
    user_content="Summarize this.",
    system_content="One sentence only.",
    model="Qwen/Qwen3.6-35B-A3B",
    temperature=0,
)
batch = client.submit_batch(requests=[req], window="24h")

Use a model supported by your sference deployment.

OpenAI-compatible responses (sync)

Standalone or stream-associated jobs via POST /v1/responses. Keys need responses:read and responses:write (default on newly issued keys).

from sference_sdk import SferenceClient

client = SferenceClient(api_key="sk_...")

created = client.create_response(
    model="Qwen/Qwen3.6-35B-A3B",
    input=[{"role": "user", "content": "Hello"}],
    metadata={"completion_window": "24h"},
)
row = client.get_response(created.id)

For a stream, add stream_id inside metadata next to completion_window.

Decisions (realtime classification)

POST /v1/decisions answers up to 64 questions about one state (text or any JSON value) in a single call. Pick a model with modality == "decisions" from list_models(). Question types: choice (pick one label), score (ordinal scale, index 0 = lowest) and noul (probability of true). Billed on input tokens only, including image tokens.

Both clients accept images=["data:image/png;base64,..."] or images=[{"content_type": "image/jpeg", "base64": "..."}] (also available as sference_sdk.DecisionImage). Images precede the state in array order and travel inline without separate uploads. Clef accepts up to 4 PNG/JPEG/WebP images, 4 MiB and 16 megapixels each, 8 MiB total decoded bytes, with a 13 MiB request body limit. Remote URLs are not accepted. Accepted images are EXIF-oriented and downscaled to at most 2,097,152 pixels before inference; usage counts the resulting image tokens. The 16-megapixel limit applies to the original uploaded image.

from sference_sdk import ChoiceQuestion, NoulQuestion, ScoreQuestion, SferenceClient

client = SferenceClient(api_key="sk_...")

decision = client.create_decision(
    model="Cloudflare/clef",
    state="I was charged twice this month, please refund me.",
    questions={
        "route": ChoiceQuestion(criteria={"billing": "Payments and invoices", "support": "Everything else"}),
        "urgency": ScoreQuestion(criteria=["Low", "Medium", "High"]),
        "refund": NoulQuestion(instructions="Does the customer ask for a refund?"),
    },
)
decision.answers["route"].choice      # "billing"
decision.answers["urgency"].score     # 0..2, expected value over the scale
decision.answers["refund"].noul       # probability of true

Questions may also be plain dicts ({"type": "noul"}). 429 means no decision capacity right now and is safe to retry; 504 means the deadline passed and nothing was charged.

OpenAI Python SDK (openai package)

If you already use the official OpenAI client, point it at sference’s /v1 endpoint and the same API key (with responses:read and responses:write).

pip install openai
import asyncio
import os

from openai import AsyncOpenAI


async def main() -> None:
    client = AsyncOpenAI(
        base_url="https://api.sference.com/v1",
        api_key=os.environ["SFERENCE_API_KEY"],
    )

    response = await client.responses.create(
        model="Qwen/Qwen3.6-35B-A3B",
        input=[{"role": "user", "content": "Hello, world!"}],
        background=True,
    )
    # Poll GET /v1/responses/{id} until terminal; your openai version may expose
    # something like await client.responses.retrieve(response.id), or use
    # AsyncSferenceClient.get_response(response.id) with the same API key.


asyncio.run(main())

Metadata: to set completion_window or stream_id like the native SDK, pass them in the request body your openai version supports (for example metadata= on create, or extra_body={"metadata": {...}} if the helper does not list those fields yet).

Async client — batches

AsyncSferenceClient uses httpx.AsyncClient so batch polling can run alongside other async I/O without blocking threads.

Use case: You already know the full set of prompts (for example a JSONL file) and want one scheduled unit of work with a clear terminal state and bulk results.

Benefits: Simple lifecycle (submit → wait → fetch results), fits large static workloads and JSONL-heavy pipelines.

import asyncio

from sference_sdk import AsyncSferenceClient


async def main() -> None:
    async with AsyncSferenceClient(api_key="sk_...") as client:
        batch = await client.submit_batch(
            input_file="./workload.jsonl",
            model="Qwen/Qwen3.6-35B-A3B",
            window="24h",
        )
        done = await client.wait_for_completion(batch.id, poll_interval=2.0, timeout=3600.0)
        results = await client.get_results(done.id)
        print(results.status, results.output_url)


asyncio.run(main())

Async client — streams

Stream-associated jobs use create_response(..., metadata={"stream_id": ..., "completion_window": "24h"}). Consume completions with list_responses_events / iter_responses_events (optional stream_id, wait_ms long-poll; optional checkpoints align with CLI sference responses tail).

Use case: Work arrives over time, or you want one id to group many responses and observe completions as they land.

Benefits: Independent submits with aggregated progress, stream-level status in the API/UI, and efficient event tailing.

import asyncio

from sference_sdk import AsyncSferenceClient


async def main() -> None:
    async with AsyncSferenceClient(api_key="sk_...") as client:
        stream = await client.create_stream(name="sdk-demo", window="24h")
        await client.create_response(
            model="Qwen/Qwen3.6-35B-A3B",
            input=[{"role": "user", "content": "Hello"}],
            metadata={"stream_id": stream.id, "completion_window": "24h"},
        )
        async for ev in client.iter_responses_events(stream_id=stream.id, checkpoint=False):
            print(ev.completion_id, ev.status)


asyncio.run(main())

CLI

For sference batch … and sference stream … commands, see the CLI README.

Metadata

Release files for sference-sdk 0.3.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for sference-sdk 0.3.2
File Size Uploaded
sference_sdk-0.3.2.tar.gz 17.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for sference-sdk 0.3.2
File Interpreter ABI Platform
sference_sdk-0.3.2-py3-none-any.whl Python 3 none any Details

Total release size: 40.2 kB

Release files / sference_sdk-0.3.2.tar.gz

Download URL sference_sdk-0.3.2.tar.gz
Size 17.4 kB
Tags Source
SHA-256 checksum
How to use checksums
9f4a9cf60705b86da64d0bae72c2b539c9c0caeaed5aeb9f10980724659e7945
BLAKE2b-256 checksum
How to use checksums
877178f602417bba050dc2f8ca217a267f937722a28e7a7679b931bcb71125f9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 10, 2026.

Transparency log

Release files / sference_sdk-0.3.2-py3-none-any.whl

Download URL sference_sdk-0.3.2-py3-none-any.whl
Size 22.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
efa34f2e18a3a2fb92fda3b45a6e3c82960fd9536f8decd3c33dc2a8c9817fed
BLAKE2b-256 checksum
How to use checksums
00b1e7cb8403df2da29e120e13b173d6b21f5a2841b8987ada61ea4c6386f3b8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 10, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.3.2 This release

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.9

2 release files

0.1.8

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

0.0.9

2 release files

0.0.8

2 release files

0.0.7

2 release files

0.0.6

2 release files

0.0.5

2 release files

0.0.4

2 release files

0.0.3

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page