Skip to main content

Bespoke Nimble Python SDK

A lightweight Python client for hosted Bespoke Nimble: ask typed questions, get decisions and candidate probabilities.

  • Package: bespokelabs-nimble
  • Import: from bespokelabs.nimble import Nimble, AsyncNimble, Choice, Noul, Score
  • Python 3.10+; runtime dependencies are HTTPX and Pydantic. No model weights or GPU libraries.
  • Uses Nimble's existing POST /v1/systemone endpoint.

Configure the client with your Nimble deployment URL. The package does not assume a public production API URL.

Install

uv pip install bespokelabs-nimble
# Or add it to a project:
uv add bespokelabs-nimble
# Or use pip:
pip install bespokelabs-nimble

From a source checkout:

uv pip install -e ./nimble-sdk

Quickstart

Point the client at your hosted Nimble server. The URL may be the service root or end in /v1; gateway path prefixes are preserved.

export BESPOKE_NIMBLE_BASE_URL="https://your-nimble-host.example"
export BESPOKE_API_KEY="your-api-key"  # Only if your gateway uses bearer authentication.
from bespokelabs.nimble import Choice, Nimble, Noul, Score

with Nimble() as client:
    result = client.system_one(
        state="I was charged twice. Please refund the duplicate payment.",
        questions={
            "refund": Noul(instructions="Does the customer request a refund?"),
            "department": Choice(
                instructions="Which department should handle this request?",
                criteria={
                    "billing": "Charges, payments, and refunds",
                    "technical": "Software bugs and outages",
                },
            ),
            "urgency": Score(
                instructions="Assess operational urgency.",
                criteria=[
                    "Service works normally",
                    "Some functionality unavailable",
                    "Complete outage",
                ],
            ),
        },
    )

print(result.nouls["refund"].noul)  # P(true), from 0 to 1
print(result.choices["department"].choice)  # e.g. "billing"
print(result.choices["department"].probabilities)  # All candidate probabilities
print(result.scores["urgency"].score)  # Expected level, from 0 to 2
print(result.usage.input_tokens)

state and instructions can also be JSON objects or lists. The server serializes them as text for the model. Raw question dictionaries with the same shape as the HTTP API are accepted.

Question and answer types

Question Input Answer
Noul Yes/no instructions; optional definitions of true and false noul: probability of true
Choice 2–26 option keys mapped to descriptions or None choice, probabilities, confidence
Score 2–26 ordered rubric descriptions, lowest first score, legend, probabilities, confidence

Noul preserves the API's name for a yes/no probability. To customize its criteria:

from bespokelabs.nimble import Noul, NoulCriteria

supported = Noul(
    instructions="Is the claim supported by the evidence?",
    criteria=NoulCriteria(yes="Supported", no="Unsupported"),
)

Score.score is sum(index * probability[index]), using zero-based indices. It is not the most likely level, and it is not normalized to 0–1 unless there are two levels. confidence measures distribution concentration, not calibrated correctness. Probability thresholds need evaluation on your own data.

Answers are accessible through result.answers or the typed result.nouls, result.choices, and result.scores dictionaries. Responses retain extra server fields through Pydantic's model_extra. result.raw_http_response exposes the HTTPX response, including timing headers. Unknown answer types or missing/mismatched answers raise APIResponseError.

Async

import asyncio
from bespokelabs.nimble import AsyncNimble, Noul


async def main():
    async with AsyncNimble() as client:
        result = await client.system_one(
            "The checkout service is down.",
            {"incident": Noul(instructions="Is there an active service incident?")},
        )
        print(result.nouls["incident"].noul)


asyncio.run(main())

Reuse one client for multiple calls. asyncio.gather works for concurrent requests; keep concurrency within await client.limits(). This SDK does not add a batch endpoint or a GPU inference runtime.

Configuration and authentication

client = Nimble(
    base_url="https://your-nimble-host.example",
    api_key="your-api-key",
    model="nimble-latest",
    timeout=120.0,
    max_retries=2,
)
# Use a context manager or call client.close() when done.
Setting Environment fallback Default
base_url BESPOKE_NIMBLE_BASE_URL Required
api_key BESPOKE_API_KEY No authorization header
model — nimble-latest
timeout — 120 seconds per HTTPX operation
max_retries — 2 retries, up to 3 attempts

An explicit argument takes precedence over its environment variable. api_key sends Authorization: Bearer .... The SDK does not create API keys or add authentication to the upstream server. Your gateway must enforce bearer authentication. Public deployments can omit the key (and leave BESPOKE_API_KEY unset).

For a private Modal proxy, pass the headers required by that deployment:

import os
from bespokelabs.nimble import Nimble

with Nimble(
    default_headers={
        "Modal-Key": os.environ["MODAL_KEY"],
        "Modal-Secret": os.environ["MODAL_SECRET"],
    }
) as client:
    print(client.health())

default_headers also supports other gateway authentication schemes. Redirects are not followed. A custom httpx.Client or httpx.AsyncClient can be passed as http_client; the caller retains ownership and must close it. Requests use the SDK's timeout and headers, plus the injected client's defaults.

Retries cover transport failures, HTTP 408, HTTP 429, and HTTP 5xx (including upstream overload status 529). Other errors fail immediately. Retries use exponential backoff with jitter and honor numeric or HTTP-date Retry-After values, capped at 60 seconds. Set max_retries=0 to disable retries. A retry may run inference again; the upstream API does not advertise idempotency support. Async cancellation propagates immediately.

Cold starts can exceed the default timeout. Increase timeout for scale-to-zero hosting, or keep the service warm. The configured timeout applies to individual network operations; it is not a deadline for all attempts combined.

Discovery and errors

from bespokelabs.nimble import APIStatusError, Nimble

with Nimble() as client:
    print(client.models())  # GET /v1/models
    print(client.limits())  # GET /v1/limits
    try:
        print(client.health())  # GET /health; unavailable services can return HTTP 503
    except APIStatusError as error:
        print(error.status_code, error.request_id)

AuthenticationError and RateLimitError specialize APIStatusError. Transport errors raise APIConnectionError or APITimeoutError; incompatible success responses raise APIResponseError. All inherit from NimbleError. Local validation raises ValueError (including Pydantic ValidationError) before sending a request. Error messages omit request bodies and credentials; HTTP failures expose error.response for explicit inspection.

The SDK validates the hosted API's limit of 64 questions and 26 candidates per question. Token limits and aggregate request budgets are enforced by the server; use client.limits() to inspect your deployment.

Hosting contract

Use the server already provided by Nimble's serving module. Its Modal deployment guide describes deploying the model. This SDK talks to that service directly.

Request:

POST /v1/systemone
Content-Type: application/json
{
  "model": "nimble-latest",
  "state": "Please refund the duplicate payment.",
  "questions": {
    "refund": {"type": "noul", "instructions": "Does the customer request a refund?"}
  }
}

Response:

{
  "model": "nimble-latest",
  "answers": {"refund": {"type": "noul", "noul": 0.99}},
  "usage": {"input_tokens": 200, "output_tokens": 2}
}

Compatibility was checked against Nimble revision f136b3f75721fda4ea961f73993cc50b08488835, its pinned openjev-sglang revision 7f84bedc169439f03379c2fa8d00ada220af2295, and its recorded deployment smoke response. Tests replay that recorded response without invoking a model. A successful local test run is not a live GPU deployment test.

The Python interface follows TypeSafe's SDK. Package/import naming follows Curator, and BESPOKE_API_KEY follows the MiniCheck SDK convention. The bespokelabs directory is an implicit namespace so this package does not add a competing root __init__.py.

Develop and release

From nimble-sdk/:

uv venv
uv pip install -e '.[dev]'
uv run --no-project pytest
uv run --no-project ruff check .
uv run --no-project ruff format --check .
uv build
uv run --no-project twine check dist/*

Before the first public release, select the repository/license metadata, confirm control of the PyPI project name, and configure PyPI publishing credentials or Trusted Publishing. Then publish the reviewed artifacts:

uv publish dist/*

Publishing is a separate release step; nothing in this checkout automatically uploads packages. After publication, verify installation in a clean environment with uv pip install bespokelabs-nimble and point it at your running deployment.

Metadata

Release files for bespokelabs-nimble 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for bespokelabs-nimble 0.1.0
File Size Uploaded
bespokelabs_nimble-0.1.0.tar.gz 13.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for bespokelabs-nimble 0.1.0
File Interpreter ABI Platform
bespokelabs_nimble-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 25.7 kB

Release files / bespokelabs_nimble-0.1.0.tar.gz

Download URL bespokelabs_nimble-0.1.0.tar.gz
Size 13.5 kB
Tags Source
SHA-256 checksum
How to use checksums
366fc7bfd07c7945891f1c4530a17d47068b31fbf1c2fb2689eafb9771cde4ff
BLAKE2b-256 checksum
How to use checksums
769a1fa8a8d14b36fc5042d897159711c7a7d9a7da87cacc36526700e678c34c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.14.6

Release files / bespokelabs_nimble-0.1.0-py3-none-any.whl

Download URL bespokelabs_nimble-0.1.0-py3-none-any.whl
Size 12.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
eb55cd1ca2b22e32a9766a3ff17661260a3a53ca9864cd4698c6e81f19e1837b
BLAKE2b-256 checksum
How to use checksums
660cf2c8d986a2cd14d3c6653d1a35a6dd40c707add68257027a2f1ce84964ae
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.14.6

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page