Skip to main content

hugging-bay

A thin, typed Python client for the Hugging Bay API: open-model discovery, byte/SHA-256 verification, server-computed safe-to-run verdicts, provenance passports, and offline lockfile checks. It does not itself judge model safety — it surfaces the server's verdicts and verifies bytes against published hashes.

The client is intentionally minimal — one method per endpoint, no client-side magic. Verdicts, fits, and identities are computed by the server and returned as parsed JSON.

Install

Install from PyPI (0.4.2 or newer — the 0.4.1 wheel is missing hugging_bay/_version.py and does not import):

python -m pip install 'hugging-bay>=0.4.2'

Or from a source checkout of this repository:

python -m pip install ./packages/python-sdk

Requires Python 3.9+. The only runtime dependency is httpx. The package is build-ready (python -m build produces a wheel + sdist); the actual PyPI upload is an owner release step.

Usage

from hugging_bay import Client

client = Client()  # defaults to https://huggingbay.xyz
# Optional auth: Client(token="sk-...") sends Authorization: Bearer <token>.
# The token is never printed in repr()/str() and never logged.

0. Get an API key (self-serve, optional)

Public discovery — search, resolve, download-plan, and hosted downloads — needs no token. You only authenticate for account workflows (saves, collections, watchlists). Minting a reader token is self-serve, with no signup and no prior auth:

from hugging_bay import Client

issued = Client().create_reader_token(display_name="my app", purpose="serving models")
token = issued["apiKey"]["token"]      # shown once — store it securely
client = Client(token=token)
print(client.me()["role"])             # validate the token

The full walkthrough (curl, JS/TS, MCP) is the one-screen quickstart at https://huggingbay.xyz/quickstart. A 401/403 from any authorized endpoint returns next.href = /quickstart#get-a-key, so there is no dead-end.

1. Find a commercial-safe model for a rig

from hugging_bay import Client

client = Client()
result = client.find_model(
    task="coding",
    gpu="rtx4090-24",
    ctx=8192,
    device_family="nvidia",
    commercial=True,
    limit=3,
)
best = result["bestPick"]
print(best["repo"], "-> safe-to-run:", best["safety"]["verdict"])

2. Verify a source URL

from hugging_bay import Client, HuggingBayError

client = Client()
try:
    resolved = client.resolve("https://huggingface.co/meta-llama/Llama-3.1-8B")
    print("resolved artifact:", resolved.get("artifactId") or resolved)
except HuggingBayError as err:
    print("could not resolve:", err.status, err.error)

3. Pull plan, then verify the lock against local files

from hugging_bay import Client, verify_lock

client = Client()
artifact_id = "art_123"

# The hosted download plan (signed manifest + per-file hashes).
plan = client.download_plan(artifact_id)

# Stream one hosted file, verify its exact plan hash, and record the completed
# pull. This stops the Hugging Bay redirect, follows the short-lived signed GCS
# URL without forwarding the API bearer, and only returns after attestation.
file = next(row for row in plan["files"] if row["path"] == "model.gguf")
client.download_file(
    artifact_id,
    file["path"],
    "./models/art_123/model.gguf",
    sha256=file["sha256"],
    size_bytes=file["sizeBytes"],
)

# Recompute local hashes and compare to the bay.lock document — pure stdlib,
# no network.
lock = client.lock(artifact_id)
report = verify_lock(lock, root_dir="./models/art_123")
if report["ok"]:
    print(f"verified {report['checked']} files")
else:
    for mismatch in report["mismatches"]:
        print("BAD:", mismatch["path"], mismatch["reason"])

4. Diagnose a failing run

from hugging_bay import Client

client = Client()
diagnosis = client.doctor(
    log_text="llama_model_load: error loading model: CUDA out of memory",
    runtime="llama.cpp",
)
if diagnosis["topFix"]:
    print("top fix:", diagnosis["topFix"])
for alt in diagnosis.get("smallerHostedAlternatives", []):
    print("smaller option:", alt)

Discovery → execution with Bay Run

BayRunClient runs the mirrored weights live on Bay Run (run.huggingbay.xyz, a separate origin). Take the next_call a catalog answer returns and run it — the client mints the anonymous $0 demo key for you and tolerates the edge's plain-text 429 "Rate exceeded.":

from hugging_bay import Client, BayRunClient

runnable = Client().find_runnable(task="embeddings", limit=1)
# hugging-bay.find-runnable.v1: the handoff lives at the top-level `next_call`
# (or under `recommendation.execution`); there is no `candidates` list.
next_call = runnable.get("next_call") or runnable["recommendation"]["execution"]

with BayRunClient() as run:            # mints a $0 demo key on first call
    out = run.run_next_call(next_call, input="text to embed")
    print(out)

# Or call the OpenAI-compatible surface directly:
with BayRunClient() as run:
    run.embeddings("BAAI/bge-small-en-v1.5", "hello world")
    run.rerank("BAAI/bge-reranker-base", query="q", documents=["a", "b"])
    run.run_pin("pin_abc123", input="...")

API surface

Method Endpoint
create_reader_token(display_name, purpose) POST /api/account/reader-token (self-serve, no auth)
me() GET /api/me
find_model(task, gpu, ctx, device_family, commercial, limit) GET /api/agents/find-model
artifact(id) GET /api/v1/artifacts/{id}
safety(id) GET /api/v1/artifacts/{id}/safety
bundle(id) GET /api/v1/artifacts/{id}/bundle
passport(id) GET /api/v1/artifacts/{id}/passport
resolve(repo_or_url) GET /api/v1/resolve?repo=
download_plan(id) GET /api/v1/artifacts/{id}/download-plan
download_file(id, path, destination, sha256, size_bytes) GCS redirect + streamed SHA-256 verification + POST /api/download-completions
lock(id) GET /api/v1/artifacts/{id}/lock
doctor(log_text, runtime) POST /api/doctor
find_runnable(task, limit, commercial, **filters) GET /api/agents/find-runnable
recipes() / recipe(slug) GET /api/recipes[/{slug}]
verify_lock(lock_dict, root_dir) offline, stdlib only
BayRunClient.free_key() POST /v1/keys/free (Bay Run)
BayRunClient.models() GET /v1/models (Bay Run)
BayRunClient.run_pin(pin_id, ...) POST /v1/run/{pin_id} (Bay Run)
BayRunClient.embeddings / rerank / classify / chat(...) OpenAI-compatible (Bay Run)
BayRunClient.run_next_call(next_call, input) executes a catalog next_call

Errors

Every Hugging Bay error is the same typed envelope on every surface: {error, code, status, message, requestId, cause?, next}. Any response with status >= 400 raises HuggingBayError, which surfaces those fields directly — .error/.code (machine code), .message, .status, .request_id, .next (the {action, href} recovery step), and the raw .body. On a 503 with a Retry-After header the client waits once and retries a single time.

resolve() follows status-based error handling: a well-formed reference with no catalog match raises status 404; malformed/unsupported input raises 400 (invalid_query). Bay Run errors raise BayRunError, which is tolerant of the edge's non-JSON bodies and exposes .retry_after.

License

MIT

Metadata

Release files for hugging-bay 0.4.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for hugging-bay 0.4.2
File Size Uploaded
hugging_bay-0.4.2.tar.gz 31.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for hugging-bay 0.4.2
File Interpreter ABI Platform
hugging_bay-0.4.2-py3-none-any.whl Python 3 none any Details

Total release size: 52.2 kB

Release files / hugging_bay-0.4.2.tar.gz

Download URL hugging_bay-0.4.2.tar.gz
Size 31.1 kB
Tags Source
SHA-256 checksum
How to use checksums
ec4a136fef3175a5ccec514b537efb2f89f60ee0affec0a5d97def80de81d954
BLAKE2b-256 checksum
How to use checksums
badf8c7e4bb432f0db00feada57dc084235f19329d4220d3eb66039ea31c911e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.3

Release files / hugging_bay-0.4.2-py3-none-any.whl

Download URL hugging_bay-0.4.2-py3-none-any.whl
Size 21.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b1161e379e32b9b82e2b67771745ba65915759eac41883eb67ecaeda875382fe
BLAKE2b-256 checksum
How to use checksums
f21b77414f7b591cdf56130fc6cd25a9ee61889f86a29c7b90837e993590ce66
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.3

Release history Release notifications | RSS feed

This release

0.4.2 This release

2 release files

0.4.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page