Skip to main content

hugging-bay

A thin, typed Python client for the Hugging Bay API: open-model discovery, byte/SHA-256 verification, server-computed safe-to-run verdicts, provenance passports, and offline lockfile checks. It does not itself judge model safety — it surfaces the server's verdicts and verifies bytes against published hashes.

The client is intentionally minimal — one method per endpoint, no client-side magic. Verdicts, fits, and identities are computed by the server and returned as parsed JSON.

Install

Install from PyPI:

python -m pip install hugging-bay

Or from a source checkout of this repository:

python -m pip install ./packages/python-sdk

Requires Python 3.9+. The only runtime dependency is httpx. The package is build-ready (python -m build produces a wheel + sdist); the actual PyPI upload is an owner release step.

Usage

from hugging_bay import Client

client = Client()  # defaults to https://huggingbay.xyz
# Optional auth: Client(token="sk-...") sends Authorization: Bearer <token>.
# The token is never printed in repr()/str() and never logged.

0. Get an API key (self-serve, optional)

Public discovery — search, resolve, download-plan, and hosted downloads — needs no token. You only authenticate for account workflows (saves, collections, watchlists). Minting a reader token is self-serve, with no signup and no prior auth:

from hugging_bay import Client

issued = Client().create_reader_token(display_name="my app", purpose="serving models")
token = issued["apiKey"]["token"]      # shown once — store it securely
client = Client(token=token)
print(client.me()["role"])             # validate the token

The full walkthrough (curl, JS/TS, MCP) is the one-screen quickstart at https://huggingbay.xyz/quickstart. A 401/403 from any authorized endpoint returns next.href = /quickstart#get-a-key, so there is no dead-end.

1. Find a commercial-safe model for a rig

from hugging_bay import Client

client = Client()
result = client.find_model(
    task="coding",
    gpu="rtx4090-24",
    ctx=8192,
    device_family="nvidia",
    commercial=True,
    limit=3,
)
best = result["bestPick"]
print(best["repo"], "-> safe-to-run:", best["safety"]["verdict"])

2. Verify a source URL

from hugging_bay import Client, HuggingBayError

client = Client()
try:
    resolved = client.resolve("https://huggingface.co/meta-llama/Llama-3.1-8B")
    print("resolved artifact:", resolved.get("artifactId") or resolved)
except HuggingBayError as err:
    print("could not resolve:", err.status, err.error)

3. Pull plan, then verify the lock against local files

from hugging_bay import Client, verify_lock

client = Client()
artifact_id = "art_123"

# The hosted download plan (signed manifest + per-file hashes).
plan = client.download_plan(artifact_id)

# Stream one hosted file, verify its exact plan hash, and record the completed
# pull. This stops the Hugging Bay redirect, follows the short-lived signed GCS
# URL without forwarding the API bearer, and only returns after attestation.
file = next(row for row in plan["files"] if row["path"] == "model.gguf")
client.download_file(
    artifact_id,
    file["path"],
    "./models/art_123/model.gguf",
    sha256=file["sha256"],
    size_bytes=file["sizeBytes"],
)

# Recompute local hashes and compare to the bay.lock document — pure stdlib,
# no network.
lock = client.lock(artifact_id)
report = verify_lock(lock, root_dir="./models/art_123")
if report["ok"]:
    print(f"verified {report['checked']} files")
else:
    for mismatch in report["mismatches"]:
        print("BAD:", mismatch["path"], mismatch["reason"])

4. Diagnose a failing run

from hugging_bay import Client

client = Client()
diagnosis = client.doctor(
    log_text="llama_model_load: error loading model: CUDA out of memory",
    runtime="llama.cpp",
)
if diagnosis["topFix"]:
    print("top fix:", diagnosis["topFix"])
for alt in diagnosis.get("smallerHostedAlternatives", []):
    print("smaller option:", alt)

Discovery → execution with Bay Run

BayRunClient runs the mirrored weights live on Bay Run (run.huggingbay.xyz, a separate origin). Take the next_call a catalog answer returns and run it — the client mints the anonymous $0 demo key for you and tolerates the edge's plain-text 429 "Rate exceeded.":

from hugging_bay import Client, BayRunClient

runnable = Client().find_runnable(task="embeddings", limit=1)
# hugging-bay.find-runnable.v1: the handoff lives at the top-level `next_call`
# (or under `recommendation.execution`); there is no `candidates` list.
next_call = runnable.get("next_call") or runnable["recommendation"]["execution"]

with BayRunClient() as run:            # mints a $0 demo key on first call
    out = run.run_next_call(next_call, input="text to embed")
    print(out)

# Or call the OpenAI-compatible surface directly:
with BayRunClient() as run:
    run.embeddings("BAAI/bge-small-en-v1.5", "hello world")
    run.rerank("BAAI/bge-reranker-base", query="q", documents=["a", "b"])
    run.run_pin("pin_abc123", input="...")

API surface

Method Endpoint
create_reader_token(display_name, purpose) POST /api/account/reader-token (self-serve, no auth)
me() GET /api/me
find_model(task, gpu, ctx, device_family, commercial, limit) GET /api/agents/find-model
artifact(id) GET /api/v1/artifacts/{id}
safety(id) GET /api/v1/artifacts/{id}/safety
bundle(id) GET /api/v1/artifacts/{id}/bundle
passport(id) GET /api/v1/artifacts/{id}/passport
resolve(repo_or_url) GET /api/v1/resolve?repo=
download_plan(id) GET /api/v1/artifacts/{id}/download-plan
download_file(id, path, destination, sha256, size_bytes) GCS redirect + streamed SHA-256 verification + POST /api/download-completions
lock(id) GET /api/v1/artifacts/{id}/lock
doctor(log_text, runtime) POST /api/doctor
find_runnable(task, limit, commercial, **filters) GET /api/agents/find-runnable
recipes() / recipe(slug) GET /api/recipes[/{slug}]
verify_lock(lock_dict, root_dir) offline, stdlib only
BayRunClient.free_key() POST /v1/keys/free (Bay Run)
BayRunClient.models() GET /v1/models (Bay Run)
BayRunClient.run_pin(pin_id, ...) POST /v1/run/{pin_id} (Bay Run)
BayRunClient.embeddings / rerank / classify / chat(...) OpenAI-compatible (Bay Run)
BayRunClient.run_next_call(next_call, input) executes a catalog next_call

Errors

Every Hugging Bay error is the same typed envelope on every surface: {error, code, status, message, requestId, cause?, next}. Any response with status >= 400 raises HuggingBayError, which surfaces those fields directly — .error/.code (machine code), .message, .status, .request_id, .next (the {action, href} recovery step), and the raw .body. On a 503 with a Retry-After header the client waits once and retries a single time.

resolve() follows status-based error handling: a well-formed reference with no catalog match raises status 404; malformed/unsupported input raises 400 (invalid_query). Bay Run errors raise BayRunError, which is tolerant of the edge's non-JSON bodies and exposes .retry_after.

License

MIT

Metadata

Release files for hugging-bay 0.4.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for hugging-bay 0.4.1
File Size Uploaded
hugging_bay-0.4.1.tar.gz 28.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for hugging-bay 0.4.1
File Interpreter ABI Platform
hugging_bay-0.4.1-py3-none-any.whl Python 3 none any Details

Total release size: 49.7 kB

Release files / hugging_bay-0.4.1.tar.gz

Download URL hugging_bay-0.4.1.tar.gz
Size 28.6 kB
Tags Source
SHA-256 checksum
How to use checksums
3e163e752da24493f0690b4223ddda50776337ad04a189a0dc0549f18243d7d3
BLAKE2b-256 checksum
How to use checksums
fdf2358b8d51fef1402cc17e9e846e3f956f5b33f5f8be5ad53edcc5322a1451
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.3

Release files / hugging_bay-0.4.1-py3-none-any.whl

Download URL hugging_bay-0.4.1-py3-none-any.whl
Size 21.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7ea1f93dcb20ed3d8e22de72c1228cfa73393bb3397ab519124004f50037bcf5
BLAKE2b-256 checksum
How to use checksums
97519ee69eb05e43312962bdd199a34cef8b7e3ef4e974decda48cf0d25dafb3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.3

Release history Release notifications | RSS feed

0.4.2

2 release files

This release

0.4.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page