hugging-bay
A thin, typed Python client for the Hugging Bay API: open-model discovery, byte/SHA-256 verification, server-computed safe-to-run verdicts, provenance passports, and offline lockfile checks. It does not itself judge model safety — it surfaces the server's verdicts and verifies bytes against published hashes.
The client is intentionally minimal — one method per endpoint, no client-side magic. Verdicts, fits, and identities are computed by the server and returned as parsed JSON.
Install
Install from PyPI:
python -m pip install hugging-bay
Or from a source checkout of this repository:
python -m pip install ./packages/python-sdk
Requires Python 3.9+. The only runtime dependency is httpx. The package is
build-ready (python -m build produces a wheel + sdist); the actual PyPI
upload is an owner release step.
Usage
from hugging_bay import Client
client = Client() # defaults to https://huggingbay.xyz
# Optional auth: Client(token="sk-...") sends Authorization: Bearer <token>.
# The token is never printed in repr()/str() and never logged.
0. Get an API key (self-serve, optional)
Public discovery — search, resolve, download-plan, and hosted downloads — needs no token. You only authenticate for account workflows (saves, collections, watchlists). Minting a reader token is self-serve, with no signup and no prior auth:
from hugging_bay import Client
issued = Client().create_reader_token(display_name="my app", purpose="serving models")
token = issued["apiKey"]["token"] # shown once — store it securely
client = Client(token=token)
print(client.me()["role"]) # validate the token
The full walkthrough (curl, JS/TS, MCP) is the one-screen quickstart at
https://huggingbay.xyz/quickstart. A 401/403 from any authorized endpoint
returns next.href = /quickstart#get-a-key, so there is no dead-end.
1. Find a commercial-safe model for a rig
from hugging_bay import Client
client = Client()
result = client.find_model(
task="coding",
gpu="rtx4090-24",
ctx=8192,
device_family="nvidia",
commercial=True,
limit=3,
)
best = result["bestPick"]
print(best["repo"], "-> safe-to-run:", best["safety"]["verdict"])
2. Verify a source URL
from hugging_bay import Client, HuggingBayError
client = Client()
try:
resolved = client.resolve("https://huggingface.co/meta-llama/Llama-3.1-8B")
print("resolved artifact:", resolved.get("artifactId") or resolved)
except HuggingBayError as err:
print("could not resolve:", err.status, err.error)
3. Pull plan, then verify the lock against local files
from hugging_bay import Client, verify_lock
client = Client()
artifact_id = "art_123"
# The hosted download plan (signed manifest + per-file hashes).
plan = client.download_plan(artifact_id)
# Stream one hosted file, verify its exact plan hash, and record the completed
# pull. This stops the Hugging Bay redirect, follows the short-lived signed GCS
# URL without forwarding the API bearer, and only returns after attestation.
file = next(row for row in plan["files"] if row["path"] == "model.gguf")
client.download_file(
artifact_id,
file["path"],
"./models/art_123/model.gguf",
sha256=file["sha256"],
size_bytes=file["sizeBytes"],
)
# Recompute local hashes and compare to the bay.lock document — pure stdlib,
# no network.
lock = client.lock(artifact_id)
report = verify_lock(lock, root_dir="./models/art_123")
if report["ok"]:
print(f"verified {report['checked']} files")
else:
for mismatch in report["mismatches"]:
print("BAD:", mismatch["path"], mismatch["reason"])
4. Diagnose a failing run
from hugging_bay import Client
client = Client()
diagnosis = client.doctor(
log_text="llama_model_load: error loading model: CUDA out of memory",
runtime="llama.cpp",
)
if diagnosis["topFix"]:
print("top fix:", diagnosis["topFix"])
for alt in diagnosis.get("smallerHostedAlternatives", []):
print("smaller option:", alt)
Discovery → execution with Bay Run
BayRunClient runs the mirrored weights live on Bay Run (run.huggingbay.xyz,
a separate origin). Take the next_call a catalog answer returns and run it —
the client mints the anonymous $0 demo key for you and tolerates the edge's
plain-text 429 "Rate exceeded.":
from hugging_bay import Client, BayRunClient
runnable = Client().find_runnable(task="embeddings", limit=1)
# hugging-bay.find-runnable.v1: the handoff lives at the top-level `next_call`
# (or under `recommendation.execution`); there is no `candidates` list.
next_call = runnable.get("next_call") or runnable["recommendation"]["execution"]
with BayRunClient() as run: # mints a $0 demo key on first call
out = run.run_next_call(next_call, input="text to embed")
print(out)
# Or call the OpenAI-compatible surface directly:
with BayRunClient() as run:
run.embeddings("BAAI/bge-small-en-v1.5", "hello world")
run.rerank("BAAI/bge-reranker-base", query="q", documents=["a", "b"])
run.run_pin("pin_abc123", input="...")
API surface
| Method | Endpoint |
|---|---|
create_reader_token(display_name, purpose) |
POST /api/account/reader-token (self-serve, no auth) |
me() |
GET /api/me |
find_model(task, gpu, ctx, device_family, commercial, limit) |
GET /api/agents/find-model |
artifact(id) |
GET /api/v1/artifacts/{id} |
safety(id) |
GET /api/v1/artifacts/{id}/safety |
bundle(id) |
GET /api/v1/artifacts/{id}/bundle |
passport(id) |
GET /api/v1/artifacts/{id}/passport |
resolve(repo_or_url) |
GET /api/v1/resolve?repo= |
download_plan(id) |
GET /api/v1/artifacts/{id}/download-plan |
download_file(id, path, destination, sha256, size_bytes) |
GCS redirect + streamed SHA-256 verification + POST /api/download-completions |
lock(id) |
GET /api/v1/artifacts/{id}/lock |
doctor(log_text, runtime) |
POST /api/doctor |
find_runnable(task, limit, commercial, **filters) |
GET /api/agents/find-runnable |
recipes() / recipe(slug) |
GET /api/recipes[/{slug}] |
verify_lock(lock_dict, root_dir) |
offline, stdlib only |
BayRunClient.free_key() |
POST /v1/keys/free (Bay Run) |
BayRunClient.models() |
GET /v1/models (Bay Run) |
BayRunClient.run_pin(pin_id, ...) |
POST /v1/run/{pin_id} (Bay Run) |
BayRunClient.embeddings / rerank / classify / chat(...) |
OpenAI-compatible (Bay Run) |
BayRunClient.run_next_call(next_call, input) |
executes a catalog next_call |
Errors
Every Hugging Bay error is the same typed envelope on every surface:
{error, code, status, message, requestId, cause?, next}. Any response with
status >= 400 raises HuggingBayError, which surfaces those fields directly —
.error/.code (machine code), .message, .status, .request_id, .next
(the {action, href} recovery step), and the raw .body. On a 503 with a
Retry-After header the client waits once and retries a single time.
resolve() follows status-based error handling: a well-formed reference with no
catalog match raises status 404; malformed/unsupported input raises 400
(invalid_query). Bay Run errors raise BayRunError, which is tolerant of the
edge's non-JSON bodies and exposes .retry_after.
License
MIT
Metadata
Release files for hugging-bay 0.4.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| hugging_bay-0.4.1.tar.gz | 28.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| hugging_bay-0.4.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 49.7 kB
Release files / hugging_bay-0.4.1.tar.gz
| Download URL | hugging_bay-0.4.1.tar.gz |
|---|---|
| Size | 28.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
3e163e752da24493f0690b4223ddda50776337ad04a189a0dc0549f18243d7d3
|
|
BLAKE2b-256 checksum How to use checksums |
fdf2358b8d51fef1402cc17e9e846e3f956f5b33f5f8be5ad53edcc5322a1451
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.3
|
Release files / hugging_bay-0.4.1-py3-none-any.whl
| Download URL | hugging_bay-0.4.1-py3-none-any.whl |
|---|---|
| Size | 21.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7ea1f93dcb20ed3d8e22de72c1228cfa73393bb3397ab519124004f50037bcf5
|
|
BLAKE2b-256 checksum How to use checksums |
97519ee69eb05e43312962bdd199a34cef8b7e3ef4e974decda48cf0d25dafb3
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.3
|