Skip to main content
Yanked

This release has been yanked by its maintainers, and will be ignored by installers, except when explicitly specified.

HiveMind Python SDK

Persistent memory and a working-context compiler for your agent, in three lines inside your own loop. Your model, your key — HiveMind never calls an LLM.

pip install hivemind-sdk

Python ≥ 3.10. The import name is hivemind. Docs: https://hivemind.militant.ai/docs

Async

Agent systems run on the event loop; so does this SDK. AsyncHiveMind mirrors the sync facade — same three-liner, nothing blocks, pooled connections underneath:

from hivemind import AsyncHiveMind

async with AsyncHiveMind(base_url="...", api_key="...") as mind:
    session = mind.session(budget_total=8192)
    result = await session.turn(user_input)
    reply = await call_your_llm(result.messages)
    await session.record(reply)

AsyncHivemindClient underneath mirrors the engine's native HivemindOperations surface name-for-name (compile_working_context, store_conversation_exchange, recall_memory_with_metadata, …) — code written against the ops layer speaks to the hosted service with the same vocabulary. The sync client remains pure standard library.

from hivemind import HiveMind

mind = HiveMind(base_url="...", api_key="...", tenant_id="...")
session = mind.session(budget_total=8192)

result = session.turn(user_input)       # store -> recall -> compile
reply = call_your_llm(result.messages)  # your model, your key
session.record(reply)                   # completes the exchange

What one turn() does

One HTTP request (POST /turn) — the service composes, in order:

  1. Stores the user message (receipted).
  2. Semantically recalls relevant memories and past conversation.
  3. Compiles local history + recalled records + active holds + your operator briefing into a token-budgeted bundle.

With record(), the whole exchange is two calls — your model runs between them.

result.messages is the entire prompt payload — send it as-is, splice nothing in. result.bundle["decisions"] explains every admission under the budget; result.receipt is the audit record. Empty recall on a young tenant is normal, not an error.

session.record(reply) stores your model's reply as the other half of the exchange, so the next turn — and every future session — remembers it.

How conversation memory recalls

Conversation is remembered as call/response exchanges — a user question, an agent's instruction, whatever the initiating text was, plus the reply it produced. Recall matches your query against both sides of every past exchange, and returns whole exchanges: one result slot per exchange, rendered call-then-response, never an answer without the message that produced it (and vice versa). Facts that appear only in a reply are just as findable as the calls that prompted them.

The recall pool is sized automatically from your session's token budget — a bigger budget_total recalls more candidates, and the compiler's budget admission decides what actually enters the bundle (with every decision receipted). Pass recall_top_k to a session only if you want to force a fixed pool.

Still worth designing around: record() files the reply and completes the exchange — treat it as part of the loop. And durable facts that should stand alone — decisions, outcomes, lessons — belong in mind.remember(...), where you control their metadata and lifecycle.

Every turn also reports where its time went: result.timings carries the server-side phase breakdown in seconds (embed_s, store_and_recall_s, completion_s, compile_s, total_s) — a slow turn names its own bottleneck.

Beyond the loop

  • mind.remember(content, metadata) — deliberately store a durable lesson, decision, fact, or outcome.
  • mind.recall(query) / mind.recall_filtered(query, metadata) — explicit recall over deliberate memories only ([] when nothing matches). Conversation history is a separate record class, recalled automatically inside the loop — or explicitly via client.recall_conversation.
  • session.hold_set(key, content) / hold_clear(key) — pin operational state ("stop-order", "API is down") into every compile until cleared.
  • mind.receipts(session_id=...) — the audit trail: what ran, what it consumed, what it produced, with lineage.
  • mind.delete_by_metadata(metadata) — destructive, audited deletion.
  • mind.client — the raw HTTP client for anything not wrapped.

Building a chat host (not a script)

The loop contract above is complete; a long-lived, restartable, streaming host needs the layer around it:

Stream first, compile second. turn() does real work — a store, two semantic recalls, and a full context compile — and at large budgets that is a multi-second operation, not a lookup. If your HTTP handler calls turn() before sending anything, the client sees pure silence for those seconds, and most streaming clients (SSE, WebSocket, fetch readers) treat prolonged silence as a dead stream and give up. So establish the stream first: send headers or a heartbeat frame, then compile, then stream your model. The compile's phase breakdown comes back in result.timings, so your own logs can say where a slow turn went without guessing:

yield sse_headers()                      # first byte now — the stream is alive
result = await session.turn(user_input)  # then the multi-second compile
yield {"tokens": result.token_count, "timings": result.timings}
# stream the model's reply, generated from result.messages
await session.record(reply)

Restart: session.hydrate(messages). A Session keeps a small local history in process memory — it is the "recent conversation" the compiler weaves together with recalled memory each turn. After a worker restart or redeploy, that history is gone, and without it the next compile sees only what semantic recall happens to surface. hydrate() rebuilds it from your own database: pass {role, content, timestamp} dicts, oldest first, with real timestamps — the compiler orders conversation by timestamp, so fabricated or missing times scramble reading order. Nothing is re-stored on the service; hydrate is purely local. The turn counter resumes automatically (or pass turn_number= explicitly). Map your conversation id to session_id so a restarted host lands back in the same session — and when the user clears a thread, start a new session id rather than hydrating the old one, so the fresh thread doesn't inherit history the user asked to be rid of.

Aborted stream: don't record. Users cancel generations; connections drop mid-stream. When that happens, call session.abort_open_exchange() and move on — do not try to "close out" the exchange. The user's message was already stored by turn() and remains recallable as an unanswered call, which is an honest record of what happened; the next turn() simply opens a fresh exchange. Never record("") to fill the gap — the service rejects empty content — and only record a partial reply if you actually want that partial remembered as what was said.

Continuation hops. role in an exchange is positional — it marks which side initiated (call) and which side answered (response), not whether a human or a machine was talking. That means a harness that feeds the model's own output back in as the next input — tool loops, inner monologue, "keep going" patterns — is using turn(prior_output) exactly as intended: the prior output is the initiating utterance of the next exchange, and recall will later find that exchange by it. If your host needs to reshape local history around such hops (say, to drop intermediate hops from the visible transcript), hydrate() is the supported way to rebuild history in the shape you want.

Authorship. Who produced a message is a separate question from which side of the exchange it sits on, and it gets a separate field: turn(text, author="operator-7") and record(reply, author="scout-agent") stamp identity as ordinary metadata. The compiler surfaces it as the chat message's name field — the slot chat APIs already accept for named participants — in both live history and recalled exchanges. This is what keeps multi-agent transcripts attributed (agent A's calls and agent B's replies each carry their own name) and what keeps a self-continuation hop from ever showing the model its own words under someone else's label. Because it is plain metadata, it also filters: recall_filtered(query, {"author": "scout-agent"}) returns only that agent's records. Untagged messages behave exactly as before.

Turn-local instructions. Everything in result.messages came from somewhere, and where it came from decides how long it lives. The ladder, shortest-lived first: artifacts_context rides one compile and is never stored or recalled — use it for document excerpts and one-shot notes. operator_briefing is also compiled-only (never stored), but the session re-sends it every turn — use it for standing instructions and mode flags. Holds persist across turns until you clear them — use hold_set/hold_clear around a tool burst or an operational condition ("API is down"). remember() is durable forever and recallable from any session — use it for decisions, outcomes, lessons. Pick the shortest lifetime that does the job; anything longer pollutes future context.

Context meter. If you show users (or your logs) a context gauge, read it from result.token_count against result.budget_total — those are the compiled packet's actual numbers, measured by the thing that built it. Local history length undercounts (it knows nothing about recall), and your LLM provider's usage.input_tokens measures a different tokenizer after your own additions. One packet, one authoritative meter.

Error policy — memory failures and chat failures are different severities, and your host should decide deliberately which is which instead of letting an exception decide for it. The cases:

Case Do this
404 (nothing stored yet) normal on young tenants — proceed
429 sleep exc.retry_after, retry
402 storage cap writes blocked, recall and compile keep working — don't fail the chat
timeout / 5xx your call: fail-open to local context or fail-closed — pick one and say so
record() fails the chat already succeeded — retry it out-of-band

Don't store model control tokens. HiveMind stores and returns text verbatim — nothing is sanitized. But recalled text re-enters your model's context on later turns, and if it contains that model's own special tokens (<|...|> markers and the like), the tokenizer may swallow them or, worse, obey them as control sequences. Strip or defang such markers before turn/record; if you need a durable marker in stored text, pick a form no tokenizer treats as special.

Base URL is the host root. The client appends /hivemind/v1 to every request itself. Pasting a URL that already contains the prefix (say, from the OpenAPI page) double-prefixes every path and produces 404s that look like a broken service. Use https://api.hivemind.militant.ai — never .../hivemind/v1.

Configuration

Constructor arguments override environment:

Env var Meaning
HIVEMIND_BASE_URL Service root (hosted or local — same API)
HIVEMIND_API_KEY Sent as Authorization: Bearer <key>
HIVEMIND_TENANT_ID Your tenant (X-Tenant-ID)
HIVEMIND_PROJECT Project binding (X-Hivemind-Project), optional
HIVEMIND_TIMEOUT Request timeout, seconds (default 120)
HIVEMIND_BUDGET_TOTAL Default compile token budget (default 4096)

Development note (this repo)

The import name hivemind collides with the service package at the repo root, so run SDK tests as their own invocation:

python -m pytest sdk/python/tests

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

hivemind_sdk-0.2.3.tar.gz (27.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

hivemind_sdk-0.2.3-py3-none-any.whl (22.5 kB view details)

Uploaded Python 3

File details

Details for the file hivemind_sdk-0.2.3.tar.gz.

File metadata

  • Download URL: hivemind_sdk-0.2.3.tar.gz
  • Upload date:
  • Size: 27.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.3

File hashes

Hashes for hivemind_sdk-0.2.3.tar.gz
Algorithm Hash digest
SHA256 95b534f3547fe69701ce71691e098da7b5c75c87fb6df7e67de41da5b2e215c2
MD5 df764b86e01376b4eaacc2e970fd5d69
BLAKE2b-256 9af6069b2d36d71da3ec309ebf116bcf3672db1107510ecea492700639ee99ba

See more details on using hashes here.

File details

Details for the file hivemind_sdk-0.2.3-py3-none-any.whl.

File metadata

  • Download URL: hivemind_sdk-0.2.3-py3-none-any.whl
  • Upload date:
  • Size: 22.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.3

File hashes

Hashes for hivemind_sdk-0.2.3-py3-none-any.whl
Algorithm Hash digest
SHA256 9768ae3dc0d849c2efad84b7d0675a2ed8324b5f48b9cc41009efae67faf97b5
MD5 7a5fbb322d5b951a389acbcda0749c52
BLAKE2b-256 4a899d46960900257fdb9f159e7cb950ca0a9efc395396e060ac833a86317da8

See more details on using hashes here.

Release history Release notifications | RSS feed

0.3.0

2 files

0.2.4

2 files

This release

0.2.3 This release

2 files

0.2.2

2 files

0.2.1

2 files

0.2.0

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page