Skip to main content

Hippo 🦛

pip install hippo-llm | Python 3.10+ | MIT | 中文文档

Run 30B models on a ¥3800 GPU at 78 tok/s. Then search through your documents without installing ChromaDB.

30-second setup

hippo-pipeline serve --model qwen3-30b-a3b-q3 --mode standalone
# → OpenAI-compatible API at localhost:8000/v1/chat/completions
import openai
client = openai.OpenAI(base_url="http://localhost:8000/v1", api_key="none")
r = client.chat.completions.create(
    model="qwen3-30b-a3b-q3",
    messages=[{"role": "user", "content": "Explain pipeline parallelism"}],
    max_tokens=500
)
print(r.choices[0].message.content)
Two-machine setup
# Machine 1
hippo-pipeline serve --model gemma-3-12b --mode pipeline --rank 0

# Machine 2
hippo-pipeline serve --model gemma-3-12b --mode pipeline --rank 1 \
  --coordinator http://192.168.1.10:9000

Split the model across machines. Run what doesn't fit on one GPU.

One install for inference + search

Most RAG setups need two services: Ollama for inference + ChromaDB for vectors. Hippo gives you both in one pip install.

from hippo.embedding import EmbeddingEngine, VectorStore

engine = EmbeddingEngine(model="nomic-embed-text")  # uses local Ollama
store = VectorStore("docs.db", mode="hybrid")  # BM25 + dense RRF fusion

# Add documents
store.add_batch([
    {"text": "Pipeline parallelism splits layers across devices", "metadata": {"source": "readme"}},
    {"text": "BM25 handles exact keyword matches", "metadata": {"source": "docs"}},
    {"text": "Speculative decoding improves latency by 2-3x", "metadata": {"source": "benchmarks"}},
], engine=engine)

# Hybrid search (BM25 + semantic, RRF fused)
results = store.search("how to run big models on small GPUs", engine=engine, top_k=5)
for doc in results:
    print(f"[{doc.score:.3f}] {doc.text}")

No external vector DB. SQLite for persistence, numpy for similarity. Works offline.

Full RAG example with local LLM
from hippo.embedding import EmbeddingEngine, VectorStore
import openai

# 1. Index your documents (one-time)
engine = EmbeddingEngine(model="nomic-embed-text")
store = VectorStore("knowledge.db", mode="hybrid")

documents = [
    "Hippo splits model layers across multiple devices using TCP.",
    "Each device only loads its shard of layers, reducing memory per device.",
    "The loop detector catches semantic repetition using Jaccard similarity.",
    "BM25 hybrid search combines keyword matching with semantic similarity.",
]
store.add_batch([{"text": d} for d in documents], engine=engine)

# 2. RAG query
query = "how does hippo handle memory?"
results = store.search(query, engine=engine, top_k=2)
context = "\n".join(doc.text for doc in results)

# 3. Generate answer with local LLM
client = openai.OpenAI(base_url="http://localhost:8000/v1", api_key="none")
response = client.chat.completions.create(
    model="qwen3-30b-a3b-q3",
    messages=[
        {"role": "system", "content": f"Answer based on this context:\n{context}"},
        {"role": "user", "content": query}
    ]
)
print(response.choices[0].message.content)

What's inside

Feature Details
Pipeline Parallelism Split any HF model across N machines. Mac + PC mixed. Plain TCP, no MPI.
Loop Detection Jaccard-similarity detector catches semantic repetition that repeat_penalty misses.
Embedding & Search Dense + BM25 + hybrid RRF fusion. SQLite-backed, sub-ms queries.
Chinese-optimized BM25 Built-in Chinese tokenizer with stop words. No jieba needed.
ANN Index Approximate nearest neighbor for large collections (>10K docs).
OpenAI-Compatible API Drop-in /v1/chat/completions. Works with LangChain, LlamaIndex, anything.
Auto Memory Budget Calculates shard splits from available VRAM automatically.

When to use Hippo

You want... Use this
Local inference on one machine --mode standalone with any GGUF model
Run a model too big for one device --mode pipeline across 2+ machines
RAG without installing ChromaDB VectorStore(mode="hybrid")
Search Chinese documents BM25 with built-in tokenizer

Install

pip install hippo-llm

Requirements: Python 3.10+, Ollama running locally for model weights and embeddings.

Roadmap

  • v0.3: ANN index for >10K document collections ✅
  • v0.4: Multi-shard support (>2 devices), automatic layer balancing
  • v0.5: Speculative decoding across shards
  • v0.6: Built-in model download + GGUF auto-conversion

Benchmarks

Setup Model Speed
Mac Mini M2 (16GB) Qwen3-4B-Q4 41 tok/s
RTX 5060 Ti (16GB) Qwen3-14B-Q4 41 tok/s
2× Mac Mini (16GB each) Qwen3-30B-A3B-Q3 78 tok/s
Mac Mini M2 (16GB) Qwen3-30B-A3B-Q3 24 tok/s

License

MIT

Author

lawcontinue — GitHub

Metadata

Release files for hippo-llm 0.3.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for hippo-llm 0.3.1
File Size Uploaded
hippo_llm-0.3.1.tar.gz 63.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for hippo-llm 0.3.1
File Interpreter ABI Platform
hippo_llm-0.3.1-py3-none-any.whl Python 3 none any Details

Total release size: 136.4 kB

Release files / hippo_llm-0.3.1.tar.gz

Download URL hippo_llm-0.3.1.tar.gz
Size 63.9 kB
Tags Source
SHA-256 checksum
How to use checksums
53b8acae00f659e35b6fed0e1a10250760f1643cdd3bdebad31fbac126e1b695
BLAKE2b-256 checksum
How to use checksums
c6e2f1c5b6515b505051c1c2754f9f6786687505d7a47e0ba81e00cdddd63c2d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jun 11, 2026.

Transparency log

Release files / hippo_llm-0.3.1-py3-none-any.whl

Download URL hippo_llm-0.3.1-py3-none-any.whl
Size 72.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
31056859e3069cb354b5e34d9f91ba7ad16a2ba43b13e9bb80a06265dde138ff
BLAKE2b-256 checksum
How to use checksums
206ccb6833a9967eac716cebf59c910b80c7c23c4d69f0ade19097b91140c647
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jun 11, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.3.1 This release

2 release files

0.3.0

2 release files

0.2.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page