Skip to main content

openframe-adapters-db-chromadb

ChromaDB vector-store adapter for the OpenFrame Microservice Suite.

Part of the openframe-adapters monorepo. Implements BaseVectorStore[T] (which extends BaseRepository[T]) from openframe-core using chromadb's native async HTTP client (chromadb.AsyncHttpClient).

This adapter targets Chroma's client-server (HTTP) mode exclusively. There is no embedded/persistent-local-mode support (chromadb.Client()/ PersistentClient()) — that mode has no server to connect to and is out of scope here, matching how this ecosystem's other adapters (Postgres, Mongo, Redis) assume a running backend server rather than an embedded/local-file mode.


Installation

pip install openframe-adapters-db-chromadb

No required env vars — every setting has a default suitable for a locally run Chroma server (chroma run, which listens on localhost:8000). You must set a collection name before using a repository, either via CHROMA_COLLECTION or the collection= constructor argument.


Async client research finding

chromadb 1.5.9 ships a genuine, documented native-async client-server API — not an experimental bolt-on over the sync client. Verified directly against the installed package:

>>> import chromadb, inspect
>>> chromadb.__version__
'1.5.9'
>>> inspect.iscoroutinefunction(chromadb.AsyncHttpClient)
True

Every method this adapter calls on the collection object AsyncClientAPI.get_or_create_collection(...) returns is a real coroutine function (inspect.iscoroutinefunction confirms this for AsyncCollection.add, .get, .update, .upsert, .delete, .query, .count, as well as AsyncClientAPI.get_or_create_collection/.create_collection/.get_collection themselves).

Because that surface genuinely exists, this adapter is built the same shape as openframe-adapters-db-postgres/-oracle — a client cached by connection-parameter tuple, native async def CRUD methods — not the run_in_executor-wrapped-sync shape a backend with no real async client would need (per ADR-003: the async strategy is decided per-adapter against the actually-installed driver, not forced). See openframe/adapters/db/chromadb/connection.py's module docstring for the full finding, including the httpx-exception classification detail below.


Quick start

Raw dict mode

from openframe.adapters.db.chromadb import ChromaDBSettings, ChromaDBRepository

settings = ChromaDBSettings(chroma_collection="items")  # reads CHROMA_* from env
repo = ChromaDBRepository(settings)

item = await repo.get("abc-123")                 # dict | None
items, total = await repo.list(10, 0)             # ([dict, ...], int)
created = await repo.create({"id": "abc-123", "vector": [0.1, 0.2, 0.3], "metadata": {"name": "x"}})
updated = await repo.update({"id": "abc-123", "vector": [0.1, 0.2, 0.3], "metadata": {"name": "y"}})
deleted = await repo.delete("abc-123")             # bool
results = await repo.search(query_vector=[0.1, 0.2, 0.3], k=5)  # list[dict], nearest first

Rows are shaped {"id": ..., "vector": [...], "metadata": {...}, "document": ...} — the plain-dict shape _row_to_entity()/_entity_to_row() transpose to/from Chroma's own columnar (parallel-list) wire format. See "Columnar wire format" below.

Typed domain mode

from dataclasses import dataclass
from openframe.adapters.db.chromadb import ChromaDBSettings, ChromaDBRepository

@dataclass
class Item:
    id: str
    vector: list[float]
    name: str

class ItemRepository(ChromaDBRepository[Item]):
    _collection = "items"

    def _row_to_entity(self, row) -> Item:
        return Item(id=row["id"], vector=row["vector"], name=row["metadata"].get("name", ""))

    def _entity_to_row(self, entity: Item) -> dict:
        return {"id": entity.id, "vector": entity.vector, "metadata": {"name": entity.name}}

settings = ChromaDBSettings()
repo = ItemRepository(settings)
item: Item | None = await repo.get("abc-123")

Columnar wire format

ChromaDB's collection.get(...)/collection.query(...) do not return row/record objects the way asyncpg/pymongo do — they return a dict of parallel lists, one list per field (ids, embeddings, metadatas, documents), all the same length and ordered together positionally. query() nests each list one level deeper (one list-of-lists per query embedding, since Chroma's query API supports batched multi-vector queries); this adapter always sends exactly one query embedding, so it always unwraps index [0].

ChromaDBRepository._get_result_to_rows()/._query_result_to_rows() transpose that columnar shape into plain per-entity row dicts before _row_to_entity() ever sees them — the single most unusual part of this adapter relative to a row-oriented driver.


create() vs upsert()

create() uses collection.add(), which fails on a duplicate id (IDAlreadyExistsError/DuplicateIDError from the server) — matching BaseRepository.create()'s "persist a NEW entity" contract, the same way a unique-constraint violation would surface in Postgres. update() uses collection.update() after a collection.get(ids=[id]) existence check, since Chroma's own update() has no documented raise-if-missing behaviour. collection.upsert() is intentionally unused by either method — it would make create() silently overwrite on a duplicate id, hiding exactly the bug class BaseRepository.create()'s contract exists to catch.


Wiring into an application

For a real service, wire ChromaDBPlugin (the BasePort-satisfying plugin class) through ApplicationBootstrap.compose() from openframe-core. This gives you proper lifecycle management — initialize() / health() / shutdown() — for free, instead of constructing ChromaDBRepository directly and managing the client yourself:

from openframe.core.runtime import ApplicationBootstrap
from openframe.core.ports import Capability
from openframe.adapters.db.chromadb import ChromaDBPlugin, ChromaDBSettings

settings = ChromaDBSettings(chroma_collection="items")
plugin = ChromaDBPlugin(settings, collection="items")

async with ApplicationBootstrap.compose(plugin) as app:
    repo = app.get(Capability.SEARCH)        # -> ChromaDBRepository
    results = await repo.search(query_vector=[0.1, 0.2, 0.3], k=5)
# client handle released automatically on exit (plugin.shutdown() ran)

compose() calls plugin.initialize() on entry and plugin.shutdown() on exit, so the collection handle is created, health-checked, and released without any manual lifecycle code. Requires openframe-core>=3.5.

Reach for a subclassed ApplicationBootstrap (with a configure() method) only when you need per-port config=/init_timeout= or conditional registration order; use app.registry as an escape hatch for anything neither tier covers. The ChromaDBRepository(settings) construction shown above under "Quick start" remains valid for tests, scripts, or any context that doesn't need plugin lifecycle management.

Optional: circuit breaker

openframe-core>=3.4 ships openframe.core.resilience. Compose CircuitBreakerProxy around the (optionally traced) repository — never the reverse, so a short-circuited call never produces a misleading adapter span for a call that never reached the adapter:

from openframe.core.resilience import CircuitBreakerProxy
from openframe.core.tracing import TracingProxy

repo = plugin.get_repository()
protected = CircuitBreakerProxy(
    TracingProxy(repo, prefix="repository.item"),
    failure_threshold=5,
    reset_timeout=30.0,
)

No adapter code needs to change to support this — both proxies wrap from the outside.


Configuration

All settings are read from environment variables. None are strictly required, but chroma_collection must be set (here or via collection=) before any repository operation will succeed.

Env var Type Default Description
CHROMA_HOST str localhost Chroma HTTP server hostname
CHROMA_PORT int 8000 Chroma HTTP server port
CHROMA_SSL bool False Use HTTPS when talking to the server
CHROMA_TENANT str default_tenant Tenant name
CHROMA_DATABASE str default_database Database name within the tenant
CHROMA_COLLECTION str "" Default collection name (required to use a repository)
CONNECTION_TIMEOUT float 30.0 Client creation timeout (s)
OPERATION_TIMEOUT float 10.0 Per-operation timeout (s)
MAX_RETRIES int 3 Max retry attempts

Health checks

ChromaDBRepository/ChromaDBPlugin implement the BasePort lifecycle from openframe-core.

health = await repo.health()   # PluginHealth(status=PluginStatus.READY, message="")

health() never raises — it returns PluginHealth(status=PluginStatus.FAILED, ...) on any internal failure.


Exception hierarchy

All exceptions are AdapterError subclasses from openframe.core.exceptions. Raw chromadb/httpx exceptions never escape the adapter.

Situation Exception
Cannot connect to Chroma / connection lost mid-call AdapterConnectionError
No collection name configured AdapterConfigurationError
Query failed (duplicate id, dimension mismatch, etc.) AdapterQueryError
Operation exceeded timeout AdapterTimeoutError

httpx.TimeoutException is checked before the broader httpx.TransportError at every call site — it is itself a TransportError subclass in the installed version, so checking order matters or every timeout would be misclassified as a generic connection failure. chromadb.errors.ChromaError and all its subclasses (IDAlreadyExistsError, InvalidDimensionException, NotFoundError, etc.) are flat — always query-class, never connection-class — confirmed against the installed 1.5.9's exception hierarchy.


Development

# from the package directory
pip install -e ".[dev]"
pytest tests/ -v

Protocol conformance

from openframe.core.ports import BaseVectorStore, BaseRepository

repo = ChromaDBRepository(settings, collection="items")
assert isinstance(repo, BaseVectorStore)  # True — structural check
assert isinstance(repo, BaseRepository)  # True — BaseVectorStore extends it

No inheritance from either Protocol is required or used.


License

MIT

Metadata

Release files for openframe-adapters-db-chromadb 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for openframe-adapters-db-chromadb 0.1.0
File Size Uploaded
openframe_adapters_db_chromadb-0.1.0.tar.gz 25.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for openframe-adapters-db-chromadb 0.1.0
File Interpreter ABI Platform
openframe_adapters_db_chromadb-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 45.4 kB

Release files / openframe_adapters_db_chromadb-0.1.0.tar.gz

Download URL openframe_adapters_db_chromadb-0.1.0.tar.gz
Size 25.2 kB
Tags Source
SHA-256 checksum
How to use checksums
2a247b44e86554fe251624396420c1777fd81bb7afc1291a4980b43a453efc4a
BLAKE2b-256 checksum
How to use checksums
6532a38bfafdfe174c4cb8766738c9d5ef019c657f7c331c1e3dbfa5074cc3b1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release files / openframe_adapters_db_chromadb-0.1.0-py3-none-any.whl

Download URL openframe_adapters_db_chromadb-0.1.0-py3-none-any.whl
Size 20.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
bc5065291a7515d0f6c1f8389132f8944a60ca8da51c7371c5d0239396105cb5
BLAKE2b-256 checksum
How to use checksums
6972a356a9463815639ea67a0cc3dfd818ca84cc8d3d8f9785737ad6489e67ac
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.14

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page