InfoLang-backed memory storage for CrewAI's unified Memory system.
Project description
infolang-crewai
InfoLang-backed storage for CrewAI's unified Memory system.
A note on CrewAI's API. Older CrewAI docs describe
memory_config={"provider": ...}and anExternalMemoryclass. As ofcrewai>=1.15, both are gone. CrewAI now has a singleMemoryclass (crewai.memory.unified_memory.Memory) backed by a pluggableStorageBackendprotocol (crewai.memory.storage.backend.StorageBackend). This package targets that interface — the one actually shipped in the installed package — not the older docs. See Known limitations for what that means in practice.
Install
The infolang SDK is on PyPI. infolang-crewai itself isn't published yet — install it
from GitHub:
pip install infolang
pip install "infolang-crewai @ git+https://github.com/InfoLang-Inc/infolang-crewai.git"
For local development against this repo:
git clone https://github.com/InfoLang-Inc/infolang-crewai.git
cd infolang-crewai
pip install -e ".[dev]"
Quickstart
from crewai import Agent, Crew, Task
from crewai.memory.unified_memory import Memory
from infolang_crewai import InfoLangEmbedder, InfoLangStorage
# One InfoLangStorage per Crew is the common case: it owns one InfoLang
# namespace (memory bank) for everything the crew remembers.
memory = Memory(
storage=InfoLangStorage(api_key="il_live_...", namespace="research-crew"),
# InfoLangEmbedder MUST be paired with InfoLangStorage -- see "Why the
# custom embedder?" below. Leaving this out (or using a real embedding
# model) makes InfoLangStorage.search() raise InfoLangCrewAIConfigError.
embedder=InfoLangEmbedder(),
)
researcher = Agent(
role="Researcher",
goal="Find and remember useful facts",
backstory="...",
)
task = Task(description="Research X and remember what you learn", agent=researcher)
crew = Crew(agents=[researcher], tasks=[task], memory=memory)
crew.kickoff()
Crew(memory=True) still works and builds CrewAI's default local (LanceDB)
store; pass a Memory(storage=InfoLangStorage(...), embedder=InfoLangEmbedder())
instance instead to route through InfoLang.
Per-agent namespaces
Two ways to isolate memory per agent, at different levels:
InfoLang-level (separate memory bank per agent) — construct one
InfoLangStorage/Memory per agent, or use namespace_resolver:
storage = InfoLangStorage(
api_key="il_live_...",
namespace="research-crew", # fallback / used by get_record() and update()
namespace_resolver=lambda scope: f"research-crew-{(scope or 'shared').strip('/').replace('/', '-') or 'shared'}",
)
namespace_resolver receives the CrewAI scope path (e.g. "/agent/researcher")
for every call that carries one (save, search, list_records, delete,
count, get_scope_info, list_scopes, list_categories, reset).
get_record() and update() take no scope argument in the CrewAI protocol,
so they always use the static namespace= — not the resolver. See
Known limitations for why namespace isolation does not
currently separate data on the hosted runtime.
CrewAI-level (within one InfoLang namespace) — use CrewAI's own
MemoryScope, which this package supports natively since it preserves the
full scope path on every record and filters on it client-side:
from crewai.memory.memory_scope import MemoryScope
researcher.memory = MemoryScope(root_path="/agent/researcher").bind(memory)
Why this exists
Reliability. Mem0's CrewAI integration has several open issues that all come down to the same thing: an unvalidated assumption about a shape (a dict key that isn't always there, a config value that silently isn't applied, a list assumed where a string arrived). This package's test suite is written specifically against that failure shape — see Testing — not as a general-purpose claim of superiority, but because those are documented, reproducible bugs and this package is tested against reproductions of them:
- Config not silently ignored. Every constructor argument is validated
at
InfoLangStorage(...)construction time (bad namespace, non-callablenamespace_resolver, conflictingclient=/api_key=, etc. all raiseInfoLangCrewAIConfigErrorimmediately, not "memory looks empty later"). - No blind indexing into API responses. Every InfoLang response is read
through the typed
infolangSDK models (or explicit multi-key lookups for the one loosely-typed route,GET /v1/memories); nothing doesresult["memory"]-style direct key access against an untyped dict. Malformed or missing fields raiseInfoLangCrewAIResponseErrorwith a clear message, or degrade gracefully to defaults on the read path — never a rawKeyError. - Malformed response shapes don't crash a call. A
hits/chunksfield that arrives as something other than a list is caught (via the SDK's own response validation) and turned into a clearInfoLangCrewAIResponseError, not a rawpydantic.ValidationErrorescapingsearch(). - Provider/auth switching is exercised. Tests cover constructing
InfoLangStorageagainst bothapi_keyanddev_keyauth, and swapping the pairedembedder=whilestorage=stays fixed, to catchAttributeError-class bugs from code that only worked under one configuration.
Why the custom embedder?
CrewAI's Memory computes an embedding from your query/content with
self.embedder before it ever calls the storage backend — StorageBackend
methods only receive list[float], never the original string
(crewai.memory.storage.backend.StorageBackend.search). InfoLang embeds
server-side from raw text (POST /v1/recall takes query: str, not a
vector). To bridge that gap without losing InfoLang's real semantic search,
InfoLangEmbedder doesn't compute a real embedding — it reversibly encodes
each string as a vector of UTF-8 byte values in [0, 1], and
InfoLangStorage decodes that vector back into text before calling InfoLang.
Always pass both together. InfoLangStorage.search() validates every
incoming vector and raises InfoLangCrewAIConfigError if it doesn't look
like InfoLangEmbedder's output (e.g. a real OpenAI embedding was wired in
by mistake) instead of silently returning wrong or empty results.
Config reference
InfoLangStorage(...) keyword arguments:
| Argument | Default | Notes |
|---|---|---|
client |
None |
Pass a pre-built infolang.InfoLang instance. Mutually exclusive with api_key/dev_key/base_url. |
api_key |
None |
Managed cloud API key (il_live_...). Falls back to INFOLANG_API_KEY env var via the SDK if omitted. |
dev_key |
None |
Self-hosted dev key (key:namespace), targets 127.0.0.1:8766. |
base_url |
None |
Override the API base URL. |
namespace |
"default" |
InfoLang bank name. Must be non-empty. Used by every call, and always used by get_record()/update(). |
namespace_resolver |
None |
Callable[[str | None], str] mapping a CrewAI scope path to a namespace, for calls that carry scope. |
source |
None |
Default source tag passed to remember() when a record doesn't set its own. |
search_pool |
50 |
top_k sent to /v1/recall is max(limit, search_pool), oversampling so client-side category/metadata filtering still has enough candidates. |
list_page |
500 |
Page size for the GET /v1/memories scans backing list_records/count/delete/update/get_record/scope introspection (see below). |
Known limitations
Hosted-runtime namespace behavior (production caveat, not a bug in this
package). POST /v1/remember on the hosted InfoLang runtime currently
ignores the namespace field in the request body — every write lands in the
default namespace regardless of what is requested. A server-side fix is
pending. Until it ships:
- Namespace-level isolation between crews/agents (the
namespace=/namespace_resolver=mechanisms above) does not actually separate data on the hosted runtime today — everything ends up indefault. Memory.recall()'s server-side namespace filter is therefore also not useful for separating crews right now, since nothing outsidedefaultgets written.- This package's
MemoryScope-based scope filtering (client-side, from data embedded in the stored text) is not affected by this bug and works correctly today, even against the shareddefaultnamespace. InfoLangStoragestill sends thenamespaceyou configure on every call — no code changes will be needed here once the server-side fix lands.
remember_batch is unsupported on the deployed runtime. The infolang
SDK exposes remember_batch() (via POST /v1/execute with a
remember_batch op), but the currently deployed runtime rejects that op.
InfoLangStorage.save() therefore issues one POST /v1/remember per record,
not a batch call — this is deliberate, not an oversight, and will not change
until the runtime accepts the op.
No native "get by id" or "filter by metadata" route. InfoLang's REST
contract is POST /v1/recall, POST /v1/remember, GET /v1/memories,
DELETE /v1/memories/{id} — there is no query-by-id or server-side metadata
filter. get_record(), update(), delete() (beyond simple cases),
list_records(), count(), get_scope_info(), list_scopes(), and
list_categories() are all built on a full namespace scan
(GET /v1/memories, page size list_page) with local filtering — O(N) per
call, the same pattern the infolang SDK itself uses internally for
reset_namespace(). This is fine for the bounded memory footprint of a
crew/agent; it is not a substitute for a real secondary index at very large
scale.
InfoLangEmbedder is not a real embedding. See
"Why the custom embedder?". Do not use it outside
this package, and do not mix it with a different embedder on the same
Memory instance.
Metadata is embedded as text. Because /v1/remember has no metadata
field, InfoLangStorage appends a compact JSON sidecar to the text it
stores, after a sentinel that keeps it out of the human-readable content.
InfoLang's server-side embedding sees that whole string, so very large
metadata dicts add non-semantic noise to what gets embedded — keep
metadata small.
Testing
pip install -e ".[dev]"
pip install -e "../sdk-python"
ruff check .
mypy
pytest
Tests use respx to mock the InfoLang HTTP API — no network access and no
real API key are required to run the suite. Coverage gate: 90%
(--cov-fail-under=90, see pyproject.toml).
Examples
See examples/minimal_crew.py for a runnable crew wired to InfoLangStorage.
License
Apache-2.0
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file infolang_crewai-0.1.0.tar.gz.
File metadata
- Download URL: infolang_crewai-0.1.0.tar.gz
- Upload date:
- Size: 308.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.5
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
24e2c0b13398453e5956161af104983a53f05b59934b15dfdb79d02a41d5c211
|
|
| MD5 |
a133fb474c8ce360745848dbc4510042
|
|
| BLAKE2b-256 |
7b162ac75f4e2f3e878b4e4679e7298ec88cb2c318361fa5870761df11f417da
|
File details
Details for the file infolang_crewai-0.1.0-py3-none-any.whl.
File metadata
- Download URL: infolang_crewai-0.1.0-py3-none-any.whl
- Upload date:
- Size: 20.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.5
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
771b946c1fbf0b67dd0189ce5f3784201012955473abca9adbc4112f87b2ac89
|
|
| MD5 |
d6da0cad58c4660b9614a1cbe931ab33
|
|
| BLAKE2b-256 |
18db3ccc3ec666fb9754a5335dc9f3879bf7ee92ec38ec45f4010523315aec04
|