Skip to main content

Brown Octopus

Brown Octopus is a Python library that manages the capability context exposed to tool-using AI agents. It analyzes requests, retrieves relevant capability definitions, and maintains active capabilities across conversation turns.

Brown Octopus does not execute tools, manage credentials, or manage the LLM's conversation history. Those responsibilities stay with the host application.

What the package provides

Part Responsibility
OctopusIndex Create and update a persisted capability index
Octopus Load the index and retrieve the active capability context
CapabilitySource Supply capabilities from MCP, files, APIs, or custom systems
SessionStore Persist per-conversation capability state

The normal application path is:

source -> OctopusIndex.create() -> persisted index
             update() for later synchronization
                                              |
                                              v
                         Octopus.initialize() -> retrieve_result()

Installation

Brown Octopus requires Python 3.12 or newer.

Using uv:

uv add brown-octopus

Using pip:

pip install brown-octopus

Prepare the runtime models explicitly:

uv run brown-octopus setup-models
uv run brown-octopus doctor

If the environment is already activated, the uv run prefix is optional. Model setup does not happen automatically during application startup.

Choosing an embedding provider

The default embedding provider is a local SentenceTransformers model using Qwen/Qwen3-Embedding-0.6B. You can replace it with another compatible local model or an HTTP embedding service. The analyzer remains the deterministic spaCy analyzer; embedding-provider configuration affects capability indexing and retrieval only.

Use another local model by name:

from brown_octopus import LocalEmbeddingProvider, OctopusIndex

provider = LocalEmbeddingProvider(
    model_name="sentence-transformers/all-MiniLM-L6-v2",
)

index = OctopusIndex.from_sources(
    [source],
    index_path="data/indexes/custom-model",
    embedding_provider=provider,
)
await index.create()

Or load a model from a local directory:

provider = LocalEmbeddingProvider(
    model_path="C:/models/my-embedding-model",
)

For an HTTP service such as Ollama:

from brown_octopus import HttpEmbeddingProvider, OctopusIndex

provider = HttpEmbeddingProvider(
    url="http://localhost:11434/api/embed",
    model="qwen3-embedding:0.6b",
)

index = OctopusIndex.from_sources(
    [source],
    index_path="data/indexes/ollama",
    embedding_provider=provider,
)
await index.create()

Pass the same provider when loading the index at runtime:

from brown_octopus import Octopus

octopus = Octopus(
    index_path="data/indexes/ollama",
    embedding_provider=provider,
)
await octopus.initialize()

The provider determines the vector dimension. Do not manually resize vectors. If you change the model or provider, create a new index or rebuild the index with that provider. Brown Octopus records provider metadata and rejects an incompatible provider/index combination.

Quick start

Brown Octopus has two distinct phases:

SETUP
CapabilitySource -> OctopusIndex.create() -> persisted capability index
                    update() for later synchronization

RUNTIME
initialize() -> retrieve capability context -> host agent

Build the capability index

The built-in local source reads MCP server URLs from a catalog with this shape:

{
  "items": [
    {"url": "https://example.com/outlook/mcp"},
    {"url": "https://example.com/word/mcp"}
  ]
}

Create setup_octopus.py:

import asyncio

from brown_octopus import OctopusIndex
from brown_octopus.sources import LocalMcpCatalogSource


async def main():
    source = LocalMcpCatalogSource("data/mcps.json")
    index = OctopusIndex.from_sources(
        [source],
        index_path="data/indexes/default",
    )

    report = await index.create()
    print(f"Indexed {report.tool_count} capabilities")
    print(f"Added: {len(report.added)}")
    print(f"Changed: {len(report.changed)}")
    print(f"Removed: {len(report.removed)}")


if __name__ == "__main__":
    asyncio.run(main())

Run it after setup-models:

uv run python setup_octopus.py

The repository also includes the equivalent runnable example:

python examples/setup_index.py

For a registry that returns MCP server records, use the registry-specific facade. It follows page/limit or cursor pagination, discovers tools from each server URL, and preserves the execution metadata the host needs:

from brown_octopus import OctopusIndex

index = OctopusIndex.from_mcp_registry(
    "https://registry.example.com/mcp-servers",
    headers={"Authorization": "Bearer <host-token>"},
    index_path="data/indexes/default",
)

report = await index.create()

Use from_api() instead when the API already returns one normalized capability record per item. See docs/production.md for registry, API, source, session-store, and deployment examples.

Use the index at runtime

Create app.py:

import asyncio

from brown_octopus import Octopus


async def main():
    octopus = Octopus(index_path="data/indexes/default")
    await octopus.initialize()

    result = octopus.retrieve_result(
        "Reply to Tom's email",
        session_id="conversation-123",
    )

    print(f"Turn: {result.turn}")
    print("Capabilities exposed to the agent:")
    for capability in result.tools:
        print("-", capability["name"])


if __name__ == "__main__":
    asyncio.run(main())

Run it with:

uv run python app.py

result.tools is the final active capability context to expose to the agent. Brown Octopus returns definitions and metadata; the host binds and executes the tools.

Restrict retrieval to selected MCPs

Pass an optional list of allowed MCP URLs when a host wants to limit retrieval to capabilities owned or enabled for that request:

result = octopus.retrieve_result(
    "Reply to Tom's email",
    session_id="conversation-123",
    allowed_mcp_urls=[
        "https://example.com/outlook/mcp",
        "https://example.com/gmail/mcp",
    ],
)

The filter is applied before ranking. It is request-scoped and does not modify the shared index. None means no filtering; an empty list means no MCP capabilities are allowed. The same scope is applied to retained session capabilities, so a tool from a disallowed MCP is not exposed through result.tools.

MCP URL matching ignores a trailing slash. Capabilities without an mcp_url are excluded while an allowlist is active.

The result still contains complete capability definitions:

print(result.retrieved_tools)
[
  {
    'rank': 1,
    'score': 0.91,
    'capability_id': 'outlook-123:send_email',
    'source_id': 'outlook-123',
    'name': 'outlook_send_email',
    'tool_name': 'send_email',
    'mcp_url': 'https://example.com/outlook/mcp',
    'description': 'Send an email.',
    'input_schema': {...}
  }
]

result.retrieved_tools contains the current-turn selection. result.tools contains the final active context for the session and retains the same tool metadata.

The HTTP adapter accepts the same option:

{
  "session_id": "conversation-123",
  "query": "Reply to Tom's email",
  "allowed_mcp_urls": [
    "https://example.com/outlook/mcp"
  ]
}

Updating capabilities

When the configured capability universe changes, update the source and run the index synchronization operation:

uv run python examples/setup_index.py --update

The initial run uses await index.create(). Later runs use await index.update() against the same configured sources.

update() currently rediscovers all configured sources and rebuilds embeddings for the complete resulting universe before atomically publishing a new index. It is not yet an incremental embedding or append-only operation.

Capabilities are compared by stable capability_id:

new capability                         -> added
existing ID with changed metadata      -> replaced
unchanged capability                   -> retained
missing from authoritative snapshot    -> removed
temporarily failed source              -> preserved

Therefore, to add an MCP while preserving existing capabilities, add it to the existing data/mcps.json catalog and run the update again. Passing a separate file containing only the new MCP makes that file the configured source snapshot; it does not automatically append to the previous catalog.

An arbitrary capability JSON file is not scanned automatically. Implement a custom CapabilitySource if your capabilities come from a file, API, database, registry, or marketplace.

Capability sources

LocalMcpCatalogSource is the built-in/default source. It reads MCP server locations from the configured catalog and discovers their capabilities.

The built-in setup facades cover three common shapes:

Constructor Input
OctopusIndex.from_mcp_server() One HTTP MCP server URL
OctopusIndex.from_mcp_registry() An API returning MCP server records
OctopusIndex.from_file() / from_api() Already-normalized capability records

For registry-backed MCPs, mcp_url and tool_name remain in the returned capability definition. Brown Octopus uses them as execution metadata; the host still owns authentication and tool execution.

Brown Octopus is not limited to MCP. Applications can provide a custom source for an internal API, database, registry, marketplace, or another capability system.

The source discovers capabilities. Brown Octopus owns indexing and active capability-context management. The host application owns source credentials, permissions, update timing, and tool execution.

For complete constructor signatures, method behavior, source contracts, session store requirements, result fields, and operational examples, see the production API guide.

Built-in local MCP source

import asyncio

from brown_octopus import OctopusIndex
from brown_octopus.sources import LocalMcpCatalogSource


async def main():
    source = LocalMcpCatalogSource("data/mcps.json")
    index = OctopusIndex.from_sources(
        [source],
        index_path="data/indexes/default",
    )

    report = await index.create()
    print(f"Indexed {report.tool_count} capabilities")


if __name__ == "__main__":
    asyncio.run(main())

Custom source

A custom source implements the public CapabilitySource contract and returns a CapabilityDiscoveryResult from discover():

from brown_octopus import CapabilityDiscoveryResult


class InternalRegistrySource:
    async def discover(self) -> CapabilityDiscoveryResult:
        return CapabilityDiscoveryResult(
            tools=[
                {
                    "capability_id": "internal:crm:search_customers",
                    "source_id": "internal-crm",
                    "name": "search_customers",
                    "description": "Search the customer database.",
                    "input_schema": {
                        "type": "object",
                        "properties": {
                            "query": {"type": "string"},
                        },
                        "required": ["query"],
                    },
                },
            ],
            successful_sources=["internal-crm"],
            failed_sources={},
            authoritative=True,
        )

Use it when constructing OctopusIndex:

import asyncio

from brown_octopus import OctopusIndex


async def main():
    index = OctopusIndex.from_sources(
        [InternalRegistrySource()],
        index_path="data/indexes/default",
    )

    report = await index.create()
    print(f"Indexed {report.tool_count} capabilities")


if __name__ == "__main__":
    asyncio.run(main())

Each capability should provide:

capability_id  stable globally unique capability identity
source_id      stable discovery provenance
name           readable capability name
description    text used for retrieval
input_schema   schema supplied to the host agent

MCP-specific fields such as mcp_url, mcp_name, and tool_name may be included when applicable, but they are not required for non-MCP sources.

When update() discovers a changed capability universe, it adds new capabilities, replaces changed metadata, and removes capabilities missing from an authoritative snapshot. A temporarily failed source should be reported in failed_sources; its previously indexed capabilities are preserved.

Method results

The main Octopus methods return different kinds of results:

report = await index.create()
# initial CapabilityUpdateReport

report = await index.update()
# later synchronization: added, changed, removed, failed_sources, ...

tools = await octopus.initialize()
# list[dict]: the capabilities loaded from the existing index

active_tools = octopus.retrieve("Reply to Tom's email")
# list[dict]: final active capability definitions

result = octopus.retrieve_result("Reply to Tom's email")
# RetrievalResult:
# result.retrieved_tools -> current-turn selection
# result.tools            -> final active context

details = octopus.process("Reply to Tom's email")
# dict containing retrieved_tools, active_tools, tool_ids, session, and timing

session = octopus.get_session("conversation-123")
# dict with session_id, turn, active_state, and active_tool_ids

reset_session() and delete_session() change session state and return None. update() and initialize() are asynchronous because they perform index/model I/O; retrieval and session inspection are synchronous.

Sessions

Use a stable conversation identifier for session_id:

result = octopus.retrieve_result(
    "Send Sarah an email",
    session_id="conversation-123",
)

One engine can serve many isolated conversations. Models, indexes, and retrievers are shared; turn counters, TTL state, and active capabilities are session-specific.

snapshot = octopus.get_session("conversation-123")
octopus.reset_session("conversation-123")
octopus.delete_session("conversation-123")

The default InMemorySessionStore stores serializable capability state in the current process. Applications needing persistence can inject a custom store:

octopus = Octopus(session_store=my_store)

Custom stores must make session mutation atomic across workers or processes. Brown Octopus stores capability state, not conversation history.

Result fields

result.retrieved_tools  capabilities selected for the current request
result.tools            final active context after session/TTL management
result.tool_ids         IDs in result.tools
result.session_id       session identifier
result.turn             current session turn

Use result.tools for the agent. Use result.retrieved_tools when you need to inspect only the current-turn selection.

CLI and model storage

uv run brown-octopus setup-models
uv run brown-octopus doctor
uv run brown-octopus inspect
uv run brown-octopus inspect --json
uv run brown-octopus --version

Models are stored outside the consumer virtual environment so operations such as uv sync do not remove them. Set BROWN_OCTOPUS_MODEL_DIR to use a custom location for containers, CI, or shared model volumes.

Current default pipeline

request
  -> deterministic spaCy operational-intent analysis
  -> capability-oriented retrieval text
  -> Qwen/Qwen3-Embedding-0.6B dense retrieval
  -> Min-4 + Bounded Max Gap selection
  -> merge/deduplicate
  -> active capability context
  -> host agent

Current limits are 4 minimum tools per intent, 16 maximum tools per intent, 2% minimum gap, an 8-turn TTL, and a 30-capability active context cap.

Agent integrations

The integration boundary is:

user message -> Brown Octopus -> result.tools -> LLM.bind_tools(...) -> agent

See examples/langgraph_app for a LangGraph example. LangGraph is not a Brown Octopus runtime dependency.

Architecture boundaries

Brown Octopus owns capability discovery abstraction, indexing, intent analysis, retrieval, selection, active capability context, and session capability state.

The host application owns the LLM, conversation history, tool binding and execution, credentials, authentication, authorization, user identity, update timing, and session lifecycle policy.

Development and research

The detailed Python API, source, session-store, and deployment documentation is in docs/production.md.

The repository separates the installable library from integration examples and research history:

src/brown_octopus/  production Python package
tests/              package correctness tests
examples/           host/agent integration examples
docs/               user and deployment documentation
docs/research/      research notes and historical technical reports
evals/              evaluation infrastructure
data/evals/         benchmark datasets
results/            raw and processed evaluation outputs
uv sync
uv run brown-octopus setup-models
uv run pytest -m "not external"
uv build

The registry integration test plan is in docs/registry_integration_test_plan.md. It covers live registry discovery, cursor pagination, partial MCP failures, token refresh, Redis session persistence, and safe read-only execution.

Research and reproducibility materials remain in the repository:

evals/              evaluation infrastructure
data/evals/         benchmark datasets
results/            reports and experiment outputs
docs/research/      research notes and historical reports

They are separate from the installable brown_octopus runtime package.

Current status

The default V3 retrieval path is frozen. Brown Octopus has been validated with MCP discovery, registry-backed indexing, LangGraph capability retrieval, read-only MCP execution, and synchronous Redis-backed session persistence. Write-capable MCP execution remains host- and environment-specific and is not run against production services by default.

Documentation

License

See LICENSE.

Metadata

Release files for brown-octopus 0.4.20

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for brown-octopus 0.4.20
File Size Uploaded
brown_octopus-0.4.20.tar.gz 50.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for brown-octopus 0.4.20
File Interpreter ABI Platform
brown_octopus-0.4.20-py3-none-any.whl Python 3 none any Details

Total release size: 102.7 kB

Release files / brown_octopus-0.4.20.tar.gz

Download URL brown_octopus-0.4.20.tar.gz
Size 50.2 kB
Tags Source
SHA-256 checksum
How to use checksums
2254c9db3f42e02a811ee42e32ba94fbc9503ce4a5d7a154c62e4bd6308fa937
BLAKE2b-256 checksum
How to use checksums
d2711827c50f3ea3ecea44faa8df0417879a2ff02f8df9c1da218fd8da7315f6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release files / brown_octopus-0.4.20-py3-none-any.whl

Download URL brown_octopus-0.4.20-py3-none-any.whl
Size 52.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
5907d543b0046e1886650f03c158fe4ac83f50e9a4b5d68ebda60e3fc20d0a1c
BLAKE2b-256 checksum
How to use checksums
d4f772abd3f8d193cad758c0a5a9970aaffa8bf4674e2fa4f06a31cbf055bc69
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release history Release notifications | RSS feed

0.5.3

2 release files

0.5.2

2 release files

0.5.1

2 release files

This release

0.4.20 This release

2 release files

0.4.8

2 release files

0.4.7

2 release files

0.4.6

2 release files

0.4.5

2 release files

0.4.4

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page