enterprise-agentic-ai-framework
An enterprise governance framework for building single- and multi-agent AI systems in Python: authorization, guardrails, observability, secrets management, LLM gateway access, and a full production evaluation suite, all as one consistent stack instead of one-off code per project.
pip install enterprise-agentic-ai-framework
The import name is agentic_ai (the PyPI distribution name is longer
for naming reasons, the package you actually import is not):
from agentic_ai.gateway import LiteLLMGateway
Status
This is an early release. The LLM gateway and the full evaluation suite are implemented today - everything else below is scaffolded (the module exists, it's empty) and not yet usable. This table will be kept current as modules land, not written once and left stale.
| Module | Status |
|---|---|
gateway - LLM gateway (LiteLLM proxy client) |
✅ Implemented |
evaluation - Agent/LLM/Tools/Multi-Agent/RAG/Security/Platform/Memory/Drift evaluation (48 metrics, see below) |
✅ Implemented |
identity - authentication |
⏳ Planned |
governance - authorization (PEP/PDP) |
⏳ Planned |
guardrails - PII/secrets/injection/jailbreak detection |
⏳ Planned |
secrets - secrets management |
⏳ Planned |
observability - distributed tracing, structured audit |
⏳ Planned |
memory - short/long-term agent memory (storage & retrieval itself) |
⏳ Planned |
context - context engineering (write/select/compress) |
⏳ Planned |
finops - LLM cost tracking |
⏳ Planned |
security - rate limiting, abuse detection (live enforcement) |
⏳ Planned |
compliance, audit, data_governance |
⏳ Planned |
monitoring, resilience, responsible_ai |
⏳ Planned |
core - agent/tool base classes, orchestrator |
⏳ Planned |
A naming note, not a contradiction: evaluation.memory and
evaluation.security are implemented; the top-level memory and
security modules are not. They're different things - evaluation.*
measures something (was a memory retrieval correct? did PII leak in
a response you already captured?) from data you already collected,
which needs no live enforcement layer underneath it. The top-level
memory/security modules would be that live layer (actually
storing conversation memory, actually rate-limiting requests) - planned,
not built yet.
Prerequisites
This library is a client, not a server. Before any of the examples
below will work, you need a LiteLLM proxy already running somewhere
reachable - agentic_ai.gateway never installs, starts, stops, or
otherwise manages that process for you. Set it up once:
1. Install LiteLLM's proxy (a separate package from this library):
pip install 'litellm[proxy]'
2. Register at least one model. Create litellm_config.yaml -
this example routes the model name gpt-4o-mini to OpenAI, reading the
real provider key from an environment variable (never hardcode it in
the YAML):
model_list:
- model_name: gpt-4o-mini
litellm_params:
model: openai/gpt-4o-mini
api_key: os.environ/OPENAI_API_KEY
Any provider LiteLLM supports works the same way - Anthropic, Azure
OpenAI, Bedrock, a local Ollama model, etc.; only litellm_params
changes. See LiteLLM's own docs for the full provider list.
3. Set the real provider key and start the proxy:
export OPENAI_API_KEY=sk-...
litellm --config litellm_config.yaml --port 4000
4. Confirm it's actually up before writing any Python against it:
curl http://localhost:4000/health/liveliness
# -> "I'm alive!"
If that curl fails, nothing below will work either - fix connectivity
to the proxy first; agentic_ai.gateway's errors will otherwise (correctly)
just tell you the same thing: it can't reach http://localhost:4000.
Only once you have a real, running, reachable LiteLLM proxy do the examples below have anything to talk to.
Quickstart: LLM Gateway
1. Connect to it
from agentic_ai.gateway import LiteLLMGateway
# No arguments needed for the common case: connects to
# http://localhost:4000, LiteLLM's own default port.
gateway = LiteLLMGateway()
reply = gateway.complete(
model="gpt-4o-mini", # must be registered on your proxy, e.g. in litellm_config.yaml
messages=[
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "Name three benefits of distributed tracing."},
],
)
print(reply)
2. Configuring host, port, and auth
from agentic_ai.gateway import LiteLLMGateway
# Custom port - your proxy isn't on LiteLLM's default 4000
gateway = LiteLLMGateway(port=5001)
# Custom host and port - a proxy running elsewhere on your network
gateway = LiteLLMGateway(host="litellm.internal", port=8080)
# Full base_url - anything host/port can't express (TLS, a path prefix)
gateway = LiteLLMGateway(base_url="https://litellm.example.com/proxy")
# A proxy that requires a virtual key
gateway = LiteLLMGateway(api_key="sk-...") # resolve this from your own
# secrets store - the gateway
# module doesn't fetch it for you
3. The full response, not just the text
complete() is a convenience wrapper around chat_completion(), which
returns the full OpenAI-compatible response body (usage, finish_reason,
etc.) when you need more than just the message content:
result = gateway.chat_completion(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Summarize this in one sentence: ..."}],
temperature=0.2,
max_tokens=200,
)
print(result["choices"][0]["message"]["content"])
print(result["usage"])
4. Handling errors
The gateway never lets a raw network exception escape - callers get one of two exceptions, so "the proxy is down" and "the proxy rejected the request" are never conflated:
from agentic_ai.gateway import GatewayConnectionError, GatewayRequestError, LiteLLMGateway
gateway = LiteLLMGateway()
try:
reply = gateway.complete("gpt-4o-mini", [{"role": "user", "content": "hi"}])
except GatewayConnectionError:
# Nothing is listening at gateway.base_url at all - is LiteLLM
# actually running? (see Prerequisites above)
...
except GatewayRequestError as e:
# The proxy responded, but with an error (bad model name, missing
# api_key, malformed request) - e includes the proxy's own message.
print(e)
5. Cleaning up
LiteLLMGateway holds an open HTTP connection pool; close it when
you're done, or use it as a context manager:
with LiteLLMGateway() as gateway:
reply = gateway.complete("gpt-4o-mini", [{"role": "user", "content": "hi"}])
# connection pool closed automatically here
Evaluation
A complete production evaluation surface for single- and multi-agent AI
systems - 48 metrics across 9 categories, organized one folder per
category under agentic_ai.evaluation:
| Category | Import | Measures |
|---|---|---|
| Agent | agentic_ai.evaluation.agent |
Task Success/Correctness, Planning, Reasoning, Execution, Recovery, Autonomy, Loops, Lifecycle |
| LLM | agentic_ai.evaluation.llm |
Response Correctness, Groundedness, Hallucination Rate, Instruction Following, Safety/Policy Violation, Latency/Tokens/Cost |
| Tools | agentic_ai.evaluation.tools |
Selection/Argument Accuracy, Success/Failure Rate, Unnecessary Calls, Latency |
| Multi-Agent | agentic_ai.evaluation.multi_agent |
Routing, Delegation, Handoff, Coordination, Duplicate Work |
| RAG | agentic_ai.evaluation.rag |
Recall@K, Context Relevance, Groundedness, Citation Accuracy |
| Security | agentic_ai.evaluation.security |
Prompt Injection, Unauthorized Execution, PII/Cross-Tenant Leakage, Authorization Violations |
| Platform | agentic_ai.evaluation.platform |
Error Rate, Timeout Rate, Cost per Successful Task, SLA Compliance |
| Memory | agentic_ai.evaluation.memory |
Retrieval Accuracy, Consistency |
| Drift | agentic_ai.evaluation.drift |
Statistical (z-score) drift on Success/Correctness/Hallucination/Latency/Cost |
Every category is deterministic, LLM-judged, or a documented mix of
both - deterministic metrics need no LLM call at all (they read fields
you already populated); judged metrics reuse the same LLMJudge from
agentic_ai.evaluation.core, built on the gateway above, nothing else.
Deterministic - no LLM call needed
from agentic_ai.evaluation.agent import AgentTrace, compute_task_execution
traces = [
AgentTrace(run_id="r1", task="find backend jobs", task_succeeded=True),
AgentTrace(run_id="r2", task="find backend jobs", task_succeeded=False),
AgentTrace(run_id="r3", task="find backend jobs", task_succeeded=True),
]
metrics = compute_task_execution(traces)
print(metrics.success_rate) # 0.6666666666666666
LLM-judged - needs a gateway, same one as above
from agentic_ai.evaluation import LLMJudge
from agentic_ai.evaluation.llm import LLMCall, judge_response_correctness
from agentic_ai.gateway import LiteLLMGateway
judge = LLMJudge(LiteLLMGateway(), model="gpt-4o-mini")
call = LLMCall(call_id="c1", model="gpt-4o-mini", prompt="What is 2+2?", response="4")
result = judge_response_correctness(judge, call)
print(result.correct, result.score)
Every judge_*() function across every category takes an optional
system_prompt override - the built-in DEFAULT_* rubric is a real,
usable starting point, not the only valid one for every domain:
from agentic_ai.evaluation.llm import judge_response_correctness
legal_rubric = "You are a strict legal-domain correctness judge. ..."
result = judge_response_correctness(judge, call, system_prompt=legal_rubric)
Everything at once, persisted, compared over time
Agent Evaluation ties every deterministic + judged category together into one report, storable and diffable:
from agentic_ai.evaluation.agent import evaluate, JSONLEvaluationStore, compare
report = evaluate(traces, judge=judge) # runs every computable category
store = JSONLEvaluationStore("eval_runs.jsonl")
store.save(report)
baseline = store.list_runs(limit=2)[1]
regressions = compare(baseline, report) # direction-aware: knows failure_rate up is bad
For statistical drift across many runs over time (not just two points),
see agentic_ai.evaluation.drift.compute_drift() and its five named
wrappers (compute_task_success_drift, compute_correctness_drift,
compute_hallucination_drift, compute_latency_drift,
compute_cost_drift).
Every category's own trace/call shape
agent, llm, multi_agent, rag, security, and memory each have
their own input model (AgentTrace, LLMCall, MultiAgentTrace,
RAGQuery, AuthorizationCheck/TenantDataCheck,
MemoryRetrieval) - populate the one your category needs from your own
agent's logging; nothing in this library runs an agent or a retriever
for you, it only evaluates the record you hand it.
Requirements
- Python 3.10+
- A LiteLLM proxy you deploy yourself (this library is a client, not a bundled server)
License
Apache-2.0
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file enterprise_agentic_ai_framework-0.2.0.tar.gz.
File metadata
- Download URL: enterprise_agentic_ai_framework-0.2.0.tar.gz
- Upload date:
- Size: 68.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.12.8
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
70fe8b1a04f5a24518d93ff056ac1e0a4695987cb4db1047aacac68be91c11a5
|
|
| MD5 |
0dd3d4ed213207c43d66d3f9145c7200
|
|
| BLAKE2b-256 |
fff251fc160ad5b4f8da8bfce56274c8f25a8381033cc0eb4c0fef4dc05b9a27
|
File details
Details for the file enterprise_agentic_ai_framework-0.2.0-py3-none-any.whl.
File metadata
- Download URL: enterprise_agentic_ai_framework-0.2.0-py3-none-any.whl
- Upload date:
- Size: 79.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.12.8
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
eeeb4e4b5fe0a19dd898ee47536aeee6278299c5fe6016d7c87e1adf223d33fb
|
|
| MD5 |
5d4c95e68069b7866418ba89fe7aab4f
|
|
| BLAKE2b-256 |
c7abc265fb4cc243df1d37c661afc0c6cdc8268117cc4b5b0b7857d6541c971e
|