Skip to main content

The Adaptive Intelligence Layer for AI Agents — eval, train, memory, environments.

Project description

Aegis

CI Repo Hygiene Docs Link Check CodeQL Python 3.11+ License: Apache 2.0

The Closed-Loop Intelligence Engine for AI Agents — eval, train, and deploy smarter agents.

Aegis by Metronis, Inc. is an open-source framework that evaluates AI agents across 100+ dimensions, identifies weaknesses, trains improved models via RL, and proves the improvement with rigorous before/after evaluation.

Quick links: Quick Start | Examples | Community Guide | Security Model | Maintainer Triage | Contributing | Governance | Releasing | Support

Benchmark Evidence Tiers

Aegis exposes three different benchmark evidence modes and they are not interchangeable:

  • internal proxy: built-in *-proxy suites are for internal regression tracking and ablations.
  • public proxy: built-in *-heldout suites use held-out slices and benchmark-native metrics, but they are still not official leaderboard-grade evidence.
  • claim-grade: only manifest-backed claim suites are intended for externally claimable benchmark results.

The built-in legal alias currently resolves to legal-heldout, which is a held-out proxy suite, not an official claim-grade legal benchmark. Treat strict built-in runs as honest proxy evidence unless the suite contract explicitly reports claim eligibility. The repo does not bundle claim-grade legal suite JSON by default; install real manifests under benchmarks/claim_suites/ or set AEGIS_BENCHMARK_SUITE_DIR to an external manifest directory.

The Pipeline

Traces In → Eval → Diagnose Weaknesses → Spin RL Environments →
Train (GRPO + Continuous Rewards) → Store Results to Memory →
Train with Tools + Self-Managed Memory → Final Eval (Prove Improvement)
flowchart LR
    A["Agent Traces"] --> B["Aegis Eval\n100+ dimensions"]
    B --> C["Weakness\nDiagnosis"]
    C --> D["RL Environments\nauto-generated"]
    D --> E["GRPO Training\ncontinuous rewards"]
    E --> F["Aegis Memory\nextract strategies"]
    F --> G["Tool + Memory\nTraining"]
    G --> H["Final Eval\nprove improvement"]
    H --> A

Three Products

Product What it does Status
Aegis Eval 100+ dimensions across 7 tiers, triangulated scoring (rule + semantic + LLM judge), legal, finance, and safety plugins Working
Aegis Train GRPO-based RL with continuous reward functions, environment factory, Observatory monitoring Building
Aegis Memory 12 operations, knowledge graph, vector store, provenance tracking, 7 time horizons Working

Installation

pip install -e ".[dev,operator]"

Install profiles:

  • .[dev,operator] — lean local dev path for the CLI, API, and hosted control plane work.
  • .[dev,research] — adds scoring and dataset tooling for eval-focused research work.
  • .[dev,launch] — adds the strict launch path dependencies for real scoring and training.
  • .[dev,all] — the heaviest internal profile; only use it when you explicitly need every optional integration.

If you are working specifically on the launch-critical strict path, prefer:

pip install -e ".[dev,launch]"

For the exact launch checklist, see docs/launch-readiness.md and docs/launch-gap-audit.md.

For a fast public-repo sanity pass before you push, run:

make oss-check

That check now also verifies that GitHub workflow actions are pinned to immutable commit SHAs instead of moving tags or branches.

To discover the main local workflows, run:

make help

To sync the repo's canonical GitHub labels from the checked-in taxonomy, run:

make labels-sync

To audit the live GitHub repository settings against the launch policy, run:

make github-launch-audit

To preview the repo-level GitHub settings patch we want before public launch, run:

make github-launch-apply

Quick Start

Evaluate an agent

from aegis import Evaluator, EvalConfig

evaluator = Evaluator(config=EvalConfig(dimensions="all"))
result = evaluator.run()

print(f"Overall score: {result.overall_score:.2%}")
for tier_name, tier_score in result.tier_scores.items():
    print(f"  {tier_name}: {tier_score:.2%}")

CLI

aegis guide                           # Show the recommended launch-grade workflows
aegis signup --name "Your Name"       # Create a hosted account and save the CLI session
aegis onboard --domain legal          # Run the hosted readiness flow and get the next commands
aegis login --json                    # Save a hosted session and emit machine-readable output
aegis config show                     # Inspect resolved URLs, state paths, and saved session metadata
aegis doctor --json                   # Run launch-path health checks for CI and scripts
aegis completion show --shell zsh     # Print an explicit completion script for your shell
aegis completion install --shell zsh  # Install completion without relying on shell autodetect
aegis app url --path training         # Resolve a hosted app URL without opening a browser
aegis app open --print-url            # Open the app or print the URL for remote terminals
aegis eval dimensions --json          # Inspect dimensions in automation-safe JSON
aegis eval trend --dimension retention_accuracy --json
aegis eval benchmark-validate --manifest benchmarks/claim.json --json
aegis trace list --json               # List stored traces with full ids for scripts
aegis trace import traces.jsonl --json
aegis data list --json                # Inspect cached datasets programmatically
aegis memory health --json            # Inspect memory subsystem status programmatically
aegis memory audit --json             # Inspect memory audit events programmatically
aegis eval run --config eval.yaml     # Run evaluation
aegis eval dimensions                 # List all registered dimensions
aegis eval benchmark-list             # Inspect proxy vs claim-grade suite status
aegis train start --domain legal --backend verl --run-now
aegis train list                      # Inspect hosted or local draft jobs
aegis train watch --job-id JOB_ID     # Follow live job state
aegis train metrics --job-id JOB_ID   # Inspect rewards and stages
aegis train adapter list              # Manage trained adapters
aegis logout                          # Clear the saved CLI session

The hosted standard backend is intentionally labeled as simulated/deterministic in CLI output. Use --backend verl on a real GPU-backed runtime for actual training instead of placeholder rollouts.

Launch Runtime Posture

make aws-launch-preflight
make aws-launch-debug-ssh
make aws-launch-refresh-ssh
make aws-launch-dns-targets
make launch-manual-gpu-sync LAUNCH_MANUAL_GPU_HOST=<gpu-host>
make launch-manual-gpu-bootstrap LAUNCH_MANUAL_GPU_HOST=<gpu-host>
make launch-manual-gpu-remote-preflight LAUNCH_MANUAL_GPU_HOST=<gpu-host>
make launch-manual-gpu-remote-baseline LAUNCH_MANUAL_GPU_HOST=<gpu-host>

For launch, AWS is the public hosting and deployment plane, while the flagship strict proof flow should be run manually from a real GPU worker outside AWS, such as a Nebius machine. The checked-in launch configs now use backend: local intentionally so they fail on laptops and non-GPU hosts rather than silently falling back to simulated training.

Architecture

src/aegis/
├── core/           # Types, config, settings (5 files)
├── cli/            # Typer CLI (1 file)
├── adapters/       # OpenAI, Anthropic, LangChain, REST (6 files)
├── api/            # FastAPI + eval/training/memory routes (9 files)
├── data/           # CUAD, LegalBench, FinanceBench loaders (5 files)
├── ingestion/      # Document parsing pipeline (6 files)
├── eval/           # Engine, 7-tier dimensions, scorers, judges (27 files)
├── environments/   # Legal & finance tool-use RL environments (4 files)
├── training/       # GRPO engine, rewards, rollouts, optimizers (16 files)
├── memory/         # 12 ops, graph, vectors, provenance (16 files)
├── plugins/        # Legal (18 dims) + Finance (20 dims) + Safety (12 dims) (6 files)
├── observatory/    # Goodhart detection, efficiency tracking (4 files)
├── security/       # RBAC (2 files)
└── store/          # SQLite, Postgres, Neo4j (5 files)

~100 source files. One pipeline. No bloat.

Adapters

Adapter Framework
OpenAIAdapter OpenAI Assistants + Chat Completions
AnthropicAdapter Anthropic Messages API
LangChainAdapter LangChain agents
RESTAdapter Any REST API

Domain Plugins

  • Legal: 18 dimensions, CUAD dataset, contract clause extraction, citation verification
  • Finance: 20 dimensions, FinanceBench + SEC EDGAR, numerical accuracy, formula validation

Documentation

Topic Link
Quick Start docs/quickstart.md
Examples docs/examples.md
Open-Source Release docs/open-source-release.md
Security Model docs/security-model.md
Maintainer Triage docs/triage.md
Architecture docs/architecture.md
Community Guide docs/community.md
Master Implementation Plan docs/master-implementation-plan.md
CLI Reference docs/cli-reference.md
API Reference docs/api-reference.md
Dimensions docs/dimensions.md
Scoring docs/scoring.md
Adapters docs/adapters.md
Plugins docs/plugins.md
Support SUPPORT.md

Contributing

See CONTRIBUTING.md.

Public community norms live in CODE_OF_CONDUCT.md, and security disclosures should follow SECURITY.md. Public support and issue-routing guidance live in SUPPORT.md. Repository governance and maintainer expectations live in GOVERNANCE.md and MAINTAINERS.md. Release procedure lives in docs/releasing.md. Public trust boundaries live in docs/security-model.md. Maintainer triage flow lives in docs/triage.md.

Open-Source Posture

This repository is intended to be public open-core source code.

That means:

  • the engine, CLI, API, reference dashboard, and self-host references live here
  • production secrets, private benchmark manifests, customer data, and managed service credentials do not
  • benchmark claims should stay aligned with the evidence tier reported by the runtime and manifests

For the release policy behind that boundary, see docs/open-source-release.md and docs/open-core-boundary.md.

License

Apache License 2.0. See LICENSE.

Built by Metronis, Inc.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

aegis_eval-0.1.1.tar.gz (2.1 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

aegis_eval-0.1.1-py3-none-any.whl (941.2 kB view details)

Uploaded Python 3

File details

Details for the file aegis_eval-0.1.1.tar.gz.

File metadata

  • Download URL: aegis_eval-0.1.1.tar.gz
  • Upload date:
  • Size: 2.1 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.13

File hashes

Hashes for aegis_eval-0.1.1.tar.gz
Algorithm Hash digest
SHA256 16d094e2e4b68b7b1bcf291dcf19ad5bf924f2e4f9d7ef191d0322ccdb7ea9bc
MD5 fe9d6ba7c2c1bc3acd1d6486fbd829d2
BLAKE2b-256 4338102d9c4406a0c462620ff8471503868f3e3509873c8ae1f624201a8a33ed

See more details on using hashes here.

File details

Details for the file aegis_eval-0.1.1-py3-none-any.whl.

File metadata

  • Download URL: aegis_eval-0.1.1-py3-none-any.whl
  • Upload date:
  • Size: 941.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.13.13

File hashes

Hashes for aegis_eval-0.1.1-py3-none-any.whl
Algorithm Hash digest
SHA256 38560fe45ce9d101231d38ca5e87ccf858baff0edb7fb6b834f00b48288aad16
MD5 e65e3f24ecdb6672d5e634eb1b6bf99c
BLAKE2b-256 b31d7971fe4c539324f914ad77ed167921434424d598a87db61568c80ff7ca71

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page