26 projects
auraone-sdk
Python SDK and CLI for AuraOne hosted AI evaluations, with sync and async clients for templates, runs, analytics, training, and governance.
auraone-agent-studio-open
Headless Agent Studio Open protocol, trace-store, sidecar, and export CLI.
rubric-studio
Install guidance and package-name reservation for Rubric Studio Open.
lerobot-quality-gates
Check LeRobot-style metadata, episodes, sensors, videos, labels, and dataset-card disclosures.
robostudio-engine
Inspect, index, QA, cluster, probe, and export robotics datasets from a headless CLI.
robot-recovery-bench
Validate robot recovery segment JSONL and summarize intervention outcomes and timing.
vla-robustness-kit
Run deterministic, simulator-free VLA perturbation diagnostics over episode metadata and instructions.
embodiment-card
Validate and render robot embodiment metadata for dataset and VLA release review.
tool-call-replay
Normalize recorded tool-call traces into deterministic local assertions and pytest regressions.
agent-trace-card
Generate local Markdown, HTML, and JSON review cards from agent trace JSON.
datasheet-ci
Validate required Datasheet, Model Card, and Data Card headings with warning-only PII pattern checks.
mcp-risk-linter
Statically lint MCP server source and docs for capability, permission, and disclosure risks.
a2a-contract-test
Validate A2A-style agent cards and recorded task lifecycle transcripts without contacting an agent.
prompt-rubric-drift
Detect deterministic prompt and rubric drift and produce pull request review reports.
eval-adapter
Export rubric-spec rubrics and normalize evaluation result shapes across common eval frameworks.
otel-eval-bridge
Convert OTLP and Phoenix GenAI trace exports into redacted local eval cases and manifests.
judge-bench
Synthetic bias and calibration probes for testing LLM-as-judge reliability.
failure-gallery
Validate and build a synthetic agent and robotics failure gallery with reproducible review records.
synthetic-disagreement
Generate deterministic synthetic reviewer disagreement and inter-annotator agreement stress curves.
judge-card
Generate, validate, and render inspectable LLM judge disclosure cards from diagnostic results.
iaa-kit
Inter-annotator agreement metrics and bootstrap confidence intervals for review and labeling workflows.
contamination-audit
Generate item-level eval contamination signals from lexical, canary, pattern, hash, and optional embedding checks.
eval-run-manifest
Build and validate eval-run provenance manifests with deterministic directory digests.
eval-conformance-suite
Run rubric-spec adapter conformance checks and render deterministic status badges.
rubric-spec
Validate, lint, diff, and convert portable LLM evaluation rubrics with AuraOne Rubric Schema v1.
auraone-evalkit
Local-first Python CLI for AI and LLM evaluation rubrics, deterministic scoring, reviewer QA, leakage checks, and evidence reports.