4 projects
aau-harness
Provider-neutral agent evaluation: seeded scenarios, exact scoring, BYO-agent adapters, public-value contracts, repeated runs, cost, and receipts.
lurescope
Deployable fraud-lure scoring API with a live adversarial-evasion demo — the serving companion to LureBench.
lurebench
A maintained benchmark and evaluation harness for detecting AI-generated fraud lures (phishing, BEC, romance / pig-butchering).
dspy-security-bench
Open mission-assurance commons for secure, authorized, and reviewable AI agents.