4 projects
aau-harness
Provider-neutral agent evaluation, release gates, evidence receipts, human baselines, and public Impact Capsules.
dspy-security-bench
Open mission-assurance commons for secure, authorized, and reviewable AI agents.
lurescope
Deployable fraud-lure scoring API with a live adversarial-evasion demo — the serving companion to LureBench.
lurebench
A maintained benchmark and evaluation harness for detecting AI-generated fraud lures (phishing, BEC, romance / pig-butchering).