Last released Aug 17, 2026
SIREN: lightweight, plug-and-play guard models reading LLM internal representations, for harmful content and for agent trajectories.