Last released May 24, 2026
A fine-tuned ML classification model that detects and blocks jailbreak attempts and prompt injection threats in LLM-based conversational systems.
Supported by