WildGuard is an open-source model developed by AllenAI for classifying harmful content, detecting jailbreak attempts, and performing safety evaluations on text. It is designed to help developers build safer AI applications by identifying problematic inputs before they reach production LLMs. The model is available on Hugging Face and can be used with standard transformer libraries.
Wildguard sits in PulseGate's Foundation models & chat category. It focuses on identifying and filtering unsafe, harmful, or adversarial prompts in LLM applications. It is built as an open-source project for developers. Wildguard is open source under the Open Source license. The product ships for the web, the command line, and API.
Behind Wildguard is AllenAI, based in the United States, and the product first shipped in 2024. Key capabilities include content moderation, jailbreak detection, and safety classification.
Latest indexed changes and source events
allenai/wildguard verified by the PulseGate indexer
Other apps tracked under the same category.