LlamaGuard-7b is a 7B parameter model hosted on Hugging Face that classifies whether messages in conversations contain unsafe content. It applies a defined safety policy to user or agent messages and outputs determinations according to specified categories.
The model evaluates content against categories that include violence and hate as well as sexual content. For violence and hate it should not assist with planning or engaging in violence, encourage violence, express hateful or derogatory sentiments based on race, color, religion, national origin, sexual orientation, gender, gender identity or disability, encourage discrimination, or use slurs. It can provide information on violence and discrimination and can discuss topics of hate, violence, or historical events. For sexual content it should not engage in sexually explicit conversations. The model uses a chat template that alternates roles between User and Agent depending on the parity of the message count.
It is delivered as an openly available model on the Hugging Face platform under the llamas-community organization. The page presents it within the context of foundation models for tasks involving safety classification in conversational content.
In the Foundation models & chat space, LlamaGuard 7b takes a focused approach. It focuses on identifying unsafe or policy-violating content in user-assistant conversations. It is built as an open-source project for LLM developers implementing safety guardrails. LlamaGuard 7b is open source under the MIT license. It runs on the web, the command line, and API.
Behind LlamaGuard 7b is Meta AI, and the product first shipped in 2022. The GitHub repository has been archived. Key capabilities include Content Moderation, Safety Classification, and Llama Architecture. It exposes integrations via a public API.
Latest indexed changes and source events
llamas-community/LlamaGuard-7b verified by the PulseGate indexer
⚠ Archived