A RoBERTa-based classifier fine-tuned on the Jigsaw Toxicity dataset to detect various forms of toxic content including threats, insults, identity attacks, and profanity. It is designed to help platforms and applications implement content moderation and create safer online environments by automatically flagging harmful text.
In the Other AI space, Roberta Toxicity Classifier takes a focused approach. Automatically identifying toxic, harmful, or inappropriate language in user-generated content. It is built as an open-source project for developers. Roberta Toxicity Classifier is open source under the Open Source license. It runs on the web and API.
garak-llm builds and maintains Roberta Toxicity Classifier, and the product first shipped in 2023.
Latest indexed changes and source events
garak-llm/roberta_toxicity_classifier verified by the PulseGate indexer