unbiased-toxic-roberta is a specialized language model based on RoBERTa, trained to identify toxic content while minimizing bias across demographic groups. Hosted on Hugging Face, it provides a more equitable alternative to standard toxicity classifiers. The model is open-source and can be integrated into applications for content moderation, research, or safety filtering.
Unbiased Toxic Roberta sits in PulseGate's Other AI category. It focuses on detecting toxic language in text while avoiding unintended demographic biases in classification. Unbiased Toxic Roberta is an open-source project aimed at AI developers and content moderators. The project is open source (Apache-2.0). The product ships for the web, the command line, and API.
Behind Unbiased Toxic Roberta is unitary, and the product first shipped in 2020. The project is developed in the open on GitHub with 1.3k stars. Among its 4 catalogued features are Toxicity Detection, Bias Reduction, and RoBERTa Architecture.
Latest indexed changes and source events
unitary/unbiased-toxic-roberta verified by the PulseGate indexer