Dialogue models trained on human conversations inadvertently learn to generate toxic responses. In addition to producing explicitly offensive utterances, these models can also implicitly insult a group or individual by aligning themselves with an offensive statement. To better understand the dynamic...
Research Assistant
AI chat, annotations, notes & similar papers
No comments yet
Be the first to share your thoughts!