We present a general approach towards controllable societal biases in natural language generation (NLG). Building upon the idea of adversarial triggers, we develop a method to induce societal biases in generated text when input prompts contain mentions of specific demographic groups. We then analyze...
Research Assistant
AI chat, annotations, notes & similar papers
No comments yet
Be the first to share your thoughts!