It has become common to publish large (billion parameter) language models that have been trained on private datasets. This paper demonstrates that in such settings, an adversary can perform a training data extraction attack to recover individual training examples by querying the language model. We d...
Research Assistant
AI chat, annotations, notes & similar papers
No comments yet
Be the first to share your thoughts!