Abstract With the success of large-scale pre-training and multilingual modeling in Natural Language Processing (NLP), recent years have seen a proliferation of large, Web-mined text datasets covering hundreds of languages. We manually audit the quality of 205 language-specific corpora released with ...
Research Assistant
AI chat, annotations, notes & similar papers
No comments yet
Be the first to share your thoughts!