Learning text-video embeddings usually requires a dataset of video clips with\nmanually provided captions. However, such datasets are expensive and time\nconsuming to create and therefore difficult to obtain on a large scale. In this\nwork, we propose instead to learn such embeddings from video data...
Research Assistant
AI chat, annotations, notes & similar papers
No comments yet
Be the first to share your thoughts!