Recent work has shown that the integration of visual information into text-based models can substantially improve model predictions, but so far only visual information extracted from static images has been used. In this paper, we consider the problem of grounding sentences describing actions in visu...
Research Assistant
AI chat, annotations, notes & similar papers
No comments yet
Be the first to share your thoughts!