A key solution to visual question answering (VQA) exists in how to fuse visual and language features extracted from an input image and question. We show that an attention mechanism that enables dense, bi-directional interactions between the two modalities contributes to boost accuracy of prediction ...
Research Assistant
AI chat, annotations, notes & similar papers
No comments yet
Be the first to share your thoughts!