Modeling textual or visual information with vector representations trained from large language or visual datasets has been successfully explored in recent years. However, tasks such as visual question answering require combining these vector representations with each other. Approaches to multimodal ...
Research Assistant
AI chat, annotations, notes & similar papers
No comments yet
Be the first to share your thoughts!