Large Vision-Language Model (LVLM) has enhanced the performance of various downstream tasks in visual-language understanding.Most existing approaches encode images and videos into separate feature spaces, which are then fed as inputs to large language models.However, due to the lack of unified token...
Research Assistant
AI chat, annotations, notes & similar papers
No comments yet
Be the first to share your thoughts!