Typical human actions last several seconds and exhibit characteristic spatio-temporal structure. Recent methods attempt to capture this structure and learn action representations with convolutional neural networks. Such representations, however, are typically learned at the level of a few video fram...
Research Assistant
AI chat, annotations, notes & similar papers
No comments yet
Be the first to share your thoughts!