While depth tends to improve network performances, it also makes gradient-based training more difficult since deeper networks tend to be more non-linear. The recently proposed knowledge distillation approach is aimed at obtaining small and fast-to-execute models, and it has shown that a student netw...
Research Assistant
AI chat, annotations, notes & similar papers
No comments yet
Be the first to share your thoughts!