Neural network training relies on our ability to find good minimizers of highly non-convex loss functions. It is well known that certain network architecture designs (e.g., skip connections) produce loss functions that train easier, and well-chosen training parameters (batch size, learning rate, opt...
Research Assistant
AI chat, annotations, notes & similar papers
No comments yet
Be the first to share your thoughts!