Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
Published in arXiv (Cornell University) • Oct 1, 2015
Authors:,,
Song Han
Huizi Mao
William J. Dally
Abstract
Neural networks are both computationally intensive and memory intensive, making them difficult to deploy on embedded systems with limited hardware resources. To address this limitation, we introduce "deep compression", a three stage pipeline: pruning, trained quantization and Huffman coding, that wo...
Finding related papers...
Discussions
(0)No comments yet
Be the first to share your thoughts!