NobleBlocks
Public

Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding

Published in arXiv (Cornell University) • Oct 1, 2015
Authors:
Song Han
,
Huizi Mao
,
William J. Dally

Abstract

Neural networks are both computationally intensive and memory intensive, making them difficult to deploy on embedded systems with limited hardware resources. To address this limitation, we introduce "deep compression", a three stage pipeline: pruning, trained quantization and Huffman coding, that wo...

Finding related papers...

Discussions

(0)

No comments yet

Be the first to share your thoughts!