Deep learning model compression
This post covers model inference optimization or compression in breadth and hopefully depth as of March 2021. This includes engineering topics like model quantization and binarization, more research-oriented topics like knowledge distillation, as well as well-known-hacks. Each year, larger and larger models are able to find methods for extracting signal from the noise in machine learning. In particular, language models get larger every day. These models are computationally expensive (in both runtime and memory), which can be both costly when served out to customers or too slow or large to function in edge environments like a phone. Researchers and practitioners have come up with many methods for optimizing neural networks to run faster or with less memory usage.
Jun-1-2021, 19:41:15 GMT
- Country:
- Asia > Middle East > Republic of Türkiye > Batman Province > Batman (0.05)
- Industry:
- Education (0.32)
- Technology: