Model Quantization with Intel Deep Learning Boost

#artificialintelligence 

The second generation of Intel Xeon Scalable processors introduced a collection of features for deep learning, packaged together as Intel Deep Learning Boost. These features include Vector Neural Network Instructions (VNNI), which increases throughput for inference applications with support for INT8 convolutions by combining multiple machine instructions from previous generations into one machine instruction.