Microsoft open sources breakthrough optimizations for transformer inference on GPU and CPU

#artificialintelligence 

One of the most popular deep learning models used for natural language processing is BERT (Bidirectional Encoder Representations from Transformers). Due to the significant computation required, inferencing BERT at high scale can be extremely costly and may not even be possible with strict latency constraints. Recently, we shared how Bing has improved BERT inference on GPU for its real-time service needs, serving more than one million BERT inferences per second within Bing's latency limits. We are excited to announce that Microsoft has open sourced enhanced versions of these optimizations into the ONNX Runtime and extended them to work on both GPU and CPU. With ONNX Runtime, AI developers can now easily productionize large transformer models with high performance across both CPU and GPU hardware, using the same technology Microsoft uses to serve their customers.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found