Microsoft Open-Sources ONNX Acceleration for BERT AI Model
With the optimizations, the model's inference latency on the SQUAD benchmark sped up 17x. Senior program manager Emma Ning gave an overview of the results in a blog post. In collaboration with engineers from Bing, the Azure researchers developed a condensed BERT model for understanding web-search queries. To improve the model's response time, the team re-implemented the model in C . Microsoft is now open-sourcing those optimizations by contributing them to ONNX Runtime, an open-source library for accelerating neural-network inference operations.
Jan-29-2020, 05:51:01 GMT
- Technology: