Cost-Driven Hardware-Software Co-Optimization of Machine Learning Pipelines

Sharma, Ravit, Romaszkan, Wojciech, Zhu, Feiqian, Gupta, Puneet, Mehta, Ankur

arXiv.org Artificial Intelligence 

The combination of Internet-of-Things (IoT) and Deep Learning (DL) trends has created an enormous demand for ultra-low footprint machine learning models, commonly referred to as TinyML [23]. On-device, or near-sensor, inference ensures privacy while avoiding high energy and latency cost of offloading computation to the cloud [8]. Enabling more complex algorithms on low-cost, microcontroller-based systems has the potential of making access to smart devices ubiquitous. Broad availability, low cost, and ease-of-use would, in turn, make it possible for people to experiment with an increasingly-broader range of applications that improve human-computer interaction, such as audio and visual wake words, context recognition, and user verification [23]. However, achieving it requires unprecedented efforts on co-optimization of algorithms and hardware to make large and computationally complex models usable on devices with very limited memory and processing power [57]. Multiple techniques have been proposed to address model compression: quantization [1, 30, 46, 57], which uses lower precision numbers for more efficient storage and computation, pruning [19, 31, 57] which removes inconsequential weights, compressed models [33, 36], and optimized software libraries [21, 43]. While all the above knobs are readily available to machine learning researchers, it is not obvious how they interact with hardware configurations, given the specific set of constraints, e.g., cost, latency, size, and user experience.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found