Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Vicent Sanz Marco

Optimizing Deep Learning Inference on Embedded Systems Through Adaptive Model Selection

Nov 09, 2019

Vicent Sanz Marco, Ben Taylor, Zheng Wang, Yehia Elkhatib

Figure 1 for Optimizing Deep Learning Inference on Embedded Systems Through Adaptive Model Selection

Figure 2 for Optimizing Deep Learning Inference on Embedded Systems Through Adaptive Model Selection

Figure 3 for Optimizing Deep Learning Inference on Embedded Systems Through Adaptive Model Selection

Figure 4 for Optimizing Deep Learning Inference on Embedded Systems Through Adaptive Model Selection

Abstract:Deep neural networks ( DNNs ) are becoming a key enabling technology for many application domains. However, on-device inference on battery-powered, resource-constrained embedding systems is often infeasible due to prohibitively long inferencing time and resource requirements of many DNNs. Offloading computation into the cloud is often unacceptable due to privacy concerns, high latency, or the lack of connectivity. While compression algorithms often succeed in reducing inferencing times, they come at the cost of reduced accuracy. This paper presents a new, alternative approach to enable efficient execution of DNNs on embedded devices. Our approach dynamically determines which DNN to use for a given input, by considering the desired accuracy and inference time. It employs machine learning to develop a low-cost predictive model to quickly select a pre-trained DNN to use for a given input and the optimization constraint. We achieve this by first off-line training a predictive model, and then using the learned model to select a DNN model to use for new, unseen inputs. We apply our approach to two representative DNN domains: image classification and machine translation. We evaluate our approach on a Jetson TX2 embedded deep learning platform and consider a range of influential DNN models including convolutional and recurrent neural networks. For image classification, we achieve a 1.8x reduction in inference time with a 7.52% improvement in accuracy, over the most-capable single DNN model. For machine translation, we achieve a 1.34x reduction in inference time over the most-capable single model, with little impact on the quality of translation.

* Accepted to be published at ACM TECS. arXiv admin note: substantial text overlap with arXiv:1805.04252

Via

Access Paper or Ask Questions

Adaptive Selection of Deep Learning Models on Embedded Systems

May 11, 2018

Ben Taylor, Vicent Sanz Marco, Willy Wolff, Yehia Elkhatib, Zheng Wang

Figure 1 for Adaptive Selection of Deep Learning Models on Embedded Systems

Figure 2 for Adaptive Selection of Deep Learning Models on Embedded Systems

Figure 3 for Adaptive Selection of Deep Learning Models on Embedded Systems

Figure 4 for Adaptive Selection of Deep Learning Models on Embedded Systems

Abstract:The recent ground-breaking advances in deep learning networks ( DNNs ) make them attractive for embedded systems. However, it can take a long time for DNNs to make an inference on resource-limited embedded devices. Offloading the computation into the cloud is often infeasible due to privacy concerns, high latency, or the lack of connectivity. As such, there is a critical need to find a way to effectively execute the DNN models locally on the devices. This paper presents an adaptive scheme to determine which DNN model to use for a given input, by considering the desired accuracy and inference time. Our approach employs machine learning to develop a predictive model to quickly select a pre-trained DNN to use for a given input and the optimization constraint. We achieve this by first training off-line a predictive model, and then use the learnt model to select a DNN model to use for new, unseen inputs. We apply our approach to the image classification task and evaluate it on a Jetson TX2 embedded deep learning platform using the ImageNet ILSVRC 2012 validation dataset. We consider a range of influential DNN models. Experimental results show that our approach achieves a 7.52% improvement in inference accuracy, and a 1.8x reduction in inference time over the most-capable single DNN model.

* Accepted to be published at LCTES 2018

Via

Access Paper or Ask Questions