Benchmarking Deep Learning Frameworks with FPGA-suitable Models on a Traffic Sign Dataset
Zhongyi Lin,Jeffrey M. Ota,John D. Owens,Pnar Muyan-Ozcelik +3 more
- 26 Jun 2018
- pp 1197-1203
TL;DR: It is discovered that Neon and MXNet deliver the best training speed and inference accuracy in general for all test cases, while Tensorflow is always among the frameworks with the highest inference accuracies.
read more
Abstract: We benchmark several widely used deep-learning frameworks for performing deep-learning-related automotive tasks (e.g., traffic sign recognition) that need to achieve realtime and high accuracy results with limited resources available on embedded platforms such as FPGAs. In our benchmarks, we use various input image sizes on models that are suitable for FPGA deployment, and investigate the training speed and inference accuracy of selected frameworks for these different sizes on a popular traffic sign recognition dataset. We report results by running the frameworks solely on the CPU as well as by turning on GPU acceleration. We also provide optimizations we apply to fine-tune the performance of the frameworks. We discover that Neon and MXNet deliver the best training speed and inference accuracy in general for all our test cases, while Tensorflow is always among the frameworks with the highest inference accuracies. We also observe that on the particular dataset we tested on (i.e., GTSRB), the image size of the region of interest does not necessarily affect the inference accuracy, and that using deep models, e.g., ResNet-32, which have longer training times, might not provide improvements to inference accuracy.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
An Efficient Task Assignment Framework to Accelerate DPU-Based Convolutional Neural Network Inference on FPGAs
TL;DR: A high performance task assignment framework built upon Xilinx hybrid CPU-FPGA MPSoC devices and integrated with optimizations to maximize performance on DPU-based CNN acceleration platform is presented.
Benchmarking Deep Learning Frameworks and Investigating FPGA Deployment for Traffic Sign Classification and Detection
Zhongyi Lin,Matthew Yih,Jeffrey M. Ota,John D. Owens,Pinar Muyan-Ozcelik +4 more
- 28 May 2019
TL;DR: This work benchmarks several widely-used deep learning frameworks and investigates the field programmable gate array (FPGA) deployment for performing traffic sign classification and detection, finding that Neon and MXNet deliver the best training speed and classification accuracy on the GPU in general for all test cases, while TensorFlow is always among the frameworks with the highest inference accuracies.
Sports image detection based on FPGA hardware system and particle swarm algorithm
Hu Jing,Xing Xiaoqiong +1 more
TL;DR: This assessment development image as a thing to ponder the usage of image acknowledgment development is set itself up as a demonstrating ground to test the suitability of the investigation framework that sees the affirmation of contenders, games affirmation, sports lead judgment, etc.
19
•Posted Content
Object Localization with a Weakly Supervised CapsNet
TL;DR: This work proposes a CapsNet architecture with object coordinate atoms and a modified routing-by-agreement algorithm with unevenly distributed initial routing probabilities, based on CapsNet but uses a routing algorithm to find the objects' approximate positions in the image coordinate system.
1
Convolutional neural network libraries benchmarking: a systematic mapping
Felipe de Almeida Florencio,Edward David Moreno +1 more
- 25 Nov 2020
TL;DR: A systematic literature mapping was conducted to analyze the scientific research in the field of convolutional neural networks (CNNs) and found 12 papers that evaluate the performance of Convolutional Neural Networks on different computer architectures as mentioned in this paper.
References
Deep Residual Learning for Image Recognition
Kaiming He,Xiangyu Zhang,Shaoqing Ren,Jian Sun +3 more
- 27 Jun 2016
TL;DR: In this article, the authors proposed a residual learning framework to ease the training of networks that are substantially deeper than those used previously, which won the 1st place on the ILSVRC 2015 classification task.
•Posted Content
Deep Residual Learning for Image Recognition
TL;DR: This work presents a residual learning framework to ease the training of networks that are substantially deeper than those used previously, and provides comprehensive empirical evidence showing that these residual networks are easier to optimize, and can gain accuracy from considerably increased depth.
117.9K
SSD: Single Shot MultiBox Detector
Wei Liu,Dragomir Anguelov,Dumitru Erhan,Christian Szegedy,Scott Reed,Cheng-Yang Fu,Alexander C. Berg +6 more
- 08 Oct 2016
TL;DR: The approach, named SSD, discretizes the output space of bounding boxes into a set of default boxes over different aspect ratios and scales per feature map location, which makes SSD easy to train and straightforward to integrate into systems that require a detection component.
•Posted Content
MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
Andrew Howard,Menglong Zhu,Bo Chen,Dmitry Kalenichenko,Weijun Wang,Tobias Weyand,M. Andreetto,Hartwig Adam +7 more
TL;DR: This work introduces two simple global hyper-parameters that efficiently trade off between latency and accuracy and demonstrates the effectiveness of MobileNets across a wide range of applications and use cases including object detection, finegrain classification, face attributes and large scale geo-localization.
18.5K
Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
Kaiming He,Xiangyu Zhang,Shaoqing Ren,Jian Sun +3 more
- 07 Dec 2015
TL;DR: In this paper, a Parametric Rectified Linear Unit (PReLU) was proposed to improve model fitting with nearly zero extra computational cost and little overfitting risk, which achieved a 4.94% top-5 test error on ImageNet 2012 classification dataset.