Journal Article10.1109/tpds.2023.3340518
Graft: Efficient Inference Serving for Hybrid Deep Learning with SLO Guarantees via DNN Re-alignment
Jing Wu,Lin Wang,Qirui Jin,Fangming Liu +3 more
14
TL;DR: This article presents Graft—an efficient inference serving system for hybrid deep learning with latency service-level objective (SLO) guarantees, and proposes efficient algorithms for merging, grouping, and re-aligning DNN fragments to maximize request batching opportunities, minimizing resource consumption while guaranteeing the inference latency SLO.
read more
Abstract: Deep neural networks (DNNs) have been widely adopted for various mobile inference tasks, yet their ever-increasing computational demands are hindering their deployment on resource-constrained mobile devices. Hybrid deep learning partitions a DNN into two parts and deploys them across the mobile device and a server, aiming to reduce inference latency or prolong battery life of mobile devices. However, such partitioning produces (non-uniform) DNN fragments which are hard to serve efficiently on the server.This paper presents Graft -- an efficient inference serving system for hybrid deep learning with latency service-level objective (SLO) guarantees. Our main insight is to mitigate the non-uniformity by a core concept called DNN re-alignment, allowing multiple heterogeneous DNN fragments to be restructured to share layers. To fully exploit the potential of DNN re-alignment, Graft employs fine-grained GPU resource sharing. Based on that, we propose efficient algorithms for merging, grouping, and re-aligning DNN fragments to maximize request batching opportunities, minimizing resource consumption while guaranteeing the inference latency SLO. We implement a Graft prototype and perform extensive experiments with five types of widely used DNNs and real-world network traces. Our results show that Graft improves resource efficiency by up to 70% compared with the state-of-the-art inference serving systems.
read more
Chat with Paper
AI Agents for this Paper
Find similar papers on Google Scholar, PubMed and Arxiv
Write a critical review of this paper
Analyze citations of this paper to find unaddressed research gaps
Citations
A Survey on Integrating Edge Computing With AI and Blockchain in Maritime Domain, Aerial Systems, IoT, and Industry 4.0
Amad Alanhdi,Laszlo Toka +1 more
TL;DR: This survey addresses extensive research on the advancements of two application domains namely maritime areas and aerial systems in terms of integration with EC architecture by discussing several experiments conducted in various fields to demonstrate the value of utilizing them in the edge computing architecture.
9
Multi-Agent Collaborative Optimization of UAV Trajectory and Latency-Aware DAG Task Offloading in UAV-Assisted MEC
Chunxiang Zheng,Kai Pan,Jiadong Dong,Lin Chen,Shunfeng Wu,Hanyun Luo,Xiaoling Zhang +6 more
5
Task Offloading With Service Migration for Satellite Edge Computing: A Deep Reinforcement Learning Approach
Haonan Wu,Xiumei Yang,Zhiyong Bu +2 more
TL;DR: This work investigates the task offloading problem with service migration for satellite edge computing (SEC) using inter-satellite cooperation and proposes a distributed scheme based on the Dueling-Double-Deep-Q-Learning (D3QN) algorithm, which can effectively reduce the service delay, and outperform the benchmark algorithms.
2
PrVFL: Pruning-Aware Verifiable Federated Learning for Heterogeneous Edge Computing
X. Wang,Haiyang Yu,Yuwen Chen,Richard Sinnott,Zhen Yang +4 more
TL;DR: This paper introduces PrVFL, a verifiable federated learning framework for heterogeneous edge computing that enables model pruning verification using zero-knowledge range proofs and delayed verification, allowing edge users to choose pruning ratios and ensuring parameter training opportunities.
1
GreenFlow: A Carbon-Efficient Scheduler for Deep Learning Workloads
Diandian Gu,Yihao Zhao,Peng Sun,Xin Jin,Xuanzhe Liu +4 more
TL;DR: GreenFlow, a GPU cluster scheduler, reduces average Job Completion Time (JCT) by up to 2.15× under a carbon emission budget, through dynamic GPU allocation, job configuration adjustment, and network packing, outperforming competitive baselines in data centers.
1
References
•Proceedings Article
Very Deep Convolutional Networks for Large-Scale Image Recognition
Karen Simonyan,Andrew Zisserman +1 more
- 04 Sep 2014
TL;DR: This work investigates the effect of the convolutional network depth on its accuracy in the large-scale image recognition setting using an architecture with very small convolution filters, which shows that a significant improvement on the prior-art configurations can be achieved by pushing the depth to 16-19 weight layers.
102.6K
ImageNet classification with deep convolutional neural networks
TL;DR: A large, deep convolutional neural network was trained to classify the 1.2 million high-resolution images in the ImageNet LSVRC-2010 contest into the 1000 different classes and employed a recently developed regularization method called "dropout" that proved to be very effective.
Going deeper with convolutions
Christian Szegedy,Wei Liu,Yangqing Jia,Pierre Sermanet,Scott Reed,Dragomir Anguelov,Dumitru Erhan,Vincent Vanhoucke,Andrew Rabinovich +8 more
- 07 Jun 2015
TL;DR: Inception as mentioned in this paper is a deep convolutional neural network architecture that achieves the new state of the art for classification and detection in the ImageNet Large-Scale Visual Recognition Challenge 2014 (ILSVRC14).
•Posted Content
Rethinking Atrous Convolution for Semantic Image Segmentation
TL;DR: The proposed `DeepLabv3' system significantly improves over the previous DeepLab versions without DenseCRF post-processing and attains comparable performance with other state-of-art models on the PASCAL VOC 2012 semantic image segmentation benchmark.
9.9K
The Emergence of Edge Computing
TL;DR: A five-video playlist demonstrating proof-of-concept implementations for three tasks: assembling 2D Lego models, freehand sketching, and playing Ping-Pong is demonstrated.
2.2K