Yinmin Zhong
13 Papers
Yinmin Zhong is an academic researcher. The author has contributed to research in topics: Computer science & Training (meteorology). The author has an hindex of 1, co-authored 3 publications.
Chat about Author
Papers
AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving
Zhuohan Li,Lianmin Zheng,Yinmin Zhong,Vincent Liu,Ying Sheng,Xin Jin,Yanping Huang,Zhi Chen,Hao Zhang,O. Gonzalez,Ionut Stoica +10 more
TL;DR: AlpaServe as discussed by the authors uses model parallelism for the statistical multiplexing of multiple devices when serving multiple models, even when a single model can fit into a single device.
64
MegaScale: Scaling Large Language Model Training to More Than 10, 000 GPUs
Ziheng Jiang,Haibin Lin,Yinmin Zhong,Qi Huang,Yangrui Chen,Zhi Zhang,Yanghua Peng,Xiang Li,Cong Xie,Shibiao Nong,Yulu Jia,Sun He,Hongmin Chen,Zhihao Bai,Qi Hou,Shipeng Yan,Ding Zhou,Yiyao Sheng,Zhuo Jiang,Haohan Xu,Haoran Wei,Zhang Zhang,Pengfei Nie,Leqi Zou,Sida Zhao,Liang Xiang,Zherui Liu,Zhe Li,X. Jia,Jia-jun Ye,Xin Jin,Xin Liu +31 more
- 23 Feb 2024
TL;DR: Researchers present MegaScale, a production system for training large language models on over 10,000 GPUs, achieving 55.2% Model FLOPs Utilization and improving efficiency by 1.34x compared to Megatron-LM, while addressing stability and fault tolerance challenges.
57
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
Yinmin Zhong,Shengyu Liu,Junda Chen,Jianbo Hu,Yibo Zhu,Xuanzhe Liu,Xin Jin,Hao Zhang +7 more
TL;DR: DistServe significantly improves LLM serving performance in terms of the maximum rate that can be served within both TTFT and TPOT constraints on each GPU, and assigns prefill and decoding computation to different GPUs, hence eliminating prefill-decoding interferences.
ElasticFlow: An Elastic Serverless Training Platform for Distributed Deep Learning
Diandian Gu,Yihao Zhao,Yinmin Zhong,Yifan Xiong,Zhenhua Han,Peng Cheng,Fan Yang,Gang Huang,Xin Jin,Xuanzhe Liu +9 more
- 27 Jan 2023
TL;DR: ElasticFlow as discussed by the authors is an elastic serverless training platform for distributed deep learning, which provides performance guarantees in terms of meeting deadlines while alleviating tedious, low-level, and manual resource management for deep learning developers.
15
LoongServe: Efficiently Serving Long-context Large Language Models with Elastic Sequence Parallelism
Bingya Wu,Shengyu Liu,Yinmin Zhong,Peng Sun,Xuanzhe Liu,Xin Jin +5 more
TL;DR: LoongServe efficiently serves long-context large language models with elastic sequence parallelism, improving throughput and reducing resource usage.
13