Characterization of Large Language Model Development in the Datacenter

Hu, Qinghao, Ye, Zhisheng, Wang, Zerui, Wang, Guoteng, Zhang, Meng, Chen, Qiaoling, Sun, Peng, Lin, Dahua, Wang, Xiaolin, Luo, Yingwei, Wen, Yonggang, Zhang, Tianwei

Apr-3-2024–arXiv.org Artificial Intelligence

Large Language Models (LLMs) have presented impressive performance across several transformative tasks. However, it is non-trivial to efficiently utilize large-scale cluster resources to develop LLMs, often riddled with numerous challenges such as frequent hardware failures, intricate parallelization strategies, and imbalanced resource utilization. In this paper, we present an in-depth characterization study of a six-month LLM development workload trace collected from our GPU datacenter Acme. Specifically, we investigate discrepancies between LLMs and prior task-specific Deep Learning (DL) workloads, explore resource utilization patterns, and identify the impact of various job failures. Our analysis summarizes hurdles we encountered and uncovers potential opportunities to optimize systems tailored for LLMs. Furthermore, we introduce our system efforts: (1) fault-tolerant pretraining, which enhances fault tolerance through LLM-involved failure diagnosis and automatic recovery. (2) decoupled scheduling for evaluation, which achieves timely performance feedback via trial decomposition and scheduling optimization.

design and implementation, workload, zhang, (15 more...)

arXiv.org Artificial Intelligence

Apr-3-2024

arXiv.org PDF

Add feedback

Country:
- North America > United States (0.14)
- Europe > Italy
  - Calabria > Catanzaro Province > Catanzaro (0.04)
- Asia
  - India > Karnataka
    - Bengaluru (0.04)
  - China > Shanghai
    - Shanghai (0.04)

Genre:
- Research Report (0.81)

Industry:
- Information Technology (1.00)
- Energy > Renewable (0.46)

Technology:
- Information Technology > Artificial Intelligence
  - Natural Language > Large Language Model (1.00)
  - Machine Learning > Neural Networks
    - Deep Learning (1.00)

Duplicate Docs Excel Report

Title
None found

Similar Docs Excel Report more

Title	Similarity	Source
None found