Breadth-First Pipeline Parallelism
–arXiv.org Artificial Intelligence
We introduce Breadth-First Pipeline Parallelism, a novel training schedule which optimizes the combination of pipeline and data parallelism. Breadth-First Pipeline Parallelism lowers training time, cost and memory usage by combining a high GPU utilization with a small batch size per GPU, and by making use of fully sharded data parallelism. Experimentally, we observed an increase of up to 43% in training throughput for a 52 billion-parameter model using a small batch size per GPU compared to Megatron-LM, which would reduce the training time and cost by the same amount on a large GPU cluster.
arXiv.org Artificial Intelligence
Jul-6-2023
- Country:
- North America
- United States > Florida
- Miami-Dade County > Miami Beach (0.04)
- Canada > Quebec
- Montreal (0.04)
- United States > Florida
- Europe > Italy
- Calabria > Catanzaro Province > Catanzaro (0.04)
- North America
- Genre:
- Research Report (1.00)
- Technology: