Efficient Fine-Grained GPU Performance Modeling for Distributed Deep Learning of LLM

Open in new window