Microsoft trains world's largest Transformer language model

#artificialintelligence 

Microsoft AI & Research today shared what it calls the largest Transformer-based language generation model ever and open-sourced a deep learning library named DeepSpeed to make distributed training of large models easier. At 17 billion parameters, Turing NLG is twice the size of Nvidia's Megatron, now the second biggest Transformer model, and includes 10 times as many parameters as OpenAI's GPT-2. Turing NLG achieves state-of-the-art results on a range of NLP tasks. Like Google's Meena and initially with GPT-2, at first Turing NLG may only be shared in private demos. Language generation models with the Transformer architecture predict the word that comes next.

Duplicate Docs Excel Report

Title
None found

Similar Docs  Excel Report  more

TitleSimilaritySource
None found