falcon-2-11b llm
Falcon2-11B Technical Report
Malartic, Quentin, Chowdhury, Nilabhra Roy, Cojocaru, Ruxandra, Farooq, Mugariya, Campesan, Giulia, Djilali, Yasser Abdelaziz Dahou, Narayan, Sanath, Singh, Ankit, Velikanov, Maksim, Boussaha, Basma El Amel, Al-Yafeai, Mohammed, Alobeidli, Hamza, Qadi, Leen Al, Seddik, Mohamed El Amine, Fedyanin, Kirill, Alami, Reda, Hacid, Hakim
The first generation of Falcon models, featuring Falcon-7B, Falcon-40B, and Falcon-180B (Almazrouei et al., 2023), made a significant contribution to the open-source community, promoting the release of advanced LLMs with permissive licenses. In this report, we introduce a second generation of models, Falcon2, focused on increased usability and integrability, towards building a multi-modal ecosystem currently composed of a large language model with 11B parameters and a corresponding vision language model. Historically, large language models first saw an important rise in performance with increased model size (Brown et al., 2020; Chowdhery et al., 2022). Updated scaling laws (Hoffmann et al., 2022) brought to light that this initial generation of large language models were most likely undertrained, highlighting the need for more training data to further increase the performance. This triggered another important paradigm shift, namely moving from large curated datasets (Gao et al., 2020; Chowdhery et al., 2022), to large-scale datasets harvesting mostly web data from the CommonCrawl project, such as RefinedWeb (Penedo et al., 2023) or RedPajama (Computer, 2023). Both these advances led to the release of large open-source models such as Llama-65B (Touvron et al., 2023a) and Falcon-180B (Almazrouei et al., 2023). More recently, the Llama2 models (Touvron et al., 2023b) showed the benefits of even more prolonged training, achieving state-of-the-art performance with smaller model sizes. This trend was followed in the past year, resulting in a number of small-sized yet highly performing models such as Qwen-7B (Bai et al., 2023), Mistral-7B (Jiang et al., 2023), Yi-6B and Yi-9B (AI et al., 2024), Gemma-7B (Team et al., 2024) and Llama3-8B (AI@Meta, 2024).