Jais and Jais-chat: Arabic-Centric Foundation and Instruction-Tuned Open Generative Large Language Models
Sengupta, Neha, Sahu, Sunil Kumar, Jia, Bokang, Katipomu, Satheesh, Li, Haonan, Koto, Fajri, Marshall, William, Gosal, Gurpreet, Liu, Cynthia, Chen, Zhiming, Afzal, Osama Mohammed, Kamboj, Samta, Pandit, Onkar, Pal, Rahul, Pradhan, Lalit, Mujahid, Zain Muhammad, Baali, Massa, Han, Xudong, Bsharat, Sondos Mahmoud, Aji, Alham Fikri, Shen, Zhiqiang, Liu, Zhengzhong, Vassilieva, Natalia, Hestness, Joel, Hock, Andy, Feldman, Andrew, Lee, Jonathan, Jackson, Andrew, Ren, Hector Xuguang, Nakov, Preslav, Baldwin, Timothy, Xing, Eric
–arXiv.org Artificial Intelligence
We introduce Jais and Jais-chat, new state-of-the-art Arabic-centric foundation and instruction-tuned open generative large language models (LLMs). The models are based on the GPT-3 decoder-only architecture and are pretrained on a mixture of Arabic and English texts, including source code in various programming languages. With 13 billion parameters, they demonstrate better knowledge and reasoning capabilities in Arabic than any existing open Arabic and multilingual models by a sizable margin, based on extensive evaluation. Moreover, the models are competitive in English compared to English-centric open models of similar size, despite being trained on much less English data. We provide a detailed description of the training, the tuning, the safety alignment, and the evaluation of the models. We release two open versions of the model -- the foundation Jais model, and an instruction-tuned Jais-chat variant -- with the aim of promoting research on Arabic LLMs. Available at https://huggingface.co/inception-mbzuai/jais-13b-chat
arXiv.org Artificial Intelligence
Sep-29-2023
- Country:
- Africa
- Ethiopia > Addis Ababa
- Addis Ababa (0.04)
- Middle East (0.04)
- Rwanda > Kigali
- Kigali (0.04)
- Ethiopia > Addis Ababa
- Asia
- China > Beijing
- Beijing (0.04)
- India (0.04)
- Japan > Honshū
- Chūbu > Toyama Prefecture > Toyama (0.04)
- Middle East
- Qatar (0.04)
- Bahrain (0.04)
- Oman (0.04)
- UAE
- Abu Dhabi Emirate > Abu Dhabi (0.04)
- Dubai Emirate > Dubai (0.04)
- Ras Al Khaimah Emirate > Ras Al Khaimah (0.04)
- Republic of Türkiye > Istanbul Province
- Istanbul (0.04)
- Iraq (0.04)
- Yemen (0.04)
- Iran (0.04)
- Kuwait (0.14)
- Saudi Arabia > Arabian Gulf (0.04)
- Jordan (0.04)
- China > Beijing
- Europe
- Belgium > Brussels-Capital Region
- Brussels (0.04)
- Ireland > Leinster
- County Dublin > Dublin (0.04)
- Middle East > Republic of Türkiye
- Istanbul Province > Istanbul (0.04)
- Ukraine > Kyiv Oblast
- Kyiv (0.04)
- Norway > Eastern Norway
- Oslo (0.04)
- Spain > Catalonia
- Barcelona Province > Barcelona (0.04)
- Slovenia (0.04)
- United Kingdom > England
- Greater Manchester > Manchester (0.04)
- Denmark > Capital Region
- Copenhagen (0.04)
- France > Provence-Alpes-Côte d'Azur
- Bouches-du-Rhône > Marseille (0.04)
- Germany
- Italy > Tuscany
- Florence (0.04)
- Pisa Province > Pisa (0.04)
- Hungary > Budapest
- Budapest (0.04)
- Belgium > Brussels-Capital Region
- Indian Ocean > Arabian Gulf (0.04)
- North America
- Canada
- British Columbia > Metro Vancouver Regional District
- Vancouver (0.04)
- Ontario > Toronto (0.04)
- British Columbia > Metro Vancouver Regional District
- Dominican Republic (0.04)
- United States
- California > Los Angeles County
- Long Beach (0.04)
- Louisiana > Orleans Parish
- New Orleans (0.04)
- Minnesota > Hennepin County
- Minneapolis (0.14)
- New York > New York County
- New York City (0.04)
- Pennsylvania > Philadelphia County
- Philadelphia (0.04)
- California > Los Angeles County
- Canada
- South America > Chile
- Africa
- Genre:
- Research Report (1.00)
- Industry:
- Banking & Finance (0.92)
- Education
- Educational Setting > Online (0.45)
- Health & Safety > School Nutrition (0.45)
- Energy (1.00)
- Government > Regional Government
- Health & Medicine
- Consumer Health (1.00)
- Therapeutic Area (1.00)
- Information Technology (0.92)
- Law (1.00)
- Leisure & Entertainment > Sports (0.67)
- Technology: