Government
The Effect of Data Partitioning Strategy on Model Generalizability: A Case Study of Morphological Segmentation
Recent work to enhance data partitioning strategies for more realistic model evaluation face challenges in providing a clear optimal choice. This study addresses these challenges, focusing on morphological segmentation and synthesizing limitations related to language diversity, adoption of multiple datasets and splits, and detailed model comparisons. Our study leverages data from 19 languages, including ten indigenous or endangered languages across 10 language families with diverse morphological systems (polysynthetic, fusional, and agglutinative) and different degrees of data availability. We conduct large-scale experimentation with varying sized combinations of training and evaluation sets as well as new test data. Our results show that, when faced with new test data: (1) models trained from random splits are able to achieve higher numerical scores; (2) model rankings derived from random splits tend to generalize more consistently.
From Bytes to Borsch: Fine-Tuning Gemma and Mistral for the Ukrainian Language Representation
Kiulian, Artur, Polishko, Anton, Khandoga, Mykola, Chubych, Oryna, Connor, Jack, Ravishankar, Raghav, Shirawalmath, Adarsh
In the rapidly advancing field of AI and NLP, generative large language models (LLMs) stand at the forefront of innovation, showcasing unparalleled abilities in text understanding and generation. However, the limited representation of low-resource languages like Ukrainian poses a notable challenge, restricting the reach and relevance of this technology. Our paper addresses this by fine-tuning the open-source Gemma and Mistral LLMs with Ukrainian datasets, aiming to improve their linguistic proficiency and benchmarking them against other existing models capable of processing Ukrainian language. This endeavor not only aims to mitigate language bias in technology but also promotes inclusivity in the digital realm. Our transparent and reproducible approach encourages further NLP research and development. Additionally, we present the Ukrainian Knowledge and Instruction Dataset (UKID) to aid future efforts in language model fine-tuning. Our research not only advances the field of NLP but also highlights the importance of linguistic diversity in AI, which is crucial for cultural preservation, education, and expanding AI's global utility. Ultimately, we advocate for a future where technology is inclusive, enabling AI to communicate effectively across all languages, especially those currently underrepresented.
JaFIn: Japanese Financial Instruction Dataset
Tanabe, Kota, Suzuki, Masahiro, Sakaji, Hiroki, Noda, Itsuki
We construct an instruction dataset for the large language model (LLM) in the Japanese finance domain. Domain adaptation of language models, including LLMs, is receiving more attention as language models become more popular. This study demonstrates the effectiveness of domain adaptation through instruction tuning. To achieve this, we propose an instruction tuning data in Japanese called JaFIn, the Japanese Financial Instruction Dataset. JaFIn is manually constructed based on multiple data sources, including Japanese government websites, which provide extensive financial knowledge. We then utilize JaFIn to apply instruction tuning for several LLMs, demonstrating that our models specialized in finance have better domain adaptability than the original models. The financial-specialized LLMs created were evaluated using a quantitative Japanese financial benchmark and qualitative response comparisons, showing improved performance over the originals.
Can AI Understand Our Universe? Test of Fine-Tuning GPT by Astrophysical Data
Wang, Yu, Zhang, Shu-Rui, Momtaz, Aidin, Moradi, Rahim, Rastegarnia, Fatemeh, Sahakyan, Narek, Shakeri, Soroush, Li, Liang
ChatGPT has been the most talked-about concept in recent months, captivating both professionals and the general public alike, and has sparked discussions about the changes that artificial intelligence (AI) will bring to the world. As physicists and astrophysicists, we are curious about if scientific data can be correctly analyzed by large language models (LLMs) and yield accurate physics. In this article, we fine-tune the generative pre-trained transformer (GPT) model by the astronomical data from the observations of galaxies, quasars, stars, gamma-ray bursts (GRBs), and the simulations of black holes (BHs), the fine-tuned model demonstrates its capability to classify astrophysical phenomena, distinguish between two types of GRBs, deduce the redshift of quasars, and estimate BH parameters. We regard this as a successful test, marking the LLM's proven efficacy in scientific research. With the ever-growing volume of multidisciplinary data and the advancement of AI technology, we look forward to the emergence of a more fundamental and comprehensive understanding of our universe. This article also shares some interesting thoughts on data collection and AI design. Using the approach of understanding the universe - looking outward at data and inward for fundamental building blocks - as a guideline, we propose a method of series expansion for AI, suggesting ways to train and control AI that is smarter than humans.
Iran warns US to 'stay away' as America shoots down drone launched at Israel
Iraqi media broadcast video that reportedly shows Iranian missiles passing through the country. Iran's mission to the United Nations argued that the country's missiles and drones fired toward Israel were justified, warning the U.S. to "stay away." "It is a conflict between Iran and the rogue Israeli regime, from which the U.S. MUST STAY AWAY!," Iran's mission to the United Nations said in a statement. Iran's warning came as U.S. officials confirmed to Fox News that the U.S. military is continuing to shoot down Iranian drones that are headed toward Israel. "U.S. forces in the region continue to shoot down Iranian-launched drones targeting Israel. Our forces remain postured to provide additional defensive support and to protect U.S. forces operating in the region," a U.S. military official said.
Iran's attack on Israel: What drones and missiles will Tehran use in its strikes?
Rear Admiral Daniel Hagari of the Israeli Defense Forces updates the public Saturday, April 13 on Iranian UAVs being launched toward the Jewish state. Iran launched dozens of unmanned aerial vehicles (UAV), also known as drones, at Israel on Saturday evening local time after a week of threatening retaliation for an attack on a consulate in Damascus. "Iran has begun an airborne attack against Israel," White House National Security Council spokesperson Adrienne Watson said in a statement Saturday. "President Biden is being regularly updated on the situation by his national security team and will meet with them this afternoon at the White House." Iran's state-run news agency Mehr News reported that the Islamic Revolutionary Guard Corps (IRGC) had announced the start of the "anti-Zionist operation," which would hit targets in the Palestinian territories, with further details of the operation announced soon.
Lawmakers send message to White House on impending Iran drone attack to Israel: 'Stand firm'
Rep. Carlos Gimenez, R-Fla., tells'Fox News Live' that China, Russia, North Korea and Iran want to establish a'new world order.' Lawmakers reacted after Iran launched drones from its own territory toward Israel late Saturday, calling for the White House to "stand firm" and "stop coddling Iran." Speaker Mike Johnson pledged America's "full resolve" to stand with Israel. "As Israel faces this vicious attack from Iran, America must show our full resolve to stand with our critical ally," Johnson said in a statement. "The world must be assured: Israel is not alone."
White House says US support for Israel is 'ironclad,' will 'support their defense' amid Iran attack
The White House vowed Saturday that the United States' support for Israel's security is "ironclad," pledging to stand with the Jewish state and "support their defense" after Iran launched an aerial drone attack towards the country Saturday afternoon. Iran launched drones from its own territory toward Israel late Saturday, days after its Supreme Leader warned it would hit back in response to an airstrike on the Iranian consulate in Syria that left several generals dead. "Iran has begun an airborne attack against Israel," White House National Security Council spokesperson Adrienne Watson said in a statement Saturday. "President Biden is being regularly updated on the situation by his national security team and will meet with them this afternoon at the White House." The White House said the president's team "is in constant communication with Israeli officials as well as other partners and allies."
Roku Breach Hits 567,000 Users
After months of delays, the US House of Representatives voted on Friday to extend a controversial warrantless wiretap program for two years. Known as Section 702, the program authorizes the US government to collect the communications of foreigners overseas. But this collection also includes reams of communications from US citizens, which are stored for years and can later be warrantlessly accessed by the FBI, which has heavily abused the program. An amendment that would require investigators to obtain such a warrant failed to pass. A group of US lawmakers on Sunday unveiled a proposal that they hope will become the country's first nationwide privacy law.
Assessing Climate Transition Risks in the Colombian Processed Food Sector: A Fuzzy Logic and Multicriteria Decision-Making Approach
Pérez-Pérez, Juan F., Gómez, Pablo Isaza, Bonet, Isis, Sánchez-Pinzón, María Solange, Caraffini, Fabio, Lochmuller, Christian
Climate risk assessment is becoming increasingly important. For organisations, identifying and assessing climate-related risks is challenging, as they can come from multiple sources. This study identifies and assesses the main climate transition risks in the colombian processed food sector. As transition risks are vague, our approach uses Fuzzy Logic and compares it to various multi-criteria decision-making methods to classify the different climate transition risks an organisation may be exposed to. This approach allows us to use linguistic expressions for risk analysis and to better describe risks and their consequences. The results show that the risks ranked as the most critical for this organisation in their order were price volatility and raw materials availability, the change to less carbon-intensive production or consumption patterns, the increase in carbon taxes and technological change, and the associated development or implementation costs. These risks show a critical risk level, which implies that they are the most significant risks for the organisation in the case study. These results highlight the importance of investments needed to meet regulatory requirements, which are the main drivers for organisations at the financial level.