Government
Trump addresses Iran attack on Israel at Pennsylvania rally: 'Would not have happened if we were in office'
Trump's third visit to the battleground state comes just one day after his election integrity joint press conference with House Speaker Mike Johnson. Following Iran's attack on Israel, former President Trump shared a heated message to those in attendance at his rally in Pennsylvania on Saturday, vowing it should not have happened. "Before going any further, I want to say God bless the people of Israel. That's because we show great weakness," Trump said to open his speech. "The weakness that we've shown is unbelievable, and it would not have happened if we were in office. They know that, and everybody knows that."
How Israel Is Defending Against Iran's Drone Attack
On Saturday, Iran launched more than 200 drones and cruise missiles at Israel. As the drones made their way across the Middle East en route to their target, Israel has invoked a number of defense systems to impede their progress. None will be more important than the Iron Dome. The Iron Dome, operational for well over a decade, comprises at least 10 missile-defense batteries strategically distributed around the country. When radar detects incoming objects, it sends that information back to a command-and-control center, which will track the threat to assess whether it's a false alarm, and where it might hit if it's not.
LuminLab: An AI-Powered Building Retrofit and Energy Modelling Platform
Credit, Kevin, Xiao, Qian, Lehane, Jack, Vazquez, Juan, Liu, Dan, De Figueiredo, Leo
This paper describes the technical and conceptual development of the LuminLab platform, an online tool that integrates a purpose-fit human-centric AI chatbot and predictive energy model into a streamlined front-end that can rapidly produce and discuss building retrofit plans in natural language. The platform provides users with the ability to engage with a range of possible retrofit pathways tailored to their individual budget and building needs on-demand. Given the complicated and costly nature of building retrofit projects, which rely on a variety of stakeholder groups with differing goals and incentives, we feel that AI-powered tools such as this have the potential to pragmatically de-silo knowledge, improve communication, and empower individual homeowners to undertake incremental retrofit projects that might not happen otherwise.
Reap the Wild Wind: Detecting Media Storms in Large-Scale News Corpora
Markus, Dror K., Levi, Effi, Sheafer, Tamir, Shenhav, Shaul R.
Media Storms, dramatic outbursts of attention to a story, are central components of media dynamics and the attention landscape. Despite their significance, there has been little systematic and empirical research on this concept due to issues of measurement and operationalization. We introduce an iterative human-in-the-loop method to identify media storms in a large-scale corpus of news articles. The text is first transformed into signals of dispersion based on several textual characteristics. In each iteration, we apply unsupervised anomaly detection to these signals; each anomaly is then validated by an expert to confirm the presence of a storm, and those results are then used to tune the anomaly detection in the next iteration. We demonstrate the applicability of this method in two scenarios: first, supplementing an initial list of media storms within a specific time frame; and second, detecting media storms in new time periods. We make available a media storm dataset compiled using both scenarios.
The Effect of Data Partitioning Strategy on Model Generalizability: A Case Study of Morphological Segmentation
Recent work to enhance data partitioning strategies for more realistic model evaluation face challenges in providing a clear optimal choice. This study addresses these challenges, focusing on morphological segmentation and synthesizing limitations related to language diversity, adoption of multiple datasets and splits, and detailed model comparisons. Our study leverages data from 19 languages, including ten indigenous or endangered languages across 10 language families with diverse morphological systems (polysynthetic, fusional, and agglutinative) and different degrees of data availability. We conduct large-scale experimentation with varying sized combinations of training and evaluation sets as well as new test data. Our results show that, when faced with new test data: (1) models trained from random splits are able to achieve higher numerical scores; (2) model rankings derived from random splits tend to generalize more consistently.
From Bytes to Borsch: Fine-Tuning Gemma and Mistral for the Ukrainian Language Representation
Kiulian, Artur, Polishko, Anton, Khandoga, Mykola, Chubych, Oryna, Connor, Jack, Ravishankar, Raghav, Shirawalmath, Adarsh
In the rapidly advancing field of AI and NLP, generative large language models (LLMs) stand at the forefront of innovation, showcasing unparalleled abilities in text understanding and generation. However, the limited representation of low-resource languages like Ukrainian poses a notable challenge, restricting the reach and relevance of this technology. Our paper addresses this by fine-tuning the open-source Gemma and Mistral LLMs with Ukrainian datasets, aiming to improve their linguistic proficiency and benchmarking them against other existing models capable of processing Ukrainian language. This endeavor not only aims to mitigate language bias in technology but also promotes inclusivity in the digital realm. Our transparent and reproducible approach encourages further NLP research and development. Additionally, we present the Ukrainian Knowledge and Instruction Dataset (UKID) to aid future efforts in language model fine-tuning. Our research not only advances the field of NLP but also highlights the importance of linguistic diversity in AI, which is crucial for cultural preservation, education, and expanding AI's global utility. Ultimately, we advocate for a future where technology is inclusive, enabling AI to communicate effectively across all languages, especially those currently underrepresented.
JaFIn: Japanese Financial Instruction Dataset
Tanabe, Kota, Suzuki, Masahiro, Sakaji, Hiroki, Noda, Itsuki
We construct an instruction dataset for the large language model (LLM) in the Japanese finance domain. Domain adaptation of language models, including LLMs, is receiving more attention as language models become more popular. This study demonstrates the effectiveness of domain adaptation through instruction tuning. To achieve this, we propose an instruction tuning data in Japanese called JaFIn, the Japanese Financial Instruction Dataset. JaFIn is manually constructed based on multiple data sources, including Japanese government websites, which provide extensive financial knowledge. We then utilize JaFIn to apply instruction tuning for several LLMs, demonstrating that our models specialized in finance have better domain adaptability than the original models. The financial-specialized LLMs created were evaluated using a quantitative Japanese financial benchmark and qualitative response comparisons, showing improved performance over the originals.
Can AI Understand Our Universe? Test of Fine-Tuning GPT by Astrophysical Data
Wang, Yu, Zhang, Shu-Rui, Momtaz, Aidin, Moradi, Rahim, Rastegarnia, Fatemeh, Sahakyan, Narek, Shakeri, Soroush, Li, Liang
ChatGPT has been the most talked-about concept in recent months, captivating both professionals and the general public alike, and has sparked discussions about the changes that artificial intelligence (AI) will bring to the world. As physicists and astrophysicists, we are curious about if scientific data can be correctly analyzed by large language models (LLMs) and yield accurate physics. In this article, we fine-tune the generative pre-trained transformer (GPT) model by the astronomical data from the observations of galaxies, quasars, stars, gamma-ray bursts (GRBs), and the simulations of black holes (BHs), the fine-tuned model demonstrates its capability to classify astrophysical phenomena, distinguish between two types of GRBs, deduce the redshift of quasars, and estimate BH parameters. We regard this as a successful test, marking the LLM's proven efficacy in scientific research. With the ever-growing volume of multidisciplinary data and the advancement of AI technology, we look forward to the emergence of a more fundamental and comprehensive understanding of our universe. This article also shares some interesting thoughts on data collection and AI design. Using the approach of understanding the universe - looking outward at data and inward for fundamental building blocks - as a guideline, we propose a method of series expansion for AI, suggesting ways to train and control AI that is smarter than humans.
Iran warns US to 'stay away' as America shoots down drone launched at Israel
Iraqi media broadcast video that reportedly shows Iranian missiles passing through the country. Iran's mission to the United Nations argued that the country's missiles and drones fired toward Israel were justified, warning the U.S. to "stay away." "It is a conflict between Iran and the rogue Israeli regime, from which the U.S. MUST STAY AWAY!," Iran's mission to the United Nations said in a statement. Iran's warning came as U.S. officials confirmed to Fox News that the U.S. military is continuing to shoot down Iranian drones that are headed toward Israel. "U.S. forces in the region continue to shoot down Iranian-launched drones targeting Israel. Our forces remain postured to provide additional defensive support and to protect U.S. forces operating in the region," a U.S. military official said.