Government
A Jumping Lunar Robot Is About to Explore a Pitch-Black Moon Crater for the First Time
A new age of commercial moon exploration is upon us, and one of the most exciting missions yet is about to launch--one laden with rovers, a drill, and even a hopper spacecraft that will try to "jump" into a permanently dark lunar crater to search for ice. The IM-2 mission, from Texas-based company Intuitive Machines, is scheduled to launch on a SpaceX Falcon 9 rocket from Cape Canaveral in Florida on Wednesday, February 26. The lander, nicknamed Athena and about the size of a car, is partially funded by NASA, as the US space agency attempts to create a new lunar economy that can support upcoming planned human missions to the moon. "NASA and the space industry is creating a new business, getting science and payloads to the surface of the moon," says Laura Forczyk, founder of the Georgia-based space consultancy firm Astralytical. "And these uncrewed missions are preparing us to send humans."
Is Xi's sudden embrace of business for real? China is left guessing
When Xi Jinping, China's leader, made his entrance at a symposium with a group of top entrepreneurs this past week, he seemed to be in good spirits. China has had a few good weeks. Artificial intelligence models by the startup DeepSeek sent U.S. stocks tumbling and Western commentators screaming, "Sputnik moment." Then, an animated film based on Chinese mythology raked in nearly 2 billion. Xi signaled that he stood behind the private sector at the meeting on Monday, pushing the Hong Kong stock market to its highest point in three years.
On Quantile Regression Forests for Modelling Mixed-Frequency and Longitudinal Data
The aim of this thesis is to extend the applications of the Quantile Regression Forest (QRF) algorithm to handle mixed-frequency and longitudinal data. To this end, standard statistical approaches have been exploited to build two novel algorithms: the Mixed- Frequency Quantile Regression Forest (MIDAS-QRF) and the Finite Mixture Quantile Regression Forest (FM-QRF). The MIDAS-QRF combines the flexibility of QRF with the Mixed Data Sampling (MIDAS) approach, enabling non-parametric quantile estimation with variables observed at different frequencies. FM-QRF, on the other hand, extends random effects machine learning algorithms to a QR framework, allowing for conditional quantile estimation in a longitudinal data setting. The contributions of this dissertation lie both methodologically and empirically. Methodologically, the MIDAS-QRF and the FM-QRF represent two novel approaches for handling mixed-frequency and longitudinal data in QR machine learning framework. Empirically, the application of the proposed models in financial risk management and climate-change impact evaluation demonstrates their validity as accurate and flexible models to be applied in complex empirical settings.
Real-time Monitoring of Economic Shocks using Company Websites
Koenig, Michael, Rauch, Jakob, Woerter, Martin
Understanding the effects of economic shocks on firms is critical for analyzing economic growth and resilience. We introduce a Web-Based Affectedness Indicator (W AI), a general-purpose tool for real-time monitoring of economic disruptions across diverse contexts. By leveraging Large Language Model (LLM) assisted classification and information extraction on texts from over five million company websites, W AI quantifies the degree and nature of firms' responses to external shocks. Using the COVID-19 pandemic as a specific application, we show that W AI is highly correlated with pandemic containment measures and reliably predicts firm performance. Unlike traditional data sources, W AI provides timely firm-level information across industries and geographies worldwide that would otherwise be unavailable due to institutional and data availability constraints. This methodology offers significant potential for monitoring and mitigating the impact of technological, political, financial, health or environmental crises, and represents a transformative tool for adaptive policy-making and economic resilience. Economic shocks, whether driven by public health crises, technological disruptions, geopolitical conflicts, or climate events, pose significant challenges to businesses and policymakers alike. Timely and accurate monitoring of these shocks is critical for crafting effective responses and enhancing economic resilience. However, traditional methods for measuring the impacts of such disruptions - such as surveys and administrative data - are often limited by costs, time lags, and coverage. In this study, we introduce the Web-Based Affectedness Indicator (W AI), a scalable and cost-effective tool for real-time monitoring of economic disruptions at the firm level. By analyzing textual data from millions of company websites, W AI provides granular insights into how firms experience and respond to external shocks. This 1 methodology overcomes traditional limitations by leveraging ubiquitous online content and state-of-the-art natural language processing (NLP) models to generate a dynamic and comprehensive view of economic affectedness. W AI can provide information on a wide range of challenges, including supply chain disruptions, financial crises, and climate-related shocks.
LongSpec: Long-Context Speculative Decoding with Efficient Drafting and Verification
Yang, Penghui, Du, Cunxiao, Zhang, Fengzhuo, Wang, Haonan, Pang, Tianyu, Du, Chao, An, Bo
Speculative decoding has become a promising technique to mitigate the high inference latency of autoregressive decoding in Large Language Models (LLMs). Despite its promise, the effective application of speculative decoding in LLMs still confronts three key challenges: the increasing memory demands of the draft model, the distribution shift between the short-training corpora and long-context inference, and inefficiencies in attention implementation. In this work, we enhance the performance of speculative decoding in long-context settings by addressing these challenges. First, we propose a memory-efficient draft model with a constant-sized Key-Value (KV) cache. Second, we introduce novel position indices for short-training data, enabling seamless adaptation from short-context training to long-context inference. Finally, we present an innovative attention aggregation method that combines fast implementations for prefix computation with standard attention for tree mask handling, effectively resolving the latency and memory inefficiencies of tree decoding. Our approach achieves strong results on various long-context tasks, including repository-level code completion, long-context summarization, and o1-like long reasoning tasks, demonstrating significant improvements in latency reduction. The code is available at https://github.com/sail-sg/LongSpec.
Diffusion Models for Tabular Data: Challenges, Current Progress, and Future Directions
Li, Zhong, Huang, Qi, Yang, Lincen, Shi, Jiayang, Yang, Zhao, van Stein, Niki, Bรคck, Thomas, van Leeuwen, Matthijs
In recent years, generative models have achieved remarkable performance across diverse applications, including image generation, text synthesis, audio creation, video generation, and data augmentation. Diffusion models have emerged as superior alternatives to Generative Adversarial Networks (GANs) and Variational Autoencoders (VAEs) by addressing their limitations, such as training instability, mode collapse, and poor representation of multimodal distributions. This success has spurred widespread research interest. In the domain of tabular data, diffusion models have begun to showcase similar advantages over GANs and VAEs, achieving significant performance breakthroughs and demonstrating their potential for addressing unique challenges in tabular data modeling. However, while domains like images and time series have numerous surveys summarizing advancements in diffusion models, there remains a notable gap in the literature for tabular data. Despite the increasing interest in diffusion models for tabular data, there has been little effort to systematically review and summarize these developments. This lack of a dedicated survey limits a clear understanding of the challenges, progress, and future directions in this critical area. This survey addresses this gap by providing a comprehensive review of diffusion models for tabular data. Covering works from June 2015, when diffusion models emerged, to December 2024, we analyze nearly all relevant studies, with updates maintained in a \href{https://github.com/Diffusion-Model-Leiden/awesome-diffusion-models-for-tabular-data}{GitHub repository}. Assuming readers possess foundational knowledge of statistics and diffusion models, we employ mathematical formulations to deliver a rigorous and detailed review, aiming to promote developments in this emerging and exciting area.
Genetics-Driven Personalized Disease Progression Model
Yang, Haoyu, Dey, Sanjoy, Meyer, Pablo
Modeling disease progression through multiple stages is critical for clinical decision-making for chronic diseases, e.g., cancer, diabetes, chronic kidney diseases, and so on. Existing approaches often model the disease progression as a uniform trajectory pattern at the population level. However, chronic diseases are highly heterogeneous and often have multiple progression patterns depending on a patient's individual genetics and environmental effects due to lifestyles. We propose a personalized disease progression model to jointly learn the heterogeneous progression patterns and groups of genetic profiles. In particular, an end-to-end pipeline is designed to simultaneously infer the characteristics of patients from genetic markers using a variational autoencoder and how it drives the disease progressions using an RNN-based state-space model based on clinical observations. Our proposed model shows improvement on real-world and synthetic clinical data.
Do Emotions Really Affect Argument Convincingness? A Dynamic Approach with LLM-based Manipulation Checks
Emotions have been shown to play a role in argument convincingness, yet this aspect is underexplored in the natural language processing (NLP) community. Unlike prior studies that use static analyses, focus on a single text domain or language, or treat emotion as just one of many factors, we introduce a dynamic framework inspired by manipulation checks commonly used in psychology and social science; leveraging LLM-based manipulation checks, this framework examines the extent to which perceived emotional intensity influences perceived convincingness. Through human evaluation of arguments across different languages, text domains, and topics, we find that in over half of cases, judgments of convincingness remain unchanged despite variations in perceived emotional intensity; when emotions do have an impact, they more often enhance rather than weaken convincingness. We further analyze how 11 LLMs behave in the same scenario, finding that while LLMs generally mirror human patterns, they struggle to capture nuanced emotional effects in individual judgments.
MAFE: Multi-Agent Fair Environments for Decision-Making Systems
Lazri, Zachary McBride, Nakra, Anirudh, Brugere, Ivan, Dervovic, Danial, Polychroniadou, Antigoni, Huang, Furong, Dachman-Soled, Dana, Wu, Min
Fairness constraints applied to machine learning (ML) models in static contexts have been shown to potentially produce adverse outcomes among demographic groups over time. To address this issue, emerging research focuses on creating fair solutions that persist over time. While many approaches treat this as a single-agent decision-making problem, real-world systems often consist of multiple interacting entities that influence outcomes. Explicitly modeling these entities as agents enables more flexible analysis of their interventions and the effects they have on a system's underlying dynamics. A significant challenge in conducting research on multi-agent systems is the lack of realistic environments that leverage the limited real-world data available for analysis. To address this gap, we introduce the concept of a Multi-Agent Fair Environment (MAFE) and present and analyze three MAFEs that model distinct social systems. Experimental results demonstrate the utility of our MAFEs as testbeds for developing multi-agent fair algorithms.
ARACNE: An LLM-Based Autonomous Shell Pentesting Agent
Nieponice, Tomas, Valeros, Veronica, Garcia, Sebastian
The complete automation of cyber-attacks is an area of growing interest since the surge of Large Language Models (LLMs) in recent years. Although the application of LLM in all areas of cybersecurity has flourished, the creation of attacking LLM agents that can act independently is among the most popular options [1]. Attacking LLM agents can perform automatic security testing of applications, lowering the cost for organizations to find vulnerabilities and misconfiguration problems and identify other security issues [2]. Existing automated attacking agents, such as PenHeal [2], AutoAttacker [3], and HackSynth [4] show promising results but with clear limitations. Agents are unable to work so far without occasional mistakes and hallucinations.