South America
Combining feature-based approaches with graph neural networks and symbolic regression for synergistic performance and interpretability
Gouvêa, Rogério Almeida, De Breuck, Pierre-Paul, Pretto, Tatiane, Rignanese, Gian-Marco, Santos, Marcos José Leite
To avoid the featuri zation bottleneck of traditional descriptors, we also leverage GNNs to generate fast, latent-space approximations of MatMiner (ℓ-MM) and Orbital Field Matrix (ℓ-OFM) features. Finally, we augment this feature set with new descriptors derived via symbolic regression. This multifac eted strategy aims to create a more robust, accurate, and versatile featurizer that capitalizes on the distinct strengths of each approach to be useful for a wider range of dataset sizes. To simplify the generation of all those features, a package was developed named MatterVial standing for MATerials fea T uR e E xtraction Via I nterpretable Artificial L earning, which, besides producing all latent-space features from the GNN models, aids i n obtaining the interpretable chemical descriptors that correlate to these high-level features. This is achieved through techniques such as SHapley Additive exPlanations (SHAP) analysi s in surrogate models and symbolic regression via Sure Independence Screening and Sparsifying Operator (SISSO) to obtain an approximate formula from the most important features. Our re sults demonstrate an overall improvement in all analyzed datasets compare d with the baseline MatMiner featurizer. In addition, it surpassed the performance of the individua l GNN models in several cases, indicating that the combination of traditional and l atent-space features leads to a more robust generalization.
AI firm plans to reconstruct lost footage from Orson Welles' masterpiece The Magnificent Ambersons
An AI company is to reconstruct the missing portions of Orson Welles' legendary mutilated masterwork The Magnificent Ambersons, it has been announced. According to the Hollywood Reporter, the Showrunner platform is planning to use its AI tools to assist in a recreation of the lost 43 minutes of Welles' 1942 film, removed and subsequently destroyed by Hollywood studio RKO. Edward Saatchi, CEO of interactive AI film-making studio Fable, which operates Showrunner, said in a statement to IndieWire: "We're starting with Orson Welles because he is the greatest storyteller of the last 200 years … So many people are rightly skeptical of AI's impact on cinema – but we hope that this gives people a sense of a positive contribution that AI can make for storytelling." Reports suggest that Showrunner is partnering with film-maker Brian Rose, who has been working since 2019 on an attempt to reconstruct the missing portions using animated sequences, as well as VFX expert Tom Clive. Welles started production in 1942 on The Magnificent Ambersons, an adaptation of Booth Tarkington's celebrated novel about a midwestern family in decline, as a follow-up to his Oscar-winning debut Citizen Kane.
Tech CEOs Praise Donald Trump at White House Dinner
The camera zooms too close to the president's face; the table at which the tech executives are seated seems far too long. Mark Zuckerberg is there, and Bill Gates and Tim Cook and Satya Nadella and Sam Altman and on and on, a baker's dozen or so of Silicon Valley's most powerful people--cutthroat competitors all--united here to pledge allegiance to Donald Trump. The introduction from Trump is characteristically both overgilded and confusing: "It's an honor to be here with this group of people. And then, about 90 seconds in, the pandering begins. This was Donald Trump's dinner with tech leaders at the State Dining Room in the White House on Thursday evening, broadcast in part for all to see on C-SPAN.
Multilinear and Linear Programs for Partially Identifiable Queries in Quasi-Markovian Structural Causal Models
Arroyo, João P., Rodrigues, João G., Lawand, Daniel, Mauá, Denis D., Lee, Junkyu, Marinescu, Radu, Gray, Alex, Laurentino, Eduardo R., Cozman, Fabio G.
We investigate partially identifiable queries in a class of causal models. We focus on acyclic Structural Causal Models that are quasi-Markovian (that is, each endogenous variable is connected with at most one exogenous confounder). We look into scenarios where endogenous variables are observed (and a distribution over them is known), while exogenous variables are not fully specified. This leads to a representation that is in essence a Bayesian network where the distribution of root variables is not uniquely determined. In such circumstances, it may not be possible to precisely compute a probability value of interest. We thus study the computation of tight probability bounds, a problem that has been solved by multilinear programming in general, and by linear programming when a single confounded component is intervened upon. We present a new algorithm to simplify the construction of such programs by exploiting input probabilities over endogenous variables. For scenarios with a single intervention, we apply column generation to compute a probability bound through a sequence of auxiliary linear integer programs, thus showing that a representation with polynomial cardinality for exogenous variables is possible. Experiments show column generation techniques to be superior to existing methods.
'Silent killer' parasitic disease spreading across multiple US states, experts warn
Fox News senior medical analyst Dr. Marc Siegel shares his perspective on whether the mosquito-borne virus in China will spread to the United States and how AI can be detrimental to children's and young adults' mental health on'Fox Report.' A little-known disease is spreading in the U.S., primarily in the state of California, health officials warn. In a new study published in the CDC journal Emerging Infectious Diseases, researchers state that human cases of Chagas disease have been confirmed in eight states, leading them to recommend that the disease is classified as "endemic." "Acknowledging the endemicity of Chagas disease in the United States is crucial for achieving global health goals," the authors wrote. The Centers for Disease Control and Prevention defines a disease as "endemic" when there is a "constant presence and/or usual prevalence" in a population within a specific geographic area -- in other words, the "baseline" level of disease within a community.
Securing AI Agents with Information-Flow Control
Costa, Manuel, Köpf, Boris, Kolluri, Aashish, Paverd, Andrew, Russinovich, Mark, Salem, Ahmed, Tople, Shruti, Wutschitz, Lukas, Zanella-Béguelin, Santiago
As AI agents become increasingly autonomous and capable, ensuring their security against vulnerabilities such as prompt injection becomes critical. This paper explores the use of information-flow control (IFC) to provide security guarantees for AI agents. We present a formal model to reason about the security and expressiveness of agent planners. Using this model, we characterize the class of properties enforceable by dynamic taint-tracking and construct a taxonomy of tasks to evaluate security and utility trade-offs of planner designs. Informed by this exploration, we present Fides, a planner that tracks confidentiality and integrity labels, deterministically enforces security policies, and introduces novel primitives for selectively hiding information. Its evaluation in AgentDojo demonstrates that this approach enables us to complete a broad range of tasks with security guarantees. A tutorial to walk readers through the the concepts introduced in the paper can be found at https://github.com/microsoft/fides
A Composite-Loss Graph Neural Network for the Multivariate Post-Processing of Ensemble Weather Forecasts
Ensemble forecasting systems have advanced meteorology by providing probabilistic estimates of future states, supporting applications from renewable energy production to transportation safety. Nonetheless, systematic biases often persist, making statistical post-processing essential. Traditional parametric post-processing techniques and machine learning-based methods can produce calibrated predictive distributions at specific locations and lead times, yet often struggle to capture dependencies across forecast dimensions. To address this, multivariate post-processing methods-such as ensemble copula coupling and the Schaake shuffle-are widely applied in a second step to restore realistic inter-variable or spatio-temporal dependencies. The aim of this study is the multivariate post-processing of ensemble forecasts using a graph neural network (dualGNN) trained with a composite loss function that combines the energy score (ES) and the variogram score (VS). The method is evaluated on two datasets: WRF-based solar irradiance forecasts over northern Chile and ECMWF visibility forecasts for Central Europe. The dualGNN consistently outperforms all empirical copula-based post-processed forecasts and shows significant improvements compared to graph neural networks trained solely on either the continuous ranked probability score (CRPS) or the ES, according to the evaluated multivariate verification metrics. Furthermore, for the WRF forecasts, the rank-order structure of the dualGNN forecasts captures valuable dependency information, enabling a more effective restoration of spatial relationships than either the raw numerical weather prediction ensemble or historical observational rank structures. By contrast, for the visibility forecasts, the GNNs trained on CRPS, ES, or the ES-VS combination outperform the calibrated reference.
Cost-Optimized Systems Engineering for IoT-Enabled Robot Nurse in Infectious Pandemic Management
Sifat, Md Mhamud Hussen, Maruf, Md, Rokunuzzaman, Md
The utilization of robotic technology has gained traction in healthcare facilities due to progress in the field that enables time and cost savings, minimizes waste, and improves patient care. Digital healthcare technologies that leverage automation, such as robotics and artificial intelligence, have the potential to enhance the sustainability and profitability of healthcare systems in the long run. However, the recent COVID-19 pandemic has amplified the need for cyber-physical robots to automate check-ups and medication administration. A robot nurse is controlled by the Internet of Things (IoT) and can serve as an automated medical assistant while also allowing supervisory control based on custom commands. This system helps reduce infection risk and improves outcomes in pandemic settings. This research presents a test case with a nurse robot that can assess a patient's health status and take action accordingly. We also evaluate the system's performance in medication administration, health-status monitoring, and life-cycle considerations.
Multi-level SSL Feature Gating for Audio Deepfake Detection
Tran, Hoan My, Lolive, Damien, Sini, Aghilas, Delhay, Arnaud, Marteau, Pierre-François, Guennec, David
Recent advancements in generative AI, particularly in speech synthesis, have enabled the generation of highly natural-sounding synthetic speech that closely mimics human voices. While these innovations hold promise for applications like assistive technologies, they also pose significant risks, including misuse for fraudulent activities, identity theft, and security threats. Current research on spoofing detection countermeasures remains limited by generalization to unseen deepfake attacks and languages. To address this, we propose a gating mechanism extracting relevant feature from the speech foundation XLS-R model as a front-end feature extractor. For downstream back-end classifier, we employ Multi-kernel gated Convolution (MultiConv) to capture both local and global speech artifacts. Additionally, we introduce Centered Kernel Alignment (CKA) as a similarity metric to enforce diversity in learned features across different MultiConv layers. By integrating CKA with our gating mechanism, we hypothesize that each component helps improving the learning of distinct synthetic speech patterns. Experimental results demonstrate that our approach achieves state-of-the-art performance on in-domain benchmarks while generalizing robustly to out-of-domain datasets, including multilingual speech samples. This underscores its potential as a versatile solution for detecting evolving speech deepfake threats.
SESGO: Spanish Evaluation of Stereotypical Generative Outputs
Robles, Melissa, Bernal, Catalina, Raigoso, Denniss, Rubio, Mateo Dulce
This paper addresses the critical gap in evaluating bias in multilingual Large Language Models (LLMs), with a specific focus on Spanish language within culturally-aware Latin American contexts. Despite widespread global deployment, current evaluations remain predominantly US-English-centric, leaving potential harms in other linguistic and cultural contexts largely underexamined. We introduce a novel, culturally-grounded framework for detecting social biases in instruction-tuned LLMs. Our approach adapts the underspecified question methodology from the BBQ dataset by incorporating culturally-specific expressions and sayings that encode regional stereotypes across four social categories: gender, race, socioeconomic class, and national origin. Using more than 4,000 prompts, we propose a new metric that combines accuracy with the direction of error to effectively balance model performance and bias alignment in both ambiguous and disambiguated contexts. To our knowledge, our work presents the first systematic evaluation examining how leading commercial LLMs respond to culturally specific bias in the Spanish language, revealing varying patterns of bias manifestation across state-of-the-art models. We also contribute evidence that bias mitigation techniques optimized for English do not effectively transfer to Spanish tasks, and that bias patterns remain largely consistent across different sampling temperatures. Our modular framework offers a natural extension to new stereotypes, bias categories, or languages and cultural contexts, representing a significant step toward more equitable and culturally-aware evaluation of AI systems in the diverse linguistic environments where they operate.