Government
SGS-GNN: A Supervised Graph Sparsification method for Graph Neural Networks
Das, Siddhartha Shankar, Arafat, Naheed Anjum, Rahman, Muftiqur, Ferdous, S M, Pothen, Alex, Halappanavar, Mahantesh M
We propose SGS-GNN, a novel supervised graph sparsifier that learns the sampling probability distribution of edges and samples sparse subgraphs of a user-specified size to reduce the computational costs required by GNNs for inference tasks on large graphs. SGS-GNN employs regularizers in the loss function to enhance homophily in sparse subgraphs, boosting the accuracy of GNNs on heterophilic graphs, where a significant number of the neighbors of a node have dissimilar labels. SGS-GNN also supports conditional updates of the probability distribution learning module based on a prior, which helps narrow the search space for sparse graphs. SGS-GNN requires fewer epochs to obtain high accuracies since it learns the search space of subgraphs more effectively than methods using fixed distributions such as random sampling. Extensive experiments using 33 homophilic and heterophilic graphs demonstrate the following: (i) with only 20% of edges retained in the sparse subgraphs, SGS-GNN improves the F1-scores by a geometric mean of 4% relative to the original graph; on heterophilic graphs, the prediction accuracy is better up to 30%. (ii) SGS-GNN outperforms state-of-the-art methods with improvement in F1-scores of 4-7% in geometric mean with similar sparsities in the sampled subgraphs, and (iii) compared to sparsifiers that employ fixed distributions, SGS-GNN requires about half the number of epochs to converge.
Large Language Models and Synthetic Data for Monitoring Dataset Mentions in Research Papers
Solatorio, Aivin V., Macalaba, Rafael, Liounis, James
Tracking how data is mentioned and used in research papers provides critical insights for improving data discoverability, quality, and production. However, manually identifying and classifying dataset mentions across vast academic literature is resource-intensive and not scalable. This paper presents a machine learning framework that automates dataset mention detection across research domains by leveraging large language models (LLMs), synthetic data, and a two-stage fine-tuning process. We employ zero-shot extraction from research papers, an LLM-as-a-Judge for quality assessment, and a reasoning agent for refinement to generate a weakly supervised synthetic dataset. The Phi-3.5-mini instruct model is pre-fine-tuned on this dataset, followed by fine-tuning on a manually annotated subset. At inference, a ModernBERT-based classifier efficiently filters dataset mentions, reducing computational overhead while maintaining high recall. Evaluated on a held-out manually annotated sample, our fine-tuned model outperforms NuExtract-v1.5 and GLiNER-large-v2.1 in dataset extraction accuracy. Our results highlight how LLM-generated synthetic data can effectively address training data scarcity, improving generalization in low-resource settings. This framework offers a pathway toward scalable monitoring of dataset usage, enhancing transparency, and supporting researchers, funders, and policymakers in identifying data gaps and strengthening data accessibility for informed decision-making.
Probing Perceptual Constancy in Large Vision Language Models
Sun, Haoran, Yu, Suyang, Li, Yijiang, Gao, Qingying, Lyu, Haiyun, Deng, Hokin, Luo, Dezhi
Perceptual constancy is the ability to maintain stable perceptions of objects despite changes in sensory input, such as variations in distance, angle, or lighting. This ability is crucial for recognizing visual information in a dynamic world, making it essential for Vision-Language Models (VLMs). However, whether VLMs are currently and theoretically capable of mastering this ability remains underexplored. In this study, we evaluated 33 VLMs using 253 experiments across three domains: color, size, and shape constancy. The experiments included single-image and video adaptations of classic cognitive tasks, along with novel tasks in in-the-wild conditions, to evaluate the models' recognition of object properties under varying conditions. We found significant variability in VLM performance, with models performance in shape constancy clearly dissociated from that of color and size constancy.
Organize the Web: Constructing Domains Enhances Pre-Training Data Curation
Wettig, Alexander, Lo, Kyle, Min, Sewon, Hajishirzi, Hannaneh, Chen, Danqi, Soldaini, Luca
Modern language models are trained on large, unstructured datasets consisting of trillions of tokens and obtained by crawling the web. The unstructured nature makes it difficult to reason about their contents and develop systematic approaches to data curation. In this paper, we unpack monolithic web corpora by developing taxonomies of their contents and organizing them into domains. We introduce WebOrganizer, a framework for organizing web pages in terms of both their topic and format. Using these two complementary notions of domains, we automatically annotate pre-training data by distilling annotations from a large language model into efficient classifiers. This allows us to study how data from different domains should be mixed to improve models on downstream tasks, and we show that we can combine insights about effective topics and formats to further boost performance. We demonstrate that our domain mixing also improves existing methods that select data based on quality. Furthermore, we study and compare how quality-based methods will implicitly change the domain mixture. Overall, our work demonstrates that constructing and mixing domains provides a valuable complement to quality-based data curation methods, opening new avenues for effective and insightful pre-training data curation.
Assortment Optimization for Patient-Provider Matching
Primary care providers are essential to the healthcare ecosystem because they are the first point of contact for many patients (Pearson and Raeke, 2000; Wu et al., 2022). Patients rely on primary care providers for routine checkups and referrals to specialists. Moreover, care continuity can instill trust and improve medication uptake rates and patient health (Wu et al., 2022). Unfortunately, high provider turnover rates frequently lead to patients without an assigned provider (Reddy et al., 2015). Provider turnover disrupts patient care and leads to worse care (Reddy et al., 2015). In principle, healthcare administrators reassign unmatched patients to other providers; however, in practice, the process takes months due to provider scarcity and the logistical burden of rematching and coordinating patient matches (Hedden et al., 2021). While many patients find their new provider quickly, others have to wait years to find a new provider due to large numbers of patients, high turnover rates, and provider scarcity (Hedden et al., 2021; Shanafelt et al., 2012). Algorithms that automatically match patients and providers can reduce logistical hassle but require balancing patient autonomy and system-wide utility. For example, while automatically assigning each patient to a provider would decrease wait times, it also reduces patient autonomy because patients cannot select their provider (Entwistle et al., 2010; Gaynor et al., 2016). 1
Russia launches fresh drone attack against Ukraine shortly after Trump-Putin phone call
Fox News senior White House correspondent Jacqui Heinrich has the latest on peace talks on'Special Report.' Ukraine's air force indicated in a Facebook post on Thursday that the Eastern European nation had been targeted in a drone attack overnight. "85 ENEMY UAVS SHOT, 52 DRONES FAILED TO REACH THEIR TARGETS (LOCATIONALLY LOST)," the top of the post read, according to a Google translation of the Ukrainian text. The announcement came after U.S. President Donald Trump noted on Wednesday that he had spoken to both Russian President Vladimir Putin and Ukrainian President Volodymyr Zelenskyy. TRUMP SAYS RUSSIA AGREES TO'IMMEDIATELY' BEGIN NEGOTIATIONS TO END WAR IN UKRAINE Ukraine's President Volodymyr Zelensky speaks during a joint press conference with the President of the European Investment Bank (EIB) in Kyiv on Feb. 10, 2025, amid the Russian invasion of Ukraine (TETIANA DZHAFAROVA/AFP via Getty Images) In a Truth Social post, the president described his call with Putin as "lengthy and highly productive."
Rogue states could use AI to do 'real harm', warns ex-Google CEO
Google's former chief executive has warned that artificial intelligence could be used by rogue states such as North Korea, Iran and Russia to "harm innocent people". Eric Schmidt, who held senior posts at Google from 2001 to 2017, told BBC Radio 4's Today programme that those countries and terrorists could adopt and misuse the technology to develop weapons to create "a bad biological attack from some evil person". The tech billionaire said: "The real fears that I have are not the ones that most people talk about AI โ I talk about extreme risk. "Think about North Korea, or Iran, or even Russia, who have some evil goal. This technology is fast enough for them to adopt that they could misuse it and do real harm."
Here's How We Can Power the AI Boom Without Building a Ton of New Gas Plants
This story was originally published on the author's substack, Field Notes with Alexander C Kaufman, to which you can subscribe here. Artificial intelligence is driving up demand for electricity--the only question is how much, and what provides the power. Over the next three years, the Lawrence Berkeley National Laboratory estimates, AI's thirst for power will double or triple. Last month, OpenAI unveiled its Stargate Project, a plan to invest 500 billion in the infrastructure for artificial intelligence over the next four years that includes adding 25 gigawatts of new electricity capacity. Right now, the most likely source of electricity to power those data centers is gas.
What the Assault on Public Education Means for Kids with Disabilities
President Donald Trump, winner of the Battle of the Billionaires at WrestleMania 23, has maintained close ties with Linda McMahon, the former C.E.O. of World Wrestling Entertainment, for decades. During the President's first term, she served for two years as head of the Small Business Administration, stepping down in 2019 to lead America First Action, a pro-Trump super PAC. Now McMahon is Trump's nominee to run the U.S. Department of Education, although she may appear to lack conventional bona fides for the position. If McMahon is confirmed by the Senate, her odd task will be to take charge of an agency in order to euthanize it. "I told Linda, 'Linda, I hope you do a great job and put yourself out of a job,' " Trump said, on February 4th.
Massive AI Stargate Project under Trump admin reveals next steps
Stargate, the massive artificial intelligence (AI) infrastructure project recently unveiled by President Donald Trump, has begun production in Texas -- with data center construction in other states expected to be announced in the coming months. OpenAI, Softbank, Oracle and other partners' total investment of 500 million in the project will produce a large-scale network of campuses. Each campus will be designed in the roughly 1 gigawatt (GW) or greater range, a measurement of electricity that can power a minimum of 750,000 homes. During a recent press briefing on The Stargate Project attended by Fox News Digital, OpenAI announced that construction on the first site is underway in Abilene, Texas. Significant progress has been made in identifying additional locations.