Government
A Dynamic Framework for Semantic Grouping of Common Data Elements (CDE) Using Embeddings and Clustering
Krishnamurthy, Madan, Korn, Daniel, Haendel, Melissa A, Mungall, Christopher J, Thessen, Anne E
This research aims to develop a dynamic and scalable framework to facilitate harmonization of Common Data Elements (CDEs) across heterogeneous biomedical datasets by addressing challenges such as semantic heterogeneity, structural variability, and context dependence to streamline integration, enhance interoperability, and accelerate scientific discovery. Our methodology leverages Large Language Models (LLMs) for context-aware text embeddings that convert CDEs into dense vectors capturing semantic relationships and patterns. These embeddings are clustered using Hierarchical Density-Based Spatial Clustering of Applications with Noise (HDBSCAN) to group semantically similar CDEs. The framework incorporates four key steps: (1) LLM-based text embedding to mathematically represent semantic context, (2) unsupervised clustering of embeddings via HDBSCAN, (3) automated labeling using LLM summarization, and (4) supervised learning to train a classifier assigning new or unclustered CDEs to labeled clusters. Evaluated on the NIH NLM CDE Repository with over 24,000 CDEs, the system identified 118 meaningful clusters at an optimized minimum cluster size of 20. The classifier achieved 90.46 percent overall accuracy, performing best in larger categories. External validation against Gravity Projects Social Determinants of Health domains showed strong agreement (Adjusted Rand Index 0.52, Normalized Mutual Information 0.78), indicating that embeddings effectively capture cluster characteristics. This adaptable and scalable approach offers a practical solution to CDE harmonization, improving selection efficiency and supporting ongoing data interoperability.
Will Agents Replace Us? Perceptions of Autonomous Multi-Agent AI
Autonomous multi-agent AI systems are poised to transform various industries, particularly software development and knowledge work. Understanding current perceptions among professionals is crucial for anticipating adoption challenges, ethical considerations, and future workforce development. This study analyzes responses from 130 participants to a survey on the capabilities, impact, and governance of AI agents. We explore expected timelines for AI replacing programmers, identify perceived barriers to deployment, and examine beliefs about responsibility when agents make critical decisions. Key findings reveal three distinct clusters of respondents. While the study explored factors associated with current AI agent deployment, the initial logistic regression model did not yield statistically significant predictors, suggesting that deployment decisions are complex and may be influenced by factors not fully captured or that a larger sample is needed. These insights highlight the need for organizations to address compliance concerns (a commonly cited barrier) and establish clear governance frameworks as they integrate autonomous agents into their workflows.
The End Of Universal Lifelong Identifiers: Identity Systems For The AI Era
Many identity systems assign a single, static identifier to an individual for life, reused across domains like healthcare, finance, and education. These Universal Lifelong Identifiers (ULIs) underpin critical workflows but now pose systemic privacy risks. We take the position that ULIs are fundamentally incompatible with the AI era and must be phased out. We articulate a threat model grounded in modern AI capabilities and show that traditional safeguards such as redaction, consent, and access controls are no longer sufficient. We define core properties for identity systems in the AI era and present a cryptographic framework that satisfies them while retaining compatibility with existing identifier workflows. Our design preserves institutional workflows, supports essential functions such as auditability and delegation, and offers a practical migration path beyond ULIs.
COMPKE: Complex Question Answering under Knowledge Editing
Cheng, Keyuan, Kan, Zijian, He, Zhixian, Zhang, Zhuoran, Ali, Muhammad Asif, Xu, Ke, Hu, Lijie, Wang, Di
Knowledge Editing, which efficiently modifies the knowledge in large language models, has gathered great attention. Current benchmarks primarily use multi-hop question answering to assess and analyze newly injected or updated knowledge. However, we argue that these benchmarks fail to effectively evaluate how well the updated models apply this knowledge in real-life scenarios, particularly when questions require complex reasoning, involving one-to-many relationships or multi-step logical intersections. To fill in this gap, we introduce a new benchmark, COMPKE: Complex Question Answering under Knowledge Editing, which includes 11,924 complex questions that reflect real-life situations. We conduct an extensive evaluation of four knowledge editing methods on COMPKE, revealing that their effectiveness varies notably across different models. For instance, MeLLo attains an accuracy of 39.47 on GPT-4O-MINI, but this drops sharply to 3.83 on QWEN2.5-3B. We further investigate the underlying causes of these disparities from both methodological and model-specific perspectives. The datasets are available at https://github.com/kzjkzj666/CompKE.
Time-R1: Towards Comprehensive Temporal Reasoning in LLMs
Liu, Zijia, Han, Peixuan, Yu, Haofei, Li, Haoru, You, Jiaxuan
Large Language Models (LLMs) demonstrate impressive capabilities but lack robust temporal intelligence, struggling to integrate reasoning about the past with predictions and plausible generations of the future. Meanwhile, existing methods typically target isolated temporal skills, such as question answering about past events or basic forecasting, and exhibit poor generalization, particularly when dealing with events beyond their knowledge cutoff or requiring creative foresight. To address these limitations, we introduce \textit{Time-R1}, the first framework to endow a moderate-sized (3B-parameter) LLM with comprehensive temporal abilities: understanding, prediction, and creative generation. Our approach features a novel three-stage development path; the first two constitute a \textit{reinforcement learning (RL) curriculum} driven by a meticulously designed dynamic rule-based reward system. This framework progressively builds (1) foundational temporal understanding and logical event-time mappings from historical data, (2) future event prediction skills for events beyond its knowledge cutoff, and finally (3) enables remarkable generalization to creative future scenario generation without any fine-tuning. Strikingly, experiments demonstrate that Time-R1 outperforms models over 200 times larger, including the state-of-the-art 671B DeepSeek-R1, on highly challenging future event prediction and creative scenario generation benchmarks. This work provides strong evidence that thoughtfully engineered, progressive RL fine-tuning allows smaller, efficient models to achieve superior temporal performance, offering a practical and scalable path towards truly time-aware AI. To foster further research, we also release \textit{Time-Bench}, a large-scale multi-task temporal reasoning dataset derived from 10 years of news data, and our series of \textit{Time-R1} checkpoints.
Indiana senator calls on WNBA, Fever to apologize to fans after accusations of racism: 'So demeaning'
Republican Sen. Jim Banks explains why Indiana Fever fans deserve an apology after the league's latest investigation during an appearance on OutKick's'Don't @ Me with Dan Dakich.' U.S. Sen. Jim Banks, R-Ind., called on the WNBA and the Indiana Fever to apologize to Fever fans after the league's investigation failed to find evidence that corroborated allegations of racial comments directed at Angel Reese during a recent game. The league investigated the allegations involving the Chicago Sky star last month after a May 17 game hosted by the Fever. Chicago Sky forward Angel Reese (5) reacts to a flagrant foul from Indiana Fever guard Caitlin Clark (22) May 17, 2025, at Gainbridge Fieldhouse in Indianapolis. "Based on information gathered to date, including from relevant fans, team and arena staff, as well as audio and video review of the game, we have not substantiated [the report,]" the league said in a statement.
California Senate passes bill that aims to make AI chatbots safer
California lawmakers on Tuesday moved one step closer to placing more guardrails around artificial intelligence-powered chatbots. The Senate passed a bill that aims to make chatbots used for companionship safer after parents raised concerns that virtual characters harmed their childrens' mental health. An artificial intelligence startup is under fire for allegedly releasing chatbots that harmed the mental health of young people. The legislation, which now heads to the California State Assembly, shows how state lawmakers are tackling safety concerns surrounding AI as tech companies release more AI-powered tools. "The country is watching again for California to lead," said Sen. Steve Padilla (D-Chula Vista), one of the lawmakers who introduced the bill, on the Senate floor.
Secret CIA program claimed to have found alien civilization on dark side of the moon: 'They look like us'
As the US prepares to send astronauts back to the moon, a CIA file has resurfaced that claims to have found life there more than 25 years ago. In the 1970s and 80s, the CIA conducted experiments with individuals who claimed they could perceive information about distant objects, events, or people, a process known as'remote viewing.' The experience of remote viewer Ingo Swann was first revealed in 1998 when he explained how his psychic episode took him to the dark side of the moon, a region that always faces away from Earth and out of sight from human eyes. That's where the remote reviewer made a shocking discovery: towers, buildings, and human-like aliens working at a secret complex on the moon's surface. Disturbingly, Swann said government officials knew the aliens had a base there, and these humanoids could actually sense his presence as he viewed them with his mind from 238,000 miles away.
Dating apps used in Mexico to lure and kidnap U.S. citizens, officials warn
U.S. citizens who visit Mexico are being warned that they may be at risk of being kidnapped by people who lure them in through dating apps, according to federal officials. The U.S. Consulate General Guadalajara warned that the victims of such schemes were kidnapped in Puerto Vallarta and Nuevo Nayarit areas in recent months, according to a news release. The consulate did not say how often this type of crime has occurred or whether any suspects have been arrested. Victims and their family members were extorted for large amounts of money in order to be released, officials said. Some of the victims met their captors in residences or hotel rooms.
FDA approves first AI tool to predict breast cancer risk
Senior medical analyst Dr. Marc Siegel discusses advancements in artificial intelligence aimed at predicting an individual's future risk of breast cancer and the increased health risks from cannabis as users age. The U.S. Food and Drug Administration (FDA) has approved the first artificial intelligence (AI) tool to predict breast cancer risk. The authorization was confirmed by digital health tech company Clairity, the developer of Clairity Breast – a novel, image-based prognostic platform designed to predict five-year breast cancer risk from a routine screening mammogram. In a press release, Clairity shared its plans to launch the AI platform across health systems through 2025. Most risk assessment models for breast cancer rely heavily on age and family history, according to Clairity.