Scientific Discovery
Strategic Hypothesis Testing
Hossain, Safwan, Chen, Yatong, Chen, Yiling
We examine hypothesis testing within a principal-agent framework, where a strategic agent, holding private beliefs about the effectiveness of a product, submits data to a principal who decides on approval. The principal employs a hypothesis testing rule, aiming to pick a p-value threshold that balances false positives and false negatives while anticipating the agent's incentive to maximize expected profitability. Building on prior work, we develop a game-theoretic model that captures how the agent's participation and reporting behavior respond to the principal's statistical decision rule. Despite the complexity of the interaction, we show that the principal's errors exhibit clear monotonic behavior when segmented by an efficiently computable critical p-value threshold, leading to an interpretable characterization of their optimal p-value threshold. We empirically validate our model and these insights using publicly available data on drug approvals. Overall, our work offers a comprehensive perspective on strategic interactions within the hypothesis testing framework, providing technical and regulatory insights.
RAISE: Enhancing Scientific Reasoning in LLMs via Step-by-Step Retrieval
Oh, Minhae, Kim, Jeonghye, Lee, Nakyung, Seo, Donggeon, Kim, Taeuk, Lee, Jungwoo
Scientific reasoning requires not only long-chain reasoning processes, but also knowledge of domain-specific terminologies and adaptation to updated findings. To deal with these challenges for scientific reasoning, we introduce RAISE, a step-by-step retrieval-augmented framework which retrieves logically relevant documents from in-the-wild corpus. RAISE is divided into three steps: problem decomposition, logical query generation, and logical retrieval. We observe that RAISE consistently outperforms other baselines on scientific reasoning benchmarks. We analyze that unlike other baselines, RAISE retrieves documents that are not only similar in terms of the domain knowledge, but also documents logically more relevant.
Splits! A Flexible Dataset and Evaluation Framework for Sociocultural Linguistic Investigation
Caplan, Eylon, Chakraborty, Tania, Goldwasser, Dan
Variation in language use, shaped by speakers' sociocultural background and specific context of use, offers a rich lens into cultural perspectives, values, and opinions. However, the computational study of these Sociocultural Linguistic Phenomena (SLP) has often been limited to bespoke analyses of specific groups or topics, hindering the pace of scientific discovery. To address this, we introduce Splits!, a 9.7 million-post dataset from Reddit designed for systematic and flexible research. The dataset contains posts from over 53,000 users across 6 demographic groups, organized into 89 discussion topics to enable comparative analysis. We validate Splits! via self-identification and by successfully replicating several known SLPs from existing literature. We complement this dataset with a framework that leverages efficient retrieval methods to rapidly validate potential SLPs (PSLPs) by automatically evaluating whether a given hypothesis is supported by our data. Crucially, to distinguish between novel and obvious insights, the framework incorporates a human-validated measure of a hypothesis's ``unexpectedness.'' We demonstrate that the two-stage process reduces the number of statistically significant findings requiring manual inspection by a factor of 1.5-1.8x, streamlining the discovery of promising phenomena for further investigation.
HIAL: A New Paradigm for Hypergraph Active Learning via Influence Maximization
Hou, Yanheng, Li, Xunkai, Li, Zhenjun, Zhou, Bing, Li, Ronghua, Wang, Guoren
In recent years, Hypergraph Neural Networks (HNNs) have demonstrated immense potential in handling complex systems with high-order interactions. However, acquiring large-scale, high-quality labeled data for these models is costly, making Active Learning (AL) a critical technique. Existing Graph Active Learning (GAL) methods, when applied to hypergraphs, often rely on techniques like "clique expansion," which destroys the high-order structural information crucial to a hypergraph's success, thereby leading to suboptimal performance. To address this challenge, we introduce HIAL (Hypergraph Active Learning), a native active learning framework designed specifically for hypergraphs. We innovatively reformulate the Hypergraph Active Learning (HAL) problem as an Influence Maximization task. The core of HIAL is a dual-perspective influence function that, based on our novel "High-Order Interaction-Aware (HOI-Aware)" propagation mechanism, synergistically evaluates a node's feature-space coverage (via Magnitude of Influence, MoI) and its topological influence (via Expected Diffusion Value, EDV). We prove that this objective function is monotone and submodular, thus enabling the use of an efficient greedy algorithm with a formal (1-1/e) approximation guarantee. Extensive experiments on seven public datasets demonstrate that HIAL significantly outperforms state-of-the-art baselines in terms of performance, efficiency, generality, and robustness, establishing an efficient and powerful new paradigm for active learning on hypergraphs.
1,000-year-old medieval sword emerges from Dutch river after chance discovery: 'Barely corroded'
SOLVA Archaeology Service in Belgium announced the recent discovery of ancient Roman artifacts and remains, including a well-preserved dog, in Velzeke. A remarkable medieval sword with rare symbols was recently put on display in a Dutch museum, over a year after it was found by construction workers unexpectedly. The discovery of the sword was announced by the Netherlands' National Museum of Antiquities (RMO) in Leiden on June 24. The artifact, named the Linschoten Sword, was found in March 2024 during "maintenance dredging activities," the museum said in a press release. Construction workers were struck by a "long piece of iron" while cleaning a small river known as the Korte Linschoten, the statement noted.
Ultra-rare first edition book from Galileo heading to auction
Breakthroughs, discoveries, and DIY tips sent every weekday. A small library's worth of rare medieval and Renaissance books are heading to auction on July 9. The expansive lot includes a portable Magna Carta, an early scientific encyclopedia, a surgical codex, and one of the oldest surviving Sephardic Torah scrolls. But according to Christies's Auction House, one manuscript is the first of its kind to go up for sale in over a century: a copy of the first pseudonymous astronomical text co-written by Galileo Galilei. The evening of October 9, 1604, offered an unexpected and ultimately revolutionary moment for astronomy.
Peer Review as Structured Commentary: Immutable Identity, Public Dialogue, and Reproducible Scholarship
This paper reconceptualises peer review as structured public commentary. Traditional academic validation is hindered by anonymity, latency, and gatekeeping. We propose a transparent, identity-linked, and reproducible system of scholarly evaluation anchored in open commentary. Leveraging blockchain for immutable audit trails and AI for iterative synthesis, we design a framework that incentivises intellectual contribution, captures epistemic evolution, and enables traceable reputational dynamics. This model empowers fields from computational science to the humanities, reframing academic knowledge as a living process rather than a static credential.
Position: Intelligent Science Laboratory Requires the Integration of Cognitive and Embodied AI
Zhang, Sha, Yang, Suorong, Xie, Tong, Xue, Xiangyuan, Hu, Zixuan, Li, Rui, Qu, Wenxi, Yin, Zhenfei, Fu, Tianfan, Hu, Di, Bran, Andres M, Ran, Nian, Hoex, Bram, Zuo, Wangmeng, Schwaller, Philippe, Ouyang, Wanli, Bai, Lei, Zhang, Yanyong, Duan, Lingyu, Tang, Shixiang, Zhou, Dongzhan
Scientific discovery has long been constrained by human limitations in expertise, physical capability, and sleep cycles. The recent rise of AI scientists and automated laboratories has accelerated both the cognitive and operational aspects of research. However, key limitations persist: AI systems are often confined to virtual environments, while automated laboratories lack the flexibility and autonomy to adaptively test new hypotheses in the physical world. Recent advances in embodied AI, such as generalist robot foundation models, diffusion-based action policies, fine-grained manipulation learning, and sim-to-real transfer, highlight the promise of integrating cognitive and embodied intelligence. This convergence opens the door to closed-loop systems that support iterative, autonomous experimentation and the possibility of serendipitous discovery. In this position paper, we propose the paradigm of Intelligent Science Laboratories (ISLs): a multi-layered, closed-loop framework that deeply integrates cognitive and embodied intelligence. ISLs unify foundation models for scientific reasoning, agent-based workflow orchestration, and embodied agents for robust physical experimentation. We argue that such systems are essential for overcoming the current limitations of scientific discovery and for realizing the full transformative potential of AI-driven science.
Diffusion-Based Hypothesis Testing and Change-Point Detection
Moushegian, Sean, Banerjee, Taposh, Tarokh, Vahid
Score-based methods have recently seen increasing popularity in modeling and generation. Methods have been constructed to perform hypothesis testing and change-point detection with score functions, but these methods are in general not as powerful as their likelihood-based peers. Recent works consider generalizing the score-based Fisher divergence into a diffusion-divergence by transforming score functions via multiplication with a matrix-valued function or a weight matrix. In this paper, we extend the score-based hypothesis test and change-point detection stopping rule into their diffusion-based analogs. Additionally, we theoretically quantify the performance of these diffusion-based algorithms and study scenarios where optimal performance is achievable. We propose a method of numerically optimizing the weight matrix and present numerical simulations to illustrate the advantages of diffusion-based algorithms.
From Data to Decision: Data-Centric Infrastructure for Reproducible ML in Collaborative eScience
Li, Zhiwei, Kesselman, Carl, Nguyen, Tran Huy, Xu, Benjamin Yixing, Bolo, Kyle, Yu, Kimberley
--Reproducibility remains a central challenge in machine learning (ML), especially in collaborative eScience projects where teams iterate over data, features, and models. Current ML workflows are often dynamic yet fragmented, relying on informal data sharing, ad hoc scripts, and loosely connected tools. This fragmentation impedes transparency, reproducibility, and the adaptability of experiments over time. This paper introduces a data-centric framework for lifecycle-aware reproducibility, centered around six structured artifacts: Dataset, Feature, Workflow, Execution, Asset, and Controlled V ocabulary. These artifacts formalize the relationships between data, code, and decisions, enabling ML experiments to be versioned, interpretable, and traceable over time. The approach is demonstrated through a clinical ML use case of glaucoma detection, illustrating how the system supports iterative exploration, improves reproducibility, and preserves the provenance of collaborative decisions across the ML lifecycle. As machine learning (ML) becomes increasingly central to scientific discovery, concerns about correctness and reproducibility have grown [1]. In eScience, ML development is typically a collaborative and iterative process involving domain experts, data engineers, and ML researchers. These teams refine models based on evolving hypotheses and new data, creating feedback loops across data curation, feature engineering, modeling, and evaluation [2]. This dynamic process frequently introduces data cascades, where early curation errors propagate downstream, compounding over time [3]. In practice, ML workflows remain fragmented: datasets are shared informally, experiments span personal and cloud environments, and data, code, and configurations are often loosely coupled [4]. While MLOps and data management tools address parts of this problem, such as code versioning, pipeline orchestration, or environment encapsulation, they often overlook the full scientific lifecycle and the socio-technical realities of collaborative ML projects [5]. In prior work, we introduced Deriva-ML [6], a socio-technical platform that extends the FAIR principles (Findable, Accessible, Interoperable, Reusable) [7] across the ML developmental lifecycle.