Government
Optimal Adaptive Experimental Design for Estimating Treatment Effect
Li, Jiachun, Simchi-Levi, David, Zhao, Yunxiao
Given n experiment subjects with potentially heterogeneous covariates and two possible treatments, namely active treatment and control, this paper addresses the fundamental question of determining the optimal accuracy in estimating the treatment effect. Furthermore, we propose an experimental design that approaches this optimal accuracy, giving a (non-asymptotic) answer to this fundamental yet still open question. The methodological contribution is listed as following. First, we establish an idealized optimal estimator with minimal variance as benchmark, and then demonstrate that adaptive experiment is necessary to achieve near-optimal estimation accuracy. Secondly, by incorporating the concept of doubly robust method into sequential experimental design, we frame the optimal estimation problem as an online bandit learning problem, bridging the two fields of statistical estimation and bandit learning. Using tools and ideas from both bandit algorithm design and adaptive statistical estimation, we propose a general low switching adaptive experiment framework, which could be a generic research paradigm for a wide range of adaptive experimental design. Through novel lower bound techniques for non-i.i.d. data, we demonstrate the optimality of our proposed experiment. Numerical result indicates that the estimation accuracy approaches optimal with as few as two or three policy updates.
UFO swarms filmed buzzing over Area 51 and other US military sites for months after 'mothership' encounter
Scores of new witnesses have emerged with more footage of the eerie'drone' UFO swarms buzzing key US military sites, including'a big fireball in a cube' over Area 51. The Las Vegas-area witness who reported this bizarre cube-shaped object claims to have observed similar strange aerial lights in the area'over 100 times' since June 2020, adding that these craft'always seem to head towards Nellis Air Force base.' Nevada's Nellis base and its sprawling complex about 40 miles northwest of Vegas -- including top secret Area 51, now legendary within UFO lore -- appear to have faced incursions by craft similar to those that plagued the Air Force in Virginia. For at least 17 nights last December, swarms of noisy small UFOs were seen'moving at rapid speeds' and displaying'flashing red, green, and white lights' within the highly restricted airspace over Virginia's Joint Base Langley–Eustis. Vegas natives have posted videos confirming they too have seen more than one red, green or white UFO that'wasn't flashing like a regular aircraft [or] like a satellite.' Another witness, who documented one September 4, 2024 case from their own 60-night experience with the odd lights, hoped coming forward might help get answers.
How a second Trump term could further enrich Elon Musk: 'There will be some quid pro quo'
Donald Trump owes his decisive 2024 presidential victory in no small part to the enthusiastic support of the world's richest man. In the months leading up to the election, Elon Musk put his full weight behind the Maga movement, advocated for Trump on major podcasts and used his influence over X to shape political discourse. Musk's America Pac injected nearly 120m into the former president's campaign. Now, Trump is looking to return the favor. Speaking with reporters last month, he said he would appoint Musk as "secretary of cost-cutting". Musk, for his part, has joked he would be interested in serving as head of the "Department of Government Efficiency" (Doge) with a stated goal of reducing government spending by 2tn.
Russia and Ukraine trade biggest drone attacks of conflict
Russia and Ukraine have both launched record drone attacks on each other overnight, with Ukrainian attacks on Moscow temporarily shutting down three of the Russian capital's airports. Russia fired 145 drones at Ukraine overnight, Ukrainian President Volodymyr Zelenskyy said on Sunday – more than in any single nighttime attack so far during their two-and-a-half-year conflict. "Last night, Russia launched a record 145 Shaheds and other strike drones against Ukraine," Zelenskyy said on social media, urging Kyiv's Western allies to do more to help Ukraine's defence. Kyiv said its air defences downed 62 of the drones. Russia also said it had downed 34 Ukrainian attack drones targeting Moscow on Sunday, the largest attempted attack on the capital since the start of the offensive in 2022, with Moscow regional Governor Andrei Vorobyov calling the attack "massive".
Russia and Ukraine launch biggest drone attacks against each other
Ukraine's attempted strike on Moscow was reportedly its largest attack on the capital since the war began, and was described as "massive" by the region's governor. One person was reported injured as drones were shot down near the Russian capital. Images on social media showed a residential building on fire. Most of the drones were downed in the Ramenskoye, Kolomna and Domodedovo districts, officials said. In September a woman was killed in a drone attack that hit Ramenskoye.
Mauritius election: Amid wiretapping scandal, what's at stake?
Some one million eligible voters in the Indian Ocean Mauritius will head out to vote on Sunday amid an explosive scandal that has implicated government figures in a covert wiretapping operation. Since independence from Britain in 1968, the southeast African country has maintained a strong, vibrant parliamentary democracy. This will be its 12th national election. Elections are usually deemed free and fair and turnout is normally high, at close to 80 percent. This time, however, the unusual drama caused by the leaked recordings has sparked national agitation and dominated the campaign season.
Field Insights for Portable Vine Robots in Urban Search and Rescue
McFarland, Ciera, Dhawan, Ankush, Kumari, Riya, Council, Chad, Coad, Margaret, Hanson, Nathaniel
Soft, growing vine robots are well-suited for exploring cluttered, unknown environments, and are theorized to be performant during structural collapse incidents caused by earthquakes, fires, explosions, and material flaws. These vine robots grow from the tip, enabling them to navigate rubble-filled passageways easily. State-of-the-art vine robots have been tested in archaeological and other field settings, but their translational capabilities to urban search and rescue (USAR) are not well understood. To this end, we present a set of experiments designed to test the limits of a vine robot system, the Soft Pathfinding Robotic Observation Unit (SPROUT), operating in an engineered collapsed structure. Our testing is driven by a taxonomy of difficulty derived from the challenges USAR crews face navigating void spaces and their associated hazards. Initial experiments explore the viability of the vine robot form factor, both ideal and implemented, as well as the control and sensorization of the system. A secondary set of experiments applies domain-specific design improvements to increase the portability and reliability of the system. SPROUT can grow through tight apertures, around corners, and into void spaces, but requires additional development in sensorization to improve control and situational awareness.
Flight Demonstration and Model Validation of a Prototype Variable-Altitude Venus Aerobot
Izraelevitz, Jacob S., Krishnamoorthy, Siddharth, Goel, Ashish, Turner, Caleb, Aiazzi, Carolina, Pauken, Michael, Carlson, Kevin, Walsh, Gerald, Leake, Carl, Quintana, Carlos, Lim, Christopher, Jain, Abhi, Dorsky, Leonard, Baines, Kevin, Cutts, James, Byrne, Paul K., Lachenmeier, Tim, Hall, Jeffery L.
This paper details a significant milestone towards maturing a buoyant aerial robotic platform, or aerobot, for flight in the Venus clouds. We describe two flights of our subscale altitude-controlled aerobot, fabricated from the materials necessary to survive Venus conditions. During these flights over the Nevada Black Rock desert, the prototype flew at the identical atmospheric densities as 54 to 55 km cloud layer altitudes on Venus. We further describe a first-principle aerobot dynamics model which we validate against the Nevada flight data and subsequently employ to predict the performance of future aerobots on Venus. The aerobot discussed in this paper is under JPL development for an in-situ mission flying multiple circumnavigations of Venus, sampling the chemical and physical properties of the planet's atmosphere and also remotely sensing surface properties.
Epistemic Integrity in Large Language Models
Ghafouri, Bijean, Mohammadzadeh, Shahrad, Zhou, James, Nair, Pratheeksha, Tian, Jacob-Junqi, Goel, Mayank, Rabbany, Reihaneh, Godbout, Jean-François, Pelrine, Kellin
Large language models are increasingly relied upon as sources of information, but their propensity for generating false or misleading statements with high confidence poses risks for users and society. In this paper, we confront the critical problem of epistemic miscalibration -- where a model's linguistic assertiveness fails to reflect its true internal certainty. We introduce a new human-labeled dataset and a novel method for measuring the linguistic assertiveness of Large Language Models (LLMs) which cuts error rates by over 50% relative to previous benchmarks. Validated across multiple datasets, our method reveals a stark misalignment between how confidently models linguistically present information and their actual accuracy. Further human evaluations confirm the severity of this miscalibration. This evidence underscores the urgent risk of the overstated certainty LLMs hold which may mislead users on a massive scale. Our framework provides a crucial step forward in diagnosing this miscalibration, offering a path towards correcting it and more trustworthy AI across domains. Large Language Models (LLMs) have markedly transformed how humans seek and consume information, becoming integral across diverse fields such as public health (Ali et al., 2023), coding (Zambrano et al., 2023), and education (Whalen & et al., 2023). Despite their growing influence, LLMs are not without shortcomings. One notable issue is the potential for generating responses that, while convincing, may be inaccurate or nonsensical--a long-standing phenomenon often referred to as "hallucinations" (Jo, 2023; Huang et al., 2023; Zhou et al., 2024b). This raises concerns about the reliability and trustworthiness of these models. A critical aspect of trustworthiness in LLMs is epistemic calibration, which represents the alignment between a model's internal confidence in its outputs and the way it expresses that confidence through natural language. Misalignment between internal certainty and external expression can lead to users being misled by overconfident or underconfident statements, posing significant risks in high-stakes domains such as legal advice, medical diagnosis, and misinformation detection. While of great normative concern, how LLMs express linguistic uncertainty has received relatively little attention to date (Sileo & Moens, 2023; Belem et al., 2024). Figures 1 and 5 illustrate the issue of epistemic calibration providing insights into the operation of certainty in the context of human interactions with LLMs. Distinct Roles of Certainty: Internal certainty and linguistic assertiveness have distinct functions within LLM interactions that shape individual beliefs. Human access to LLM certainty: Linguistic assertiveness holds a critical role as the primary form of certainty available to users.
Probabilistic Consensus through Ensemble Validation: A Framework for LLM Reliability
Large Language Models (LLMs) have shown significant advances in text generation but often lack the reliability needed for autonomous deployment in high-stakes domains like healthcare, law, and finance. Existing approaches rely on external knowledge or human oversight, limiting scalability. We introduce a novel framework that repurposes ensemble methods for content validation through model consensus. In tests across 78 complex cases requiring factual accuracy and causal consistency, our framework improved precision from 73.1% to 93.9% with two models (95% CI: 83.5%-97.9%) and to 95.6% with three models (95% CI: 85.2%-98.8%). Statistical analysis indicates strong inter-model agreement ($\kappa$ > 0.76) while preserving sufficient independence to catch errors through disagreement. We outline a clear pathway to further enhance precision with additional validators and refinements. Although the current approach is constrained by multiple-choice format requirements and processing latency, it offers immediate value for enabling reliable autonomous AI systems in critical applications.