Law
'It brings you closer to the natural world': the rise of the Merlin birdsong identifying app
'It brings you closer to the natural world': the rise of the Merlin birdsong identifying app W hen Natasha Walter first became curious about the birds around her, she recorded their songs on her phone and arduously tried to match each song with online recordings. After a friend recommended Merlin Bird ID, a free app, she tried it in her London garden and was delighted to discover the birds she assumed were female blackbirds - "this is how bad a birder I was" - were actually song thrushes and mistle thrushes. "I'm obsessed with Merlin - it's wonderful and it's been a joy to me," says Walter, a writer and human rights activist. "This is what AI and machine-learning have been invented for. Merlin is having a moment. The app, developed by the Cornell Lab of Ornithology in New York, which listens for birdsong and identifies the species singing, has been downloaded 33m times, in 240 countries and territories around the world. Britain has the second highest total number of users - more than 1.5 million in 2024, an 88% increase from 2023. Every month, there has been a 30% increase in new users of the app, whose sound identification function was launched in 2021. Merlin has been trained to identify the songs of more than 1,300 species around the world, with more birds added twice a year. Different songs make distinct patterns on spectrograms and Merlin is trained to recognise these different shapes and attribute them to a species. For latecomers to birding, or those lacking a knowledgeable friend, the app has become their teacher. "My fear at first was I wouldn't actually learn because I'm outsourcing my understanding of birds to this app," says Walter. "But that hasn't come to pass.
The Environmental and Human Rights Costs of China's Clean Energy Investments Abroad
If a major disaster like Fukushima or Chornobyl ever happens again, the world would know almost straight away, thanks to an array of government and DIY radiation-monitoring programs running globally. Why Don't Norwegians Hate Tesla Like the Rest of Europe Does? November's Tesla registrations were down in France, Sweden, Denmark, and Germany. Norway, however, is bucking the trend--thanks to a tax incentive system that will soon be rolled back.
Differentiable sorting for censored time-to-event data.
Survival analysis is a crucial semi-supervised task in machine learning with significant real-world applications, especially in healthcare. The most common approach to survival analysis, Cox's partial likelihood, can be interpreted as a ranking model optimized on a lower bound of the concordance index. We follow these connections further, with listwise ranking losses that allow for a relaxation of the pairwise independence assumption. Given the inherent transitivity of ranking, we explore differentiable sorting networks as a means to introduce a stronger transitive inductive bias during optimization.
AMDP: An Adaptive Detection Procedure for False Discovery Rate Control in High-Dimensional Mediation Analysis
High-dimensional mediation analysis is often associated with a multiple testing problem for detecting significant mediators. Assessing the uncertainty of this detecting process via false discovery rate (FDR) has garnered great interest. To control the FDR in multiple testing, two essential steps are involved: ranking and selection. Existing approaches either construct p-values without calibration or disregard the joint information across tests, leading to conservation in FDR control or non-optimal ranking rules for multiple hypotheses. In this paper, we develop an adaptive mediation detection procedure (referred to as AMDP) to identify relevant mediators while asymptotically controlling the FDR in high-dimensional mediation analysis. AMDP produces the optimal rule for ranking hypotheses and proposes a data-driven strategy to determine the threshold for mediator selection. This novel method captures information from the proportions of composite null hypotheses and the distribution of p-values, which turns the high dimensionality into an advantage instead of a limitation. The numerical studies on synthetic and real data sets illustrate the performances of AMDP compared with existing approaches.
Temporal Causal Mediation through a Point Process: Direct and Indirect Effects of Healthcare Interventions
Deciding on an appropriate intervention requires a causal model of a treatment, the outcome, and potential mediators. Causal mediation analysis lets us distinguish between direct and indirect effects of the intervention, but has mostly been studied in a static setting. In healthcare, data come in the form of complex, irregularly sampled time-series, with dynamic interdependencies between a treatment, outcomes, and mediators across time. Existing approaches to dynamic causal mediation analysis are limited to regular measurement intervals, simple parametric models, and disregard long-range mediator--outcome interactions. To address these limitations, we propose a non-parametric mediator--outcome model where the mediator is assumed to be a temporal point process that interacts with the outcome process. With this model, we estimate the direct and indirect effects of an external intervention on the outcome, showing how each of these affects the whole future trajectory. We demonstrate on semi-synthetic data that our method can accurately estimate direct and indirect effects.
The Harvard USPTO Patent Dataset: A Large-Scale, Well-Structured, and Multi-Purpose Corpus of Patent Applications
Innovation is a major driver of economic and social development, and information about many kinds of innovation is embedded in semi-structured data from patents and patent applications. Though the impact and novelty of innovations expressed in patent data are difficult to measure through traditional means, machine learning offers a promising set of techniques for evaluating novelty, summarizing contributions, and embedding semantics. In this paper, we introduce the Harvard USPTO Patent Dataset (HUPD), a large-scale, well-structured, and multi-purpose corpus of English-language patent applications filed to the United States Patent and Trademark Office (USPTO) between 2004 and 2018. With more than 4.5 million patent documents, HUPD is two to three times larger than comparable corpora. Unlike other NLP patent datasets, HUPD contains the inventor-submitted versions of patent applications, not the final versions of granted patents, allowing us to study patentability at the time of filing using NLP methods for the first time.