Africa
Robustly estimating heterogeneity in factorial data using Rashomon Partitions
Venkateswaran, Aparajithan, Sankar, Anirudh, Chandrasekhar, Arun G., McCormick, Tyler H.
Many statistical analyses, in both observational data and randomized control trials, ask: how does the outcome of interest vary with combinations of observable covariates? How do various drug combinations affect health outcomes, or how does technology adoption depend on incentives and demographics? Our goal is to partition this factorial space into ``pools'' of covariate combinations where the outcome differs across the pools (but not within a pool). Existing approaches (i) search for a single ``optimal'' partition under assumptions about the association between covariates or (ii) sample from the entire set of possible partitions. Both these approaches ignore the reality that, especially with correlation structure in covariates, many ways to partition the covariate space may be statistically indistinguishable, despite very different implications for policy or science. We develop an alternative perspective, called Rashomon Partition Sets (RPSs). Each item in the RPS partitions the space of covariates using a tree-like geometry. RPSs incorporate all partitions that have posterior values near the maximum a posteriori partition, even if they offer substantively different explanations, and do so using a prior that makes no assumptions about associations between covariates. This prior is the $\ell_0$ prior, which we show is minimax optimal. Given the RPS we calculate the posterior of any measurable function of the feature effects vector on outcomes, conditional on being in the RPS. We also characterize approximation error relative to the entire posterior and provide bounds on the size of the RPS. Simulations demonstrate this framework allows for robust conclusions relative to conventional regularization techniques. We apply our method to three empirical settings: price effects on charitable giving, chromosomal structure (telomere length), and the introduction of microfinance.
Preventing Model Collapse in Gaussian Process Latent Variable Models
Li, Ying, Lin, Zhidi, Yin, Feng, Zhang, Michael Minyi
Gaussian process latent variable models (GPLVMs) are a versatile family of unsupervised learning models, commonly used for dimensionality reduction. However, common challenges in modeling data with GPLVMs include inadequate kernel flexibility and improper selection of the projection noise, which leads to a type of model collapse characterized primarily by vague latent representations that do not reflect the underlying structure of the data. This paper addresses these issues by, first, theoretically examining the impact of the projection variance on model collapse through the lens of a linear GPLVM. Second, we address the problem of model collapse due to inadequate kernel flexibility by integrating the spectral mixture (SM) kernel and a differentiable random Fourier feature (RFF) kernel approximation, which ensures computational scalability and efficiency through off-the-shelf automatic differentiation tools for learning the kernel hyperparameters, projection variance, and latent representations within the variational inference framework. The proposed GPLVM, named advisedRFLVM, is evaluated across diverse datasets and consistently outperforms various salient competing models, including state-of-the-art variational autoencoders (VAEs) and GPLVM variants, in terms of informative latent representations and missing data imputation.
Thailand's economy stumbles as Philippines, Vietnam, Indonesia race ahead
Bangkok, Thailand – Sheltering from the sun on a street corner, Kridsada Ahjed rues the day he got involved with the loan sharks who now gobble up most of his daily earnings. "I went to the loan sharks because people like me – with no assets or savings – cannot qualify to get help from legitimate banks," Ahjed, a 40-year-old motorcycle taxi driver, told Al Jazeera. "Now almost everything I make in a day goes towards paying the interest on my debt." Kridsada is far from alone. Thailand's household debt reached nearly 87 percent of gross domestic product last year, according to the Bank of Thailand, among the highest on earth.
Retired Admiral William McRaven on Why U.S. Leadership Matters
Retired Navy Adm. William McRaven's nearly 40-year career in the U.S. military has spanned everything from deployments as a Navy SEAL, hunting down high-value targets overseas, commanding U.S Special Operations forces in Iraq and Afghanistan, and advising Presidents George W. Bush and Barack Obama. But McRaven is best known for planning and overseeing the 2011 raid that ended with the death of Osama bin Laden. In December that year, McRaven was named as a runner-up for TIME's Person of the Year for his role in the operation. "There is nobody in the U.S. government that thinks we can kill our way to victory, certainly not the special-operations guys," he told TIME in 2011, "but what happens is, by capturing and killing some of these high-value targets, we buy space and time for the rest of the government to work." After retiring from the U.S. military in 2014, McRaven served as the chancellor of the University of Texas System and has written several books on leadership.
OpenAI debuts voice cloning tool, but deems it too risky for public release
OpenAI has unveiled a tool for cloning people's voices but is holding back on its public release due to concerns about possible misuse in a key election year. Voice Engine can replicate a person's voice based on a 15-second audio sample, according to an OpenAI blog post demonstrating the tool. But the ChatGPT creator is "taking a cautious and informed approach" to the technology and hopes to start a dialogue on "the responsible deployment of synthetic voices", the company said in the blog post published on Friday. "We recognize that generating speech that resembles people's voices has serious risks, which are especially top of mind in an election year," the San Francisco-based start-up said. "We are engaging with U.S. and international partners from across government, media, entertainment, education, civil society and beyond to ensure we are incorporating their feedback as we build."
Dialogue with Robots: Proposals for Broadening Participation and Research in the SLIVAR Community
Kennington, Casey, Alikhani, Malihe, Pon-Barry, Heather, Atwell, Katherine, Bisk, Yonatan, Fried, Daniel, Gervits, Felix, Han, Zhao, Inan, Mert, Johnston, Michael, Korpan, Raj, Litman, Diane, Marge, Matthew, Matuszek, Cynthia, Mead, Ross, Mohan, Shiwali, Mooney, Raymond, Parde, Natalie, Sinapov, Jivko, Stewart, Angela, Stone, Matthew, Tellex, Stefanie, Williams, Tom
The ability to interact with machines using natural human language is becoming not just commonplace, but expected. The next step is not just text interfaces, but speech interfaces and not just with computers, but with all machines including robots. In this paper, we chronicle the recent history of this growing field of spoken dialogue with robots and offer the community three proposals, the first focused on education, the second on benchmarks, and the third on the modeling of language when it comes to spoken interaction with robots. The three proposals should act as white papers for any researcher to take and build upon.
Are large language models superhuman chemists?
Mirza, Adrian, Alampara, Nawaf, Kunchapu, Sreekanth, Emoekabu, Benedict, Krishnan, Aswanth, Wilhelmi, Mara, Okereke, Macjonathan, Eberhardt, Juliane, Elahi, Amir Mohammad, Greiner, Maximilian, Holick, Caroline T., Gupta, Tanya, Asgari, Mehrdad, Glaubitz, Christina, Klepsch, Lea C., Köster, Yannik, Meyer, Jakob, Miret, Santiago, Hoffmann, Tim, Kreth, Fabian Alexander, Ringleb, Michael, Roesner, Nicole, Schubert, Ulrich S., Stafast, Leanne M., Wonanke, Dinga, Pieler, Michael, Schwaller, Philippe, Jablonka, Kevin Maik
Large language models (LLMs) have gained widespread interest due to their ability to process human language and perform tasks on which they have not been explicitly trained. This is relevant for the chemical sciences, which face the problem of small and diverse datasets that are frequently in the form of text. LLMs have shown promise in addressing these issues and are increasingly being harnessed to predict chemical properties, optimize reactions, and even design and conduct experiments autonomously. However, we still have only a very limited systematic understanding of the chemical reasoning capabilities of LLMs, which would be required to improve models and mitigate potential harms. Here, we introduce "ChemBench," an automated framework designed to rigorously evaluate the chemical knowledge and reasoning abilities of state-of-the-art LLMs against the expertise of human chemists. We curated more than 7,000 question-answer pairs for a wide array of subfields of the chemical sciences, evaluated leading open and closed-source LLMs, and found that the best models outperformed the best human chemists in our study on average. The models, however, struggle with some chemical reasoning tasks that are easy for human experts and provide overconfident, misleading predictions, such as about chemicals' safety profiles. These findings underscore the dual reality that, although LLMs demonstrate remarkable proficiency in chemical tasks, further research is critical to enhancing their safety and utility in chemical sciences. Our findings also indicate a need for adaptations to chemistry curricula and highlight the importance of continuing to develop evaluation frameworks to improve safe and useful LLMs.
Unveiling Divergent Inductive Biases of LLMs on Temporal Data
Unraveling the intricate details of events in natural language necessitates a subtle understanding of temporal dynamics. Despite the adeptness of Large Language Models (LLMs) in discerning patterns and relationships from data, their inherent comprehension of temporal dynamics remains a formidable challenge. This research meticulously explores these intrinsic challenges within LLMs, with a specific emphasis on evaluating the performance of GPT-3.5 and GPT-4 models in the analysis of temporal data. Employing two distinct prompt types, namely Question Answering (QA) format and Textual Entailment (TE) format, our analysis probes into both implicit and explicit events. The findings underscore noteworthy trends, revealing disparities in the performance of GPT-3.5 and GPT-4. Notably, biases toward specific temporal relationships come to light, with GPT-3.5 demonstrating a preference for "AFTER'' in the QA format for both implicit and explicit events, while GPT-4 leans towards "BEFORE''. Furthermore, a consistent pattern surfaces wherein GPT-3.5 tends towards "TRUE'', and GPT-4 exhibits a preference for "FALSE'' in the TE format for both implicit and explicit events. This persistent discrepancy between GPT-3.5 and GPT-4 in handling temporal data highlights the intricate nature of inductive bias in LLMs, suggesting that the evolution of these models may not merely mitigate bias but may introduce new layers of complexity.
TS-CausalNN: Learning Temporal Causal Relations from Non-linear Non-stationary Time Series Data
Faruque, Omar, Ali, Sahara, Zheng, Xue, Wang, Jianwu
The growing availability and importance of time series data across various domains, including environmental science, epidemiology, and economics, has led to an increasing need for time-series causal discovery methods that can identify the intricate relationships in the non-stationary, non-linear, and often noisy real world data. However, the majority of current time series causal discovery methods assume stationarity and linear relations in data, making them infeasible for the task. Further, the recent deep learning-based methods rely on the traditional causal structure learning approaches making them computationally expensive. In this paper, we propose a Time-Series Causal Neural Network (TS-CausalNN) - a deep learning technique to discover contemporaneous and lagged causal relations simultaneously. Our proposed architecture comprises (i) convolutional blocks comprising parallel custom causal layers, (ii) acyclicity constraint, and (iii) optimization techniques using the augmented Lagrangian approach. In addition to the simple parallel design, an advantage of the proposed model is that it naturally handles the non-stationarity and non-linearity of the data. Through experiments on multiple synthetic and real world datasets, we demonstrate the empirical proficiency of our proposed approach as compared to several state-of-the-art methods. The inferred graphs for the real world dataset are in good agreement with the domain understanding.
mChartQA: A universal benchmark for multimodal Chart Question Answer based on Vision-Language Alignment and Reasoning
Wei, Jingxuan, Xu, Nan, Chang, Guiyong, Luo, Yin, Yu, BiHui, Guo, Ruifeng
The goal of multimodal chart question answering is to automatically answer a natural language question about a chart to facilitate visual data analysis (Hoque et al., 2022), where the ability to understand and interact with visual data is essential (Masry et al., 2022). It has emerged as a crucial intersection of computer vision and natural language processing, addressing the growing demand for intelligent systems capable of interpreting complex visual data in charts (Masry et al., 2022). Beyond its general applications, multimodal chart question-answering plays a pivotal role in sectors requiring precise and rapid analysis of visual data. In the financial domain, it is indispensable for tasks such as financial report analysis (Wang et al., 2023a), decision support (Kafle et al., 2020), invoice parsing (Gerling and Lessmann, 2023), and contract review (Jie et al., 2023). Similarly, in the medical field, it significantly contributes to the digitization of patient records (Xu et al., 2021), medical insurance review (Meskó, 2023), diagnostic assistance (Othmani and Zeghina, 2022), and quality control (Schilcher et al., 2024) of medical records. Due to the richness and ambiguities of natural language and complex visual reasoning, multimodal chart question answering task requires to predict the answer in the intersection of information visualization, natural language processing, and human computer interactions (Hoque et al., 2022). Early approaches apply natural language processing techniques by largely depending on heuristics or grammarbased parsing techniques (Setlur et al., 2016; Srinivasan and Stasko, 2017; Hoque et al., 2017; Gao et al., 2015). Thanks to insufficient processing of complex linguistic phenomena, over-reliance on grammatical rules, and limited depth of understanding natural language, deep learning models have been introduced for understanding natural language queries about visualizations (Chaudhry et al., 2020; Singh and Shekhar, 2020; Reddy et al., 2019).