Government
PAC Neural Prediction Set Learning to Quantify the Uncertainty of Generative Language Models
Uncertainty learning and quantification of models are crucial tasks to enhance the trustworthiness of the models. Importantly, the recent surge of generative language models (GLMs) emphasizes the need for reliable uncertainty quantification due to the concerns on generating hallucinated facts. In this paper, we propose to learn neural prediction set models that comes with the probably approximately correct (PAC) guarantee for quantifying the uncertainty of GLMs. Unlike existing prediction set models, which are parameterized by a scalar value, we propose to parameterize prediction sets via neural networks, which achieves more precise uncertainty quantification but still satisfies the PAC guarantee. We demonstrate the efficacy of our method on four types of language datasets and six types of models by showing that our method improves the quantified uncertainty by $63\%$ on average, compared to a standard baseline method.
Exploring acceptance of autonomous vehicle policies using KeyBERT and SNA: Targeting engineering students
This study aims to explore user acceptance of Autonomous Vehicle (AV) policies with improved text-mining methods. Recently, South Korean policymakers have viewed Autonomous Driving Car (ADC) and Autonomous Driving Robot (ADR) as next-generation means of transportation that will reduce the cost of transporting passengers and goods. They support the construction of V2I and V2V communication infrastructures for ADC and recognize that ADR is equivalent to pedestrians to promote its deployment into sidewalks. To fill the gap where end-user acceptance of these policies is not well considered, this study applied two text-mining methods to the comments of graduate students in the fields of Industrial, Mechanical, and Electronics-Electrical-Computer. One is the Co-occurrence Network Analysis (CNA) based on TF-IWF and Dice coefficient, and the other is the Contextual Semantic Network Analysis (C-SNA) based on both KeyBERT, which extracts keywords that contextually represent the comments, and double cosine similarity. The reason for comparing these approaches is to balance interest not only in the implications for the AV policies but also in the need to apply quality text mining to this research domain. Significantly, the limitation of frequency-based text mining, which does not reflect textual context, and the trade-off of adjusting thresholds in Semantic Network Analysis (SNA) were considered. As the results of comparing the two approaches, the C-SNA provided the information necessary to understand users' voices using fewer nodes and features than the CNA. The users who pre-emptively understood the AV policies based on their engineering literacy and the given texts revealed potential risks of the AV accident policies. This study adds suggestions to manage these risks to support the successful deployment of AVs on public roads.
Machine Learning Enhanced Hankel Dynamic-Mode Decomposition
Curtis, Christopher W., Alford-Lago, D. Jay, Bollt, Erik, Tuma, Andrew
While the acquisition of time series has become more straightforward, developing dynamical models from time series is still a challenging and evolving problem domain. Within the last several years, to address this problem, there has been a merging of machine learning tools with what is called the dynamic mode decomposition (DMD). This general approach has been shown to be an especially promising avenue for accurate model development. Building on this prior body of work, we develop a deep learning DMD based method which makes use of the fundamental insight of Takens' Embedding Theorem to build an adaptive learning scheme that better approximates higher dimensional and chaotic dynamics. We call this method the Deep Learning Hankel DMD (DLHDMD). We likewise explore how our method learns mappings which tend, after successful training, to significantly change the mutual information between dimensions in the dynamics. This appears to be a key feature in enhancing the DMD overall, and it should help provide further insight for developing other deep learning methods for time series analysis and model generation. This work uses machine learning to develop an accurate method for generating models of chaotic dynamical systems using measurements alone.
A Human Word Association based model for topic detection in social networks
Khadivi, Mehrdad Ranjbar, Akbarpour, Shahin, Feizi-Derakhshi, Mohammad-Reza, Anari, Babak
With the widespread use of social networks, detecting the topics discussed in these networks has become a significant challenge. The current works are mainly based on frequent pattern mining or semantic relations, and the language structure is not considered. The meaning of language structural methods is to discover the relationship between words and how humans understand them. Therefore, this paper uses the Concept of the Imitation of the Mental Ability of Word Association to propose a topic detection framework in social networks. This framework is based on the Human Word Association method. A special extraction algorithm has also been designed for this purpose. The performance of this method is evaluated on the FA-CUP dataset. It is a benchmark dataset in the field of topic detection. The results show that the proposed method is a good improvement compared to other methods, based on the Topic-recall and the keyword F1 measure. Also, most of the previous works in the field of topic detection are limited to the English language, and the Persian language, especially microblogs written in this language, is considered a low-resource language. Therefore, a data set of Telegram posts in the Farsi language has been collected. Applying the proposed method to this dataset also shows that this method works better than other topic detection methods.
On the Interpretability and Significance of Bias Metrics in Texts: a PMI-based Approach
Valentini, Francisco, Rosati, Germรกn, Blasi, Damiรกn, Slezak, Diego Fernandez, Altszyler, Edgar
In recent years, word embeddings have been widely used to measure biases in texts. Even if they have proven to be effective in detecting a wide variety of biases, metrics based on word embeddings lack transparency and interpretability. We analyze an alternative PMI-based metric to quantify biases in texts. It can be expressed as a function of conditional probabilities, which provides a simple interpretation in terms of word co-occurrences. We also prove that it can be approximated by an odds ratio, which allows estimating confidence intervals and statistical significance of textual biases. This approach produces similar results to metrics based on word embeddings when capturing gender gaps of the real world embedded in large corpora.
NASA's James Webb telescope catches glimpse of possible 'dark stars' for the first time - which could solve one of the universe's biggest mysteries
NASA's James Webb Space Telescope has detected what were believed to be fabled'dark stars' that could solve one of the universe's biggest mysteries. A team of astronomers led by The University of Texas (UT) at Austin identified three potential'dark stars' that formed about 320 million years after the Big Bang, making them the earliest stars ever seen by human eyes. The image shows three fuzzy dots glowing in the blackness of space, but astronomers believe the tiny specs could lead to uncovering the elusive dark matter. Dark stars could only exist if dark matter creates heat at the core, preventing the stars from collapsing and causing them to puff up, which the team found in JWST's observations. Although dark matter makes up about 85 percent of the universe, its nature has eluded scientists.
The Chutzpah of the Self-Driving Car Company That Says "Humans Are Terrible Drivers"
Traffic deaths have been tumbling across the rich world, with Japan and Norway among the countries recently reaching postwar lows. The notable outlier is the United States. American crash fatalities hit a 16-year high in 2021 before barely budging last year. An American is now two to five times more likely to die in a collision than citizens of peer nations. Those expressing concern about this trend include Transportation Secretary Pete Buttigieg (who has called it "a national crisis"), roadway safety advocates, and newspaper editorial pages.
How judges, not politicians, could dictate America's AI rules
If these cases prove successful, they could force OpenAI, Meta, Microsoft, and others to change the way AI is built, trained, and deployed so that it is more fair and equitable. They could also create new ways for artists, authors, and others to be compensated for having their work used as training data for AI models, through a system of licensing and royalties. The generative AI boom has revived American politicians' enthusiasm for passing AI-specific laws. However, we're unlikely to see any such legislation pass in the next year, given the split Congress and intense lobbying from tech companies, says Ben Winters, senior counsel at the Electronic Privacy Information Center. Even the most prominent attempt to create new AI rules, Senator Chuck Schumer's SAFE Innovation framework, does not include any specific policy proposals.
House Committee Targets U.C. Berkeley Program for China Ties
A congressional committee focused on national security threats from China said it had "grave concerns" about a research partnership between the University of California, Berkeley, and several Chinese entities, claiming that the collaboration's advanced research could help the Chinese government gain an economic, technological or military advantage. In a letter sent last week to Berkeley's president and chancellor, the House Select Committee on the Chinese Communist Party requested extensive information about the Tsinghua-Berkeley Shenzhen Institute, a collaboration set up in 2014 with China's prestigious Tsinghua University and the Chinese city of Shenzhen. The letter pointed to the institute's research into certain "dual-use technologies" that are employed by both civilian and military institutions, like advanced semiconductors and imaging technology used for mapping terrain or driving autonomous cars. The committee also questioned whether Berkeley had properly disclosed Chinese funding for the institute, and cited its collaborations with Chinese universities and companies that have been the subjects of sanctions by the United States in recent years, like the National University of Defense Technology, the telecom firm Huawei and the Chinese drone maker DJI.
House Committee Targets U.C. Berkeley Program for China Ties
A congressional committee focused on national security threats from China said it had "grave concerns" about a research partnership between the University of California, Berkeley, and several Chinese entities, claiming that the collaboration's advanced research could help the Chinese government gain an economic, technological or military advantage. In a letter sent last week to Berkeley's president and chancellor, the House Select Committee on the Chinese Communist Party requested extensive information about the Tsinghua-Berkeley Shenzhen Institute, a collaboration set up in 2014 with China's prestigious Tsinghua University and the Chinese city of Shenzhen. The letter pointed to the institute's research into certain "dual-use technologies" that are employed by both civilian and military institutions, like advanced semiconductors and imaging technology used for mapping terrain or driving autonomous cars. The committee also questioned whether Berkeley had properly disclosed Chinese funding for the institute, and cited its collaborations with Chinese universities and companies that have been the subjects of sanctions by the United States in recent years, like the National University of Defense Technology, the telecom firm Huawei and the Chinese drone maker DJI.