Goto

Collaborating Authors

 Government


Probabilistic Inference in Language Models via Twisted Sequential Monte Carlo

arXiv.org Machine Learning

Numerous capability and safety techniques of Large Language Models (LLMs), including RLHF, automated red-teaming, prompt engineering, and infilling, can be cast as sampling from an unnormalized target distribution defined by a given reward or potential function over the full sequence. In this work, we leverage the rich toolkit of Sequential Monte Carlo (SMC) for these probabilistic inference problems. In particular, we use learned twist functions to estimate the expected future value of the potential at each timestep, which enables us to focus inference-time computation on promising partial sequences. We propose a novel contrastive method for learning the twist functions, and establish connections with the rich literature of soft reinforcement learning. As a complementary application of our twisted SMC framework, we present methods for evaluating the accuracy of language model inference techniques using novel bidirectional SMC bounds on the log partition function. These bounds can be used to estimate the KL divergence between the inference and target distributions in both directions. We apply our inference evaluation techniques to show that twisted SMC is effective for sampling undesirable outputs from a pretrained model (a useful component of harmlessness training and automated red-teaming), generating reviews with varied sentiment, and performing infilling tasks.


Belarus says it thwarted attempted Lithuanian drone strikes; Vilnius rebuffs claims

FOX News

Fox News Flash top headlines are here. Check out what's clicking on Foxnews.com. A top security official in Belarus claimed Thursday that the country has prevented attempted drone strikes from Lithuania targeting the Belarusian capital and surrounding areas. He did not present evidence for the claim or give any details. He also said that "radicals" in Lithuania and Poland are producing drones to attack Belarus.


Alphabet hails 'once-in-a-generation' AI opportunity as revenue rises

The Guardian

Shares in Alphabet, the owner of Google and YouTube, surged after it issued its first ever dividend and revealed that profits had surged in the last quarter. Sundar Pichai, CEO, hailed the transition to artificial intelligence as a "once-in-a-generation opportunity" as his company races to integrate the technology across its business. Investors cheered the firm's earnings, and news of a 70bn stock buyback. Google posted 80.5bn in revenue for the first quarter of 2024, up 15% on the same period last year, and reported 1.89 in earnings per share, up from 1.17 – surpassing analysts' expectations on both counts. Shares in Alphabet were up roughly 15% in after-hours trading.


UFO or drone? 'Flying cylinder' spotted soaring over New York City's LaGuardia Airport baffles passenger

Daily Mail - Science & tech

A woman has claimed that she witnessed a possible UFO while flying in a passenger airplane over New York City. Michelle Reyes shared the video online, which she capture from the window seat, showing a'flying cylinder' whizz by as she traveled over LaGuardia Airport. She told NewsNation that she observed the black object moving at high speeds - much faster than the airplane - and that another passenger had also witnessed it. A UFO expert analyzed the clip, determining no evidence that the video was fake or a hoax - but some have suggested the object was a drone. Michelle Reyes spoke NewsMax's Ashleigh Banfield about the mysterious object she spotted while flying over New York City'The first thing I did was email the FAA to let them know what I saw,' Reyes told NewsMax's Ashleigh Banfield, noting she has yet to receive a response.


The world's biggest 3D printer can a make a house in under 80 hours

Engadget

The University of Maine just unveiled the world's largest polymer 3D printer. The new printer, named Factory of the Future 1.0 (FoF 1.0), can print objects as large as 96 feet long by 32 feet wide by 18 feet high. It's also quite speedy, relatively speaking, as it can print up to 500 pounds per hour. It can dynamically switch between printing techniques to suit different aspects of complex jobs. The printer can flip between large-scale additive manufacturing, subtractive manufacturing, continuous tape layup and robot arm operations.


Tech CEOs Say Ethical A.I. and Innovation Are 'Two Sides of the Same Coin'

TIME - Tech

CEOs of start-ups and big tech companies spoke at the TIME100 Summit on Wednesday about innovating with artificial intelligence in an ethical way, just moments before a spirited debate on the future of the technology. "Regulation and innovation are two sides of the same coin," said Rosanne Kincaid-Smith, Group Chief Operating Officer of Northern Data Group, which is a signature partner of the TIME100 Summit. She added tech companies and industry leaders should work towards better regulation. "Not actively contributing through lobbying would be a huge miss for us," she said. Kincaid-Smith stressed the benefits of AI and suggested that questions about whether AI is "evil" and going to negatively impact the workforce are misguided.


Evaluating Collaborative Autonomy in Opposed Environments using Maritime Capture-the-Flag Competitions

arXiv.org Artificial Intelligence

The objective of this work is to evaluate multi-agent artificial intelligence methods when deployed on teams of unmanned surface vehicles (USV) in an adversarial environment. Autonomous agents were evaluated in real-world scenarios using the Aquaticus test-bed, which is a Capture-the-Flag (CTF) style competition involving teams of USV systems. Cooperative teaming algorithms of various foundations in behavior-based optimization and deep reinforcement learning (RL) were deployed on these USV systems in two versus two teams and tested against each other during a competition period in the fall of 2023. Deep reinforcement learning applied to USV agents was achieved via the Pyquaticus test bed, a lightweight gymnasium environment that allows simulated CTF training in a low-level environment. The results of the experiment demonstrate that rule-based cooperation for behavior-based agents outperformed those trained in Deep-reinforcement learning paradigms as implemented in these competitions. Further integration of the Pyquaticus gymnasium environment for RL with MOOS-IvP in terms of configuration and control schema will allow for more competitive CTF games in future studies. As the development of experimental deep RL methods continues, the authors expect that the competitive gap between behavior-based autonomy and deep RL will be reduced. As such, this report outlines the overall competition, methods, and results with an emphasis on future works such as reward shaping and sim-to-real methodologies and extending rule-based cooperation among agents to react to safety and security events in accordance with human experts intent/rules for executing safety and security processes.


Fiper: a Visual-based Explanation Combining Rules and Feature Importance

arXiv.org Artificial Intelligence

Artificial Intelligence algorithms have now become pervasive in multiple high-stakes domains. However, their internal logic can be obscure to humans. Explainable Artificial Intelligence aims to design tools and techniques to illustrate the predictions of the so-called black-box algorithms. The Human-Computer Interaction community has long stressed the need for a more user-centered approach to Explainable AI. This approach can benefit from research in user interface, user experience, and visual analytics. This paper proposes a visual-based method to illustrate rules paired with feature importance. A user study with 15 participants was conducted comparing our visual method with the original output of the algorithm and textual representation to test its effectiveness with users.


Unleashing the Potential of Fractional Calculus in Graph Neural Networks with FROND

arXiv.org Artificial Intelligence

We introduce the FRactional-Order graph Neural Dynamical network (FROND), a new continuous graph neural network (GNN) framework. Unlike traditional continuous GNNs that rely on integer-order differential equations, FROND employs the Caputo fractional derivative to leverage the non-local properties of fractional calculus. This approach enables the capture of long-term dependencies in feature updates, moving beyond the Markovian update mechanisms in conventional integer-order models and offering enhanced capabilities in graph representation learning. We offer an interpretation of the node feature updating process in FROND from a non-Markovian random walk perspective when the feature updating is particularly governed by a diffusion process. We demonstrate analytically that oversmoothing can be mitigated in this setting. Experimentally, we validate the FROND framework by comparing the fractional adaptations of various established integer-order continuous GNNs, demonstrating their consistently improved performance and underscoring the framework's potential as an effective extension to enhance traditional continuous GNNs. The code is available at \url{https://github.com/zknus/ICLR2024-FROND}.


CLARE: Cognitive Load Assessment in REaltime with Multimodal Data

arXiv.org Artificial Intelligence

We present a novel multimodal dataset for Cognitive Load Assessment in REaltime (CLARE). The dataset contains physiological and gaze data from 24 participants with self-reported cognitive load scores as ground-truth labels. The dataset consists of four modalities, namely, Electrocardiography (ECG), Electrodermal Activity (EDA), Electroencephalogram (EEG), and Gaze tracking. To map diverse levels of mental load on participants during experiments, each participant completed four nine-minutes sessions on a computer-based operator performance and mental workload task (the MATB-II software) with varying levels of complexity in one minute segments. During the experiment, participants reported their cognitive load every 10 seconds. For the dataset, we also provide benchmark binary classification results with machine learning and deep learning models on two different evaluation schemes, namely, 10-fold and leave-one-subject-out (LOSO) cross-validation. Benchmark results show that for 10-fold evaluation, the convolutional neural network (CNN) based deep learning model achieves the best classification performance with ECG, EDA, and Gaze. In contrast, for LOSO, the best performance is achieved by the deep learning model with ECG, EDA, and EEG.