Goto

Collaborating Authors

 Personal


AI firms 'should include members of public on boards to protect society'

The Guardian

Companies developing powerful artificial intelligence systems must have independent board members representing the "interests of society", according to an expert regarded as one of the modern godfathers of the technology. Yoshua Bengio, a co-winner of the 2018 Turing Award – referred to as the "Nobel prize of computing" – said AI firms must have oversight from members of the public, as advances in the technology accelerate rapidly. Speaking in the wake of the boardroom upheaval at the ChatGPT developer OpenAI, including the exit and return of its chief executive, Sam Altman, Bengio said a "democratic process" was needed to monitor developments in the field. "How do we make sure that these advances are happening in a way that doesn't endanger the public? How do we make sure that they're not abused for increasing one's power?" the AI pioneer told the Guardian. "To me, the answer is obvious in principle.


ZeroNLG: Aligning and Autoencoding Domains for Zero-Shot Multimodal and Multilingual Natural Language Generation

arXiv.org Artificial Intelligence

Natural Language Generation (NLG) accepts input data in the form of images, videos, or text and generates corresponding natural language text as output. Existing NLG methods mainly adopt a supervised approach and rely heavily on coupled data-to-text pairs. However, for many targeted scenarios and for non-English languages, sufficient quantities of labeled data are often not available. To relax the dependency on labeled data of downstream tasks, we propose an intuitive and effective zero-shot learning framework, ZeroNLG, which can deal with multiple NLG tasks, including image-to-text (image captioning), video-to-text (video captioning), and text-to-text (neural machine translation), across English, Chinese, German, and French within a unified framework. ZeroNLG does not require any labeled downstream pairs for training. During training, ZeroNLG (i) projects different domains (across modalities and languages) to corresponding coordinates in a shared common latent space; (ii) bridges different domains by aligning their corresponding coordinates in this space; and (iii) builds an unsupervised multilingual auto-encoder to learn to generate text by reconstructing the input text given its coordinate in shared latent space. Consequently, during inference, based on the data-to-text pipeline, ZeroNLG can generate target sentences across different languages given the coordinate of input data in the common space. Within this unified framework, given visual (imaging or video) data as input, ZeroNLG can perform zero-shot visual captioning; given textual sentences as input, ZeroNLG can perform zero-shot machine translation. We present the results of extensive experiments on twelve NLG tasks, showing that, without using any labeled downstream pairs for training, ZeroNLG generates high-quality and believable outputs and significantly outperforms existing zero-shot methods.


Improving Activation Steering in Language Models with Mean-Centring

arXiv.org Artificial Intelligence

Recent work in activation steering has demonstrated the potential to better control the outputs of Large Language Models (LLMs), but it involves finding steering vectors. This is difficult because engineers do not typically know how features are represented in these models. We seek to address this issue by applying the idea of mean-centring to steering vectors. We find that taking the average of activations associated with a target dataset, and then subtracting the mean of all training activations, results in effective steering vectors. We test this method on a variety of models on natural language tasks by steering away from generating toxic text, and steering the completion of a story towards a target genre. We also apply mean-centring to extract function vectors, more effectively triggering the execution of a range of natural language tasks by a significant margin (compared to previous baselines). This suggests that mean-centring can be used to easily improve the effectiveness of activation steering in a wide range of contexts.


OneLLM: One Framework to Align All Modalities with Language

arXiv.org Artificial Intelligence

Multimodal large language models (MLLMs) have gained significant attention due to their strong multimodal understanding capability. However, existing works rely heavily on modality-specific encoders, which usually differ in architecture and are limited to common modalities. In this paper, we present OneLLM, an MLLM that aligns eight modalities to language using a unified framework. We achieve this through a unified multimodal encoder and a progressive multimodal alignment pipeline. In detail, we first train an image projection module to connect a vision encoder with LLM. Then, we build a universal projection module (UPM) by mixing multiple image projection modules and dynamic routing. Finally, we progressively align more modalities to LLM with the UPM. To fully leverage the potential of OneLLM in following instructions, we also curated a comprehensive multimodal instruction dataset, including 2M items from image, audio, video, point cloud, depth/normal map, IMU and fMRI brain activity. OneLLM is evaluated on 25 diverse benchmarks, encompassing tasks such as multimodal captioning, question answering and reasoning, where it delivers excellent performance. Code, data, model and online demo are available at https://github.com/csuhan/OneLLM


Completeness, Recall, and Negation in Open-World Knowledge Bases: A Survey

arXiv.org Artificial Intelligence

General-purpose knowledge bases (KBs) are a cornerstone of knowledge-centric AI. Many of them are constructed pragmatically from Web sources, and are thus far from complete. This poses challenges for the consumption as well as the curation of their content. While several surveys target the problem of completing incomplete KBs, the first problem is arguably to know whether and where the KB is incomplete in the first place, and to which degree. In this survey we discuss how knowledge about completeness, recall, and negation in KBs can be expressed, extracted, and inferred. We cover (i) the logical foundations of knowledge representation and querying under partial closed-world semantics; (ii) the estimation of this information via statistical patterns; (iii) the extraction of information about recall from KBs and text; (iv) the identification of interesting negative statements; and (v) relaxed notions of relative recall. This survey is targeted at two types of audiences: (1) practitioners who are interested in tracking KB quality, focusing extraction efforts, and building quality-aware downstream applications; and (2) data management, knowledge base and semantic web researchers who wish to understand the state of the art of knowledge bases beyond the open-world assumption. Consequently, our survey presents both fundamental methodologies and their working, and gives practice-oriented recommendations on how to choose between different approaches for a problem at hand.


Let AI Jimmy Stewart put you to sleep with a new Calm bedtime story

Engadget

Jimmy Stewart can now send you off to a blissful night's rest with a Calm bedtime story. The mindfulness app is known for its Sleep Stories, read by celebrities including Harry Styles and Idris Elba, to help users drift off to dreamland. To revive Stewart's iconic voice Calm has collaborated with AI company Respeecher. The new It's a Wonderful Sleep Story, which Calm has dubbed "a heartwarming new holiday tale," is now available for Premium subscribers. Stewart starred in several major films (including It's a Wonderful Life) and was known for his signature drawl and calming voice.


Invariant Descriptors of Motion and Force Trajectories for Interpreting Object Manipulation Tasks in Contact

arXiv.org Artificial Intelligence

Invariant descriptors of point and rigid-body motion trajectories have been proposed in the past as representative task models for motion recognition and generalization. Currently, no invariant descriptor exists for representing force trajectories, which appear in contact tasks. This paper introduces invariant descriptors for force trajectories by exploiting the duality between motion and force. Two types of invariant descriptors are presented depending on whether the trajectories consist of screw or vector coordinates. Methods and software are provided for robustly calculating the invariant descriptors from noisy measurements using optimal control. Using experimental human demonstrations of 3D contour following and peg-on-hole alignment tasks, invariant descriptors are shown to result in task representations that do not depend on the calibration of reference frames or sensor locations. The tuning process for the optimal control problems is shown to be fast and intuitive. Similar to motions in free space, the proposed invariant descriptors for motion and force trajectories may prove useful for the recognition and generalization of constrained motions, such as during object manipulation in contact.


Privacy-Aware Data Acquisition under Data Similarity in Regression Markets

arXiv.org Artificial Intelligence

Data markets facilitate decentralized data exchange for applications such as prediction, learning, or inference. The design of these markets is challenged by varying privacy preferences as well as data similarity among data owners. Related works have often overlooked how data similarity impacts pricing and data value through statistical information leakage. We demonstrate that data similarity and privacy preferences are integral to market design and propose a query-response protocol using local differential privacy for a two-party data acquisition mechanism. In our regression data market model, we analyze strategic interactions between privacy-aware owners and the learner as a Stackelberg game over the asked price and privacy factor. Finally, we numerically evaluate how data similarity affects market participation and traded data value. A. Context and Motivation In recent years, there has been a surge in Internet of Things (IoT) devices with sensing and computing capabilities, leading to an abundance of IoT data. Shashi Raj Pandey and Petar Popovski are with the Connectivity Section, Department of Electronic Systems, Aalborg University, Denmark. Pierre Pinson has primary affiliation with Dyson School of Design Engineering, Imperial College London, UK. He is also affiliated to the Technical University of Denmark, Department of Technology, Management and Economics, as well as with Halfspace This work was supported by the Villum Investigator Grant "WATER" from the Velux Foundation, Denmark.


A Dating App Tried to Update Its Interface. Unbridled, Horny Chaos Ensued.

Slate

When Aaron* logged on to the kinky, nonmonogamy-focused dating app Feeld on Thursday to finalize plans with a match, the interface wouldn't load. As a middle-aged man in an ethically nonmonogamous relationship, Aaron considers Feeld a great way to meet other like-minded people in his area--and that's exactly what he was hoping to do this past Friday. Someone he had a connection with was in town for one night only, and he wanted to take advantage. He tried logging in again and changing his Wi-Fi connection, but nothing seemed to do the trick. Flummoxed, he took to X, the platform formerly known as Twitter, to see if there was any explanation.


As the last vanguards of the Greatest Generation pass, 7 things to know when caring for a parent

FOX News

Fox News' Martha MacCallum has the latest on her new Fox Nation documentary on'The Story.' My father-in-law passed away last month, days away from his 99th birthday. He lived with us for 13 years. He was a great man, a World War II veteran who loved his wife and raised three children. As his vascular dementia worsened – unlike Alzheimer's, his long-term memory remained intact almost until the end – my wife would set him up with a familiar film. "The Godfather" played most frequently, followed by "Patton."