Goto

Collaborating Authors

 Media


Can companies make decisions with AI?

#artificialintelligence

AI can play many roles in the technology stack of a modern enterprise. Its performance as a neutral, data-based, analytical advisor could allow businesses to use algorithms to predict whether a decision is the right one. AI-based decisions are part of an arsenal of tools leveraged by technology high performers. Businesses led by digitally savvy leaders, those who champion emerging technologies such as AI, outperform other like-sized businesses by 48% on valuation and revenue growth, according to one MIT research study. "The integration of traditional decisioning into AI is really just starting to hit its stride right now," said Rowan Curran, analyst at Forrester.


CorefDiffs: Co-referential and Differential Knowledge Flow in Document Grounded Conversations

arXiv.org Artificial Intelligence

Knowledge-grounded dialog systems need to incorporate smooth transitions among knowledge selected for generating responses, to ensure that dialog flows naturally. For document-grounded dialog systems, the inter- and intra-document knowledge relations can be used to model such conversational flows. We develop a novel Multi-Document Co-Referential Graph (Coref-MDG) to effectively capture the inter-document relationships based on commonsense and similarity and the intra-document co-referential structures of knowledge segments within the grounding documents. We propose CorefDiffs, a Co-referential and Differential flow management method, to linearize the static Coref-MDG into conversational sequence logic. CorefDiffs performs knowledge selection by accounting for contextual graph structures and the knowledge difference sequences. CorefDiffs significantly outperforms the state-of-the-art by 9.5\%, 7.4\%, and 8.2\% on three public benchmarks. This demonstrates that the effective modeling of co-reference and knowledge difference for dialog flows are critical for transitions in document-grounded conversation


Comprint: Image Forgery Detection and Localization using Compression Fingerprints

arXiv.org Artificial Intelligence

Manipulation tools that realistically edit images are widely available, making it easy for anyone to create and spread misinformation. In an attempt to fight fake news, forgery detection and localization methods were designed. However, existing methods struggle to accurately reveal manipulations found in images on the internet, i.e., in the wild. That is because the type of forgery is typically unknown, in addition to the tampering traces being damaged by recompression. This paper presents Comprint, a novel forgery detection and localization method based on the compression fingerprint or comprint. It is trained on pristine data only, providing generalization to detect different types of manipulation. Additionally, we propose a fusion of Comprint with the state-of-the-art Noiseprint, which utilizes a complementary camera model fingerprint. We carry out an extensive experimental analysis and demonstrate that Comprint has a high level of accuracy on five evaluation datasets that represent a wide range of manipulation types, mimicking in-the-wild circumstances. Most notably, the proposed fusion significantly outperforms state-of-the-art reference methods. As such, Comprint and the fusion Comprint+Noiseprint represent a promising forensics tool to analyze in-the-wild tampered images.


Pay Self-Attention to Audio-Visual Navigation

arXiv.org Artificial Intelligence

Audio-visual embodied navigation, as a hot research topic, aims training a robot to reach an audio target using egocentric visual (from the sensors mounted on the robot) and audio (emitted from the target) input. The audio-visual information fusion strategy is naturally important to the navigation performance, but the state-of-the-art methods still simply concatenate the visual and audio features, potentially ignoring the direct impact of context. Moreover, the existing approaches requires either phase-wise training or additional aid (e.g. topology graph and sound semantics). Up till this date, the work that deals with the more challenging setup with moving target(s) is still rare. As a result, we propose an end-to-end framework FSAAVN (feature self-attention audio-visual navigation) to learn chasing after a moving audio target using a context-aware audio-visual fusion strategy implemented as a self-attention module. Our thorough experiments validate the superior performance (both quantitatively and qualitatively) of FSAAVN in comparison with the state-of-the-arts, and also provide unique insights about the choice of visual modalities, visual/audio encoder backbones and fusion patterns.


TartanCalib: Iterative Wide-Angle Lens Calibration using Adaptive SubPixel Refinement of AprilTags

arXiv.org Artificial Intelligence

Wide-angle cameras are uniquely positioned for mobile robots, by virtue of the rich information they provide in a small, light, and cost-effective form factor. An accurate calibration of the intrinsics and extrinsics is a critical pre-requisite for using the edge of a wide-angle lens for depth perception and odometry. Calibrating wide-angle lenses with current state-of-the-art techniques yields poor results due to extreme distortion at the edge, as most algorithms assume a lens with low to medium distortion closer to a pinhole projection. In this work we present our methodology for accurate wide-angle calibration. Our pipeline generates an intermediate model, and leverages it to iteratively improve feature detection and eventually the camera parameters. We test three key methods to utilize intermediate camera models: (1) undistorting the image into virtual pinhole cameras, (2) reprojecting the target into the image frame, and (3) adaptive subpixel refinement. Combining adaptive subpixel refinement and feature reprojection significantly improves reprojection errors by up to 26.59 %, helps us detect up to 42.01 % more features, and improves performance in the downstream task of dense depth mapping. Finally, TartanCalib is open-source and implemented into an easy-to-use calibration toolbox. We also provide a translation layer with other state-of-the-art works, which allows for regressing generic models with thousands of parameters or using a more robust solver. To this end, TartanCalib is the tool of choice for wide-angle calibration. Project website and code: http://tartancalib.com.


Bias amplification in experimental social networks is reduced by resampling

arXiv.org Artificial Intelligence

Large-scale social networks are thought to contribute to polarization by amplifying people's biases. However, the complexity of these technologies makes it difficult to identify the mechanisms responsible and to evaluate mitigation strategies. Here we show under controlled laboratory conditions that information transmission through social networks amplifies motivational biases on a simple perceptual decision-making task. Participants in a large behavioral experiment showed increased rates of biased decision-making when part of a social network relative to asocial participants, across 40 independently evolving populations. Drawing on techniques from machine learning and Bayesian statistics, we identify a simple adjustment to content-selection algorithms that is predicted to mitigate bias amplification. This algorithm generates a sample of perspectives from within an individual's network that is more representative of the population as a whole. In a second large experiment, this strategy reduced bias amplification while maintaining the benefits of information sharing. For example, social networks often lead to "echo-chambers" of like-minded individuals In this paper, we use an experimental paradigm to study how information sharing affects bias in judgment and decision-making. This new experimental paradigm allowed us to evaluate a mathematical theory of bias amplification and test a mitigation strategy based on this theory. Participants received a monetary reward for every correct answer. However, certain participants were offered an additional monetary reward for every green or blue dot in each stimulus ("motivated color" was randomized across participants, Our experimental paradigm consisted of arranging participants into an ordered set of groups, called "waves". At each wave t participants in social conditions observed judgments made by the participants in wave t 1. Participants in asocial conditions did not observe any social information. Each colored circle at the top of the image represents a participant. Stimuli consisted of 100 randomly positioned and sized blue and green dots displayed for one second. After viewing a stimulus, participants indicated whether they thought the stimulus had more green or more blue dots. Participants received feedback after each judgment on practice trials, and at the end of the experiment on test trials. All participants received a bonus on every trial if their judgment was correct. Participants in motivated conditions (shown here) received an additional bonus on every trial for every dot of their motivated color (green in both plots) regardless of whether their judgment was correct.


Continual Meta-Reinforcement Learning for UAV-Aided Vehicular Wireless Networks

arXiv.org Artificial Intelligence

An important use case is offered by vehicular ground users are static and have known locations. The same wireless networks, in which UABSs serve as relays authors in [28] extended their previous work by considering between vehicular users and the network, enabling the users multiple UABSs. Unlike these previous works, in this paper, to upload data collected by on-board sensors [5]-[11]. Such we consider traffic conditions characterized by vehicular users user-generated data are collected by the network, and then with a priori unknown locations and we move beyond conventional forwarded to other vehicles by means of BSs or road side meta-RL by accounting for the constraint that simulators units (RSUs). Being able to offer stronger, possibly line-ofsight for previous traffic configurations cannot be revisited. The (LoS), links to vehicles as compared to (static) ground rest of the paper is organized as follows. The system model BSs, UABSs can support demanding vehicle-to-everything and the problem formulation are described in Section II.


PaLM: Scaling Language Modeling with Pathways

arXiv.org Artificial Intelligence

Large language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed to adapt the model to a particular application. To further our understanding of the impact of scale on few-shot learning, we trained a 540-billion parameter, densely activated, Transformer language model, which we call Pathways Language Model PaLM. We trained PaLM on 6144 TPU v4 chips using Pathways, a new ML system which enables highly efficient training across multiple TPU Pods. We demonstrate continued benefits of scaling by achieving state-of-the-art few-shot learning results on hundreds of language understanding and generation benchmarks. On a number of these tasks, PaLM 540B achieves breakthrough performance, outperforming the finetuned state-of-the-art on a suite of multi-step reasoning tasks, and outperforming average human performance on the recently released BIG-bench benchmark. A significant number of BIG-bench tasks showed discontinuous improvements from model scale, meaning that performance steeply increased as we scaled to our largest model. PaLM also has strong capabilities in multilingual tasks and source code generation, which we demonstrate on a wide array of benchmarks. We additionally provide a comprehensive analysis on bias and toxicity, and study the extent of training data memorization with respect to model scale. Finally, we discuss the ethical considerations related to large language models and discuss potential mitigation strategies.


Controllable Data Generation by Deep Learning: A Review

arXiv.org Artificial Intelligence

Designing and generating new data under targeted properties has been attracting various critical applications such as molecule design, image editing and speech synthesis. Traditional hand-crafted approaches heavily rely on expertise experience and intensive human efforts, yet still suffer from the insufficiency of scientific knowledge and low throughput to support effective and efficient data generation. Recently, the advancement of deep learning induces expressive methods that can learn the underlying representation and properties of data. Such capability provides new opportunities in figuring out the mutual relationship between the structural patterns and functional properties of the data and leveraging such relationship to generate structural data given the desired properties. This article provides a systematic review of this promising research area, commonly known as controllable deep data generation. Firstly, the potential challenges are raised and preliminaries are provided. Then the controllable deep data generation is formally defined, a taxonomy on various techniques is proposed and the evaluation metrics in this specific domain are summarized. After that, exciting applications of controllable deep data generation are introduced and existing works are experimentally analyzed and compared. Finally, the promising future directions of controllable deep data generation are highlighted and five potential challenges are identified.


Tesla AI Day: Optimus Bot Was Better Than Anyone Expected

#artificialintelligence

Today I bring you an analysis of Tesla's autonomous humanoid robot, Optimus, unveiled on Friday at AI day 2022. As I always try to do, this is a nuanced take that highlights the good -- and the bad. Last year's AI day was especially exciting because Musk revealed Tesla was working on Optimus, a robot intended to "eliminate dangerous, repetitive, and boring tasks," and capable of following orders expressed in natural language like, "pick up that bolt and attach it to the car with that wrench." Musk also promised they'd have a working prototype for 2022's AI day, which -- given the unrealistic deadline --, hyped some and reminded others of his tendency to overpromise and underdeliver. I was in the latter group. Before we dive into it, let me remind you that AI day is explicitly intended for recruiting purposes: The target audience isn't journalists or investors, but engineers. This means that most news outlets' reviews of the event will be limited in describing the implications and non-incisive in analyzing the shortcomings.