Goto

Collaborating Authors

 Pacific Ocean


Chinese Discourse Annotation Reference Manual

arXiv.org Artificial Intelligence

This document provides extensive guidelines and examples for Rhetorical Structure Theory (RST) annotation in Mandarin Chinese. The guideline is divided into three sections. We first introduce preprocessing steps to prepare data for RST annotation. Secondly, we discuss syntactic criteria to segment texts into Elementary Discourse Units (EDUs). Lastly, we provide examples to define and distinguish discourse relations in different genres. We hope that this reference manual can facilitate RST annotations in Chinese and accelerate the development of the RST framework across languages.


Deep Mind's AlphaZero Game Playing AI -- Reduces Compute Time, Cuts Costs & Saves Energy

#artificialintelligence

CIOs & CTOs, the MIT Technology Review reported on October 5, 2022 reported that "DeepMind's game-playing AI AlphaZero has beaten a 50-year-old record in computer science. DeepMind has used its board-game playing AI to discover a faster way to solve a fundamental math problem in computer science, beating a record that has stood for more than 50 years. The problem, matrix multiplication, is a crucial type of calculation at the heart of many different applications, from displaying images on a screen to simulating complex physics. It is also fundamental to machine learning itself. Speeding up this calculation could have a big impact on thousands of everyday computer tasks, cutting costs and saving energy."


Learning to Reconstruct Missing Data from Spatiotemporal Graphs with Sparse Observations

arXiv.org Artificial Intelligence

Modeling multivariate time series as temporal signals over a (possibly dynamic) graph is an effective representational framework that allows for developing models for time series analysis. In fact, discrete sequences of graphs can be processed by autoregressive graph neural networks to recursively learn representations at each discrete point in time and space. Spatiotemporal graphs are often highly sparse, with time series characterized by multiple, concurrent, and long sequences of missing data, e.g., due to the unreliable underlying sensor network. In this context, autoregressive models can be brittle and exhibit unstable learning dynamics. The objective of this paper is, then, to tackle the problem of learning effective models to reconstruct, i.e., impute, missing data points by conditioning the reconstruction only on the available observations. In particular, we propose a novel class of attention-based architectures that, given a set of highly sparse discrete observations, learn a representation for points in time and space by exploiting a spatiotemporal propagation architecture aligned with the imputation task. Representations are trained end-to-end to reconstruct observations w.r.t. the corresponding sensor and its neighboring nodes. Compared to the state of the art, our model handles sparse data without propagating prediction errors or requiring a bidirectional model to encode forward and backward time dependencies. Empirical results on representative benchmarks show the effectiveness of the proposed method.


PaLM: Scaling Language Modeling with Pathways

arXiv.org Artificial Intelligence

Large language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed to adapt the model to a particular application. To further our understanding of the impact of scale on few-shot learning, we trained a 540-billion parameter, densely activated, Transformer language model, which we call Pathways Language Model PaLM. We trained PaLM on 6144 TPU v4 chips using Pathways, a new ML system which enables highly efficient training across multiple TPU Pods. We demonstrate continued benefits of scaling by achieving state-of-the-art few-shot learning results on hundreds of language understanding and generation benchmarks. On a number of these tasks, PaLM 540B achieves breakthrough performance, outperforming the finetuned state-of-the-art on a suite of multi-step reasoning tasks, and outperforming average human performance on the recently released BIG-bench benchmark. A significant number of BIG-bench tasks showed discontinuous improvements from model scale, meaning that performance steeply increased as we scaled to our largest model. PaLM also has strong capabilities in multilingual tasks and source code generation, which we demonstrate on a wide array of benchmarks. We additionally provide a comprehensive analysis on bias and toxicity, and study the extent of training data memorization with respect to model scale. Finally, we discuss the ethical considerations related to large language models and discuss potential mitigation strategies.


Learning Spatially-Adaptive Squeeze-Excitation Networks for Image Synthesis and Image Recognition

arXiv.org Artificial Intelligence

Learning light-weight yet expressive deep networks in both image synthesis and image recognition remains a challenging problem. Inspired by a more recent observation that it is the data-specificity that makes the multi-head self-attention (MHSA) in the Transformer model so powerful, this paper proposes to extend the widely adopted light-weight Squeeze-Excitation (SE) module to be spatially-adaptive to reinforce its data specificity, as a convolutional alternative of the MHSA, while retaining the efficiency of SE and the inductive basis of convolution. It presents two designs of spatially-adaptive squeeze-excitation (SASE) modules for image synthesis and image recognition respectively. For image synthesis tasks, the proposed SASE is tested in both low-shot and one-shot learning tasks. It shows better performance than prior arts. For image recognition tasks, the proposed SASE is used as a drop-in replacement for convolution layers in ResNets and achieves much better accuracy than the vanilla ResNets, and slightly better than the MHSA counterparts such as the Swin-Transformer and Pyramid-Transformer in the ImageNet-1000 dataset, with significantly smaller models.


Celebrate over 20 years of AI/ML at Innovation Day

#artificialintelligence

Be our guest as we celebrate 20 years of AI/ML innovation on October 25, 2022, 9:00 AM โ€“ 10:30 AM PT. The first 1,500 people to register will receive $50 of AWS credits. Over the past 20 years, Amazon has delivered many world firsts for artificial intelligence (AI) and machine learning (ML). ML is an integral part of Amazon and is used for everything from applying personalization models at checkout, to forecasting the demand for products globally, to creating autonomous flight for Amazon Prime Air drones, to natural language processing (NLP) on Alexa. And the use of ML isn't slowing down anytime soon, because ML helps Amazon exceed customer expectations for convenience, cost, and delivery speed.


Cloud Classification with Unsupervised Deep Learning

arXiv.org Artificial Intelligence

We present a framework for cloud characterization that leverages modern unsupervised deep learning technologies. While previous neural network-based cloud classification models have used supervised learning methods, unsupervised learning allows us to avoid restricting the model to artificial categories based on historical cloud classification schemes and enables the discovery of novel, more detailed classifications. Our framework learns cloud features directly from radiance data produced by NASA's Moderate Resolution Imaging Spectroradiometer (MODIS) satellite instrument, deriving cloud characteristics from millions of images without relying on pre-defined cloud types during the training process. We present preliminary results showing that our method extracts physically relevant information from radiance data and produces meaningful cloud classes.


The grandfather of AI art, DALL-E, is now free for you to try

PCWorld

For months, the "first" AI art program, DALL-E, has been hidden behind a beta wall that has limited access. Now it's open to everyone to try out, with a generous amount of credits, to boot. Each signup adds 50 credits to your account, with each credit generating four 1024 1024 images from a single prompt from the OpenAI server. You'll get 15 new credits per month, though the credits do not roll over. OpenAI also has placed content limits on the type of images you can generate, forbidding violence, sexual acts (including nudity), politicians, and public figures.


Ships are turning whales into 'ocean roadkill'. This AI system is trying to stop it

#artificialintelligence

Fran was a celebrity whale โ€“ the most photographed humpback in the San Francisco Bay, with 277 recorded sightings since 2005. Last month, she was hit by a ship and killed. Her death marked a grim milestone: Fran was the fifth whale to be killed by a ship strike in the area this year, according to the Marine Mammal Center. Collisions with ships are one of the leading causes of death for endangered whales, who breed, eat and travel in deep channels in the same busy waters that cargo ships frequent. Whales that spend their lives near the surface โ€“ such as humpbacks and right whales โ€“ are especially at risk. One 2019 study likened their plight to those of land animals forced to criss-cross the highways that cut through their habitats.


EgoSpeed-Net: Forecasting Speed-Control in Driver Behavior from Egocentric Video Data

arXiv.org Artificial Intelligence

Speed-control forecasting, a challenging problem in driver behavior analysis, aims to predict the future actions of a driver in controlling vehicle speed such as braking or acceleration. In this paper, we try to address this challenge solely using egocentric video data, in contrast to the majority of works in the literature using either third-person view data or extra vehicle sensor data such as GPS, or both. To this end, we propose a novel graph convolutional network (GCN) based network, namely, EgoSpeed-Net. We are motivated by the fact that the position changes of objects over time can provide us very useful clues for forecasting the speed change in future. We first model the spatial relations among the objects from each class, frame by frame, using fully-connected graphs, on top of which GCNs are applied for feature extraction. Then we utilize a long short-term memory network to fuse such features per class over time into a vector, concatenate such vectors and forecast a speed-control action using a multilayer perceptron classifier. We conduct extensive experiments on the Honda Research Institute Driving Dataset and demonstrate the superior performance of EgoSpeed-Net.