Goto

Collaborating Authors

 Oceania


The Foreseeable Future: Self-Supervised Learning to Predict Dynamic Scenes for Indoor Navigation

arXiv.org Artificial Intelligence

Abstract--We present a method for generating, predicting, and using Spatiotemporal Occupancy Grid Maps (SOGM), which embed future semantic information of real dynamic scenes. We present an auto-labeling process that creates SOGMs from noisy real navigation data. We use a 3D-2D feedforward architecture, trained to predict the future time steps of SOGMs, given 3D lidar frames as input. Our pipeline is entirely self-supervised, thus enabling lifelong learning for real robots. The network is composed of a 3D back-end that extracts rich features and enables the semantic segmentation of the lidar frames, and a 2D front-end that predicts the future information embedded in the SOGM representation, potentially capturing the complexities and uncertainties of real-world multi-agent, multi-future interactions. We also design a navigation system that uses these predicted SOGMs within planning, after they have been transformed into Spatiotemporal Risk Maps (SRMs). We verify our navigation system's abilities in simulation, validate it on a real robot, study SOGM predictions on real data in various circumstances, and Time is represented as a color, from red (now) to yellow (future). REDICTING the future has always fascinated humanity. In this paper, we provide a detailed curiosity for the unknown has never faded. But we tend to description of the collection of algorithms required for these forget that we already predict the future constantly in our daily various tasks, for a complete view of the overall approach, as lives, only it is for a short horizon. Walking in the street, illustrated in Figure 2. catching a falling object, or driving a car, all these actions Some of the algorithms we use have already been introduced require a certain level of anticipation. In the first one [1], we described can become quite good at predicting what might happen for how to automatically annotate 3D lidar points, and train a the next few seconds in many situations; what about robots? In the second one We study this question in the context of a concrete example: [2], our system learned to predict the future of dynamic a robot learning on its own to navigate among humans or scenes as SOGMs. Until now, we only evaluated results in dynamic objects in an indoor space. Our approach allows the a simulated environment.


Nearest Neighbor Non-autoregressive Text Generation

arXiv.org Artificial Intelligence

Non-autoregressive (NAR) models can generate sentences with less computation than autoregressive models but sacrifice generation quality. Previous studies addressed this issue through iterative decoding. This study proposes using nearest neighbors as the initial state of an NAR decoder and editing them iteratively. We present a novel training strategy to learn the edit operations on neighbors to improve NAR text generation. Experimental results show that the proposed method (NeighborEdit) achieves higher translation quality (1.69 points higher than the vanilla Transformer) with fewer decoding iterations (one-eighteenth fewer iterations) on the JRC-Acquis En-De dataset, the common benchmark dataset for machine translation using nearest neighbors. We also confirm the effectiveness of the proposed method on a data-to-text task (WikiBio). In addition, the proposed method outperforms an NAR baseline on the WMT'14 En-De dataset. We also report analysis on neighbor examples used in the proposed method.


Towards Higher-order Topological Consistency for Unsupervised Network Alignment

arXiv.org Artificial Intelligence

--Network alignment task, which aims to identify corresponding nodes in different networks, is of great significance for many subsequent applications. Without the need for labeled anchor links, unsupervised alignment methods have been attracting more and more attention. However, the topological consistency assumptions defined by existing methods are generally low-order and less accurate because only the edge-indiscriminative topological pattern is considered, which is especially risky in an unsupervised setting. T o reposition the focus of the alignment process from low-order to higher-order topological consistency, in this paper, we propose a fully unsupervised network alignment framework named HTC. The proposed higher-order topological consistency is formulated based on edge orbits, which is merged into the information aggregation process of a graph convolutional network so that the alignment consistencies are transformed into the similarity of node embeddings. Furthermore, the encoder is trained to be multi-orbit-aware and then be refined to identify more trusted anchor links. Node correspondence is comprehensively evaluated by integrating all different orders of consistency. In addition to sound theoretical analysis, the superiority of the proposed method is also empirically demonstrated through extensive experimental evaluation. On three pairs of real-world datasets and two pairs of synthetic datasets, our HTC consistently outperforms a wide variety of unsupervised and supervised methods with the least or comparable time consumption. It also exhibits robustness to structural noise as a result of our multi-orbit-aware training mechanism. Network alignment task, which aims to identify entity correspondence across different networks, is usually the very first step of many downstream analyzing tasks. For instance, recognizing the same user on different social networks can facilitate friend suggestion, item recommendation, personalized advertisement [1]-[5]. Similar scenarios also exist widely in other fields, such as protein network analysis [6], knowledge discovery [7], etc. Identifying corresponding nodes across different networks is an extremely hard task, even for humans. Manually labelling correspondence can be prohibitively challenging, expensive (in human efforts, time, and money costs), and tedious [8]. Due to such obstacles, in some cases, it may be impractical to get access to sufficient labels for training well-performed supervised or even semi-supervised models [4], [9]. By contrast, unsupervised models can be trained without the need for labeled data, which is more flexible and practical in real-world application scenarios. Thus, unsupervised alignment methods have been drawing a surge of interest recently [10]-[12].


Cross-lingual Transfer Learning for Fake News Detector in a Low-Resource Language

arXiv.org Artificial Intelligence

Development of methods to detect fake news (FN) in low-resource languages has been impeded by a lack of training data. In this study, we solve the problem by using only training data from a high-resource language. Our FN-detection system permitted this strategy by applying adversarial learning that transfers the detection knowledge through languages. To assist the knowledge transfer, our system judges the reliability of articles by exploiting source information, which is a cross-lingual feature that represents the credibility of the speaker. In experiments, our system got 3.71% higher accuracy than a system that uses a machine-translated training dataset. In addition, our suggested cross-lingual feature exploitation for fake news detection improved accuracy by 3.03%.


FooDI-ML: a large multi-language dataset of food, drinks and groceries images and descriptions

arXiv.org Artificial Intelligence

In this paper we introduce the FooDI-ML dataset. This dataset contains over 1.5M unique images and over 9.5M store names, product names descriptions, and collection sections gathered from the Glovo application. The data made available corresponds to food, drinks and groceries products from 37 countries in Europe, the Middle East, Africa and Latin America. The dataset comprehends 33 languages, including 870K samples of languages of countries from Eastern Europe and Western Asia such as Ukrainian and Kazakh, which have been so far underrepresented in publicly available visio-linguistic datasets. The dataset also includes widely spoken languages such as Spanish and English. To assist further research, we include benchmarks over two tasks: text-image retrieval and conditional image generation.


Coalescing Global and Local Information for Procedural Text Understanding

arXiv.org Artificial Intelligence

Procedural text understanding is a challenging language reasoning task that requires models to track entity states across the development of a narrative. A complete procedural understanding solution should combine three core aspects: local and global views of the inputs, and global view of outputs. Prior methods considered a subset of these aspects, resulting in either low precision or low recall. In this paper, we propose Coalescing Global and Local Information (CGLI), a new model that builds entity- and timestep-aware input representations (local input) considering the whole context (global input), and we jointly model the entity states with a structured prediction objective (global output). Thus, CGLI simultaneously optimizes for both precision and recall. We extend CGLI with additional output layers and integrate it into a story reasoning framework. Extensive experiments on a popular procedural text understanding dataset show that our model achieves state-of-the-art results; experiments on a story reasoning benchmark show the positive impact of our model on downstream reasoning.


Engineering Manager, ML Infrastructure

#artificialintelligence

Affirm is reinventing credit to make it more honest and friendly, giving consumers the flexibility to buy now and pay later without any hidden fees or compounding interest. We are looking for an Engineering Manager to lead projects and initiatives on our new ML Developer Productivity team within the ML Platform group. In this role, you will lead and grow a hard-working distributed team to drive the vision and roadmap of the machine learning developer experience. Our mission is to build a self-service, easy-to-use foundation for developing and delivering robust models to production. This is a new team that will focus on creating internal tools used by ML Engineers for fast paced ML development.


Senior Software Engineer, Machine Learning (Credit Engineering)

#artificialintelligence

Affirm is reinventing credit to make it more honest and friendly, giving consumers the flexibility to buy now and pay later without any hidden fees or compounding interest. The International Machine Learning Underwriting (INML) team builds Affirm essential real-time underwriting models used outside of the US. These models determine which users are approved for an Affirm loan and how much they are approved for. These models are central to Affirm's continued success and growth, and need to be explainable to the user and scalable across regions. We offer a competitive package, with some highlights listed below.


After 25 years, we still don't see bicycle kicks at the RoboCup

Engadget

This year's RoboCup symposium held in Bangkok, Thailand marks the 25th anniversary of the event, an international competition dedicated to the advancement of robotic and artificial intelligence technologies. The original goal of the event was to get the state of robotics in robust enough shape that one might field a team of robotic soccer players capable of beating a World Cup champion (human) team by 2050 – but a lot has changed since 1997. Both the event and its mechanical entrants have evolved by leaps and bounds in the intervening years. The number of teams participating has ballooned tenfold since the inaugural event, from 38 to more than 300, with competitors now coming from more than 40 nations worldwide. And rather than fall down stairs, today's cutting-edge humanoid constructs are backflipping off them.


Development of Sleep State Trend (SST), a bedside measure of neonatal sleep state fluctuations based on single EEG channels

arXiv.org Machine Learning

Objective: To develop and validate an automated method for bedside monitoring of sleep state fluctuations in neonatal intensive care units. Methods: A deep learning -based algorithm was designed and trained using 53 EEG recordings from a long-term (a)EEG monitoring in 30 near-term neonates. The results were validated using an external dataset from 30 polysomnography recordings. In addition to training and validating a single EEG channel quiet sleep detector, we constructed Sleep State Trend (SST), a bedside-ready means for visualizing classifier outputs. Results: The accuracy of quiet sleep detection in the training data was 90%, and the accuracy was comparable (85-86%) in all bipolar derivations available from the 4-electrode recordings. The algorithm generalized well to an external dataset, showing 81% overall accuracy despite different signal derivations. SST allowed an intuitive, clear visualization of the classifier output. Conclusions: Fluctuations in sleep states can be detected at high fidelity from a single EEG channel, and the results can be visualized as a transparent and intuitive trend in the bedside monitors. Significance: The Sleep State Trend (SST) may provide caregivers a real-time view of sleep state fluctuations and its cyclicity.