Oceania
Symbolic Knowledge Extraction from Opaque Predictors Applied to Cosmic-Ray Data Gathered with LISA Pathfinder
Sabbatini, Federico, Grimani, Catia
Machine learning models are nowadays ubiquitous in space missions, performing a wide variety of tasks ranging from the prediction of multivariate time series through the detection of specific patterns in the input data. Adopted models are usually deep neural networks or other complex machine learning algorithms providing predictions that are opaque, i.e., human users are not allowed to understand the rationale behind the provided predictions. Several techniques exist in the literature to combine the impressive predictive performance of opaque machine learning models with human-intelligible prediction explanations, as for instance the application of symbolic knowledge extraction procedures. In this paper are reported the results of different knowledge extractors applied to an ensemble predictor capable of reproducing cosmic-ray data gathered on board the LISA Pathfinder space mission. A discussion about the readability/fidelity trade-off of the extracted knowledge is also presented.
Senior Software Engineer, Decisions Foundations (ML Platform)
Affirm is reinventing credit to make it more honest and friendly, giving consumers the flexibility to buy now and pay later without any hidden fees or compounding interest. We are looking for a Senior Engineer to lead projects and initiatives on our newest subteam within ML Platform: The ML Developer Productivity team. As an early team member, you will contribute extensively to setting and executing on a vision for increasing ML velocity at Affirm through a focus on developer productivity. Our mission is to build a self-service, easy-to-use foundation for developing and delivering robust models to production. This is a new team that will focus on creating internal tools used by ML Engineers for fast paced ML development.
Can TinyML really provide on-device learning? - Stacey on IoT
Imagine if your smart speaker could be trained to recognize your accent, or if a pair of running shoes could alert you in real time if your gait changed, indicating fatigue. Or if, in the industrial world, sensors could parse vibration information from a machine that changed location and function often in real time, halting the machine if that information suggested there was a problem. We often write about the value of on-device machine learning (ML), but what we're generally discussing is running existing models on a device and matching incoming data against the established model. This is known as inference. So when you say the name "Alexa," your smart speaker matches the pattern and wakes up.
Understanding Functions in AI
Every single data transformation we do in Artificial intelligence seeks to convert input-data to the most representative format required for the task we aim to solve… This conversion is done through functions. A machine-learning model transforms its input data into meaningful outputs. A process that is "learned" from exposure to known examples of inputs and outputs. Thus, the ML-model "learns a function" that maps its input data to the expected output. We have a table of a few data points, some belong to a "white" class and others to a "black" class.
Making AI accessible leads to greater innovation
It's difficult to visualise the true scale of AI, as it's almost certainly more than you imagine – it's going to contribute more to the global economy than the current GDP of India and China combined. PwC research suggests that AI could contribute as much as $15.7 trillion by 2030, and by singularly responsible for a 26 per cent boost in the GDP of local economies. That would place it as one of the most fundamentally transformational changes in human history and, PwC notes, there is the opportunity for emerging economies to leapfrog developed ones by being faster on the AI uptake. However, for all the promise of AI, there remain challenges. Gartner research suggest that only 54 per cent of AI projects make it from pilot to production. The challenge, Garter says, has to do with scale.
Portail Emploi CNRS - Offre d'emploi - Bioimage Analyst Position(M/W)
The main mission of the engineer is to participate in the project-based creation of automated image analysis tools. He will also guide and advise facility users in matters of image analysis, and promote the use of best practices in biological image analysis. Activities Design and implement tools for automated image analysis and processing using ImageJ/Fiji, java, python and other platforms. Set-up advanced workflows, including: image segmentation, quantification of intracellular protein distribution, pattern recognition, tracking of dynamic particles Train and integrate deep-learning methods into image analysis workflows Work with scientists to write and implement algorithms for solving image processing problems from multidimensional fluorescence microscopy datasets. While no formal biology training is needed, a strong interest in biology would facilitate the interaction with biologist users.
MICO: Selective Search with Mutual Information Co-training
Wang, Zhanyu, Zhang, Xiao, Yun, Hyokun, Teo, Choon Hui, Chilimbi, Trishul
In contrast to traditional exhaustive search, selective search first clusters documents into several groups before all the documents are searched exhaustively by a query, to limit the search executed within one group or only a few groups. Selective search is designed to reduce the latency and computation in modern large-scale search systems. In this study, we propose MICO, a Mutual Information CO-training framework for selective search with minimal supervision using the search logs. After training, MICO does not only cluster the documents, but also routes unseen queries to the relevant clusters for efficient retrieval. In our empirical experiments, MICO significantly improves the performance on multiple metrics of selective search and outperforms a number of existing competitive baselines.
T-NER: An All-Round Python Library for Transformer-based Named Entity Recognition
Ushio, Asahi, Camacho-Collados, Jose
Language model (LM) pretraining has led to consistent improvements in many NLP downstream tasks, including named entity recognition (NER). In this paper, we present T-NER (Transformer-based Named Entity Recognition), a Python library for NER LM finetuning. In addition to its practical utility, T-NER facilitates the study and investigation of the cross-domain and cross-lingual generalization ability of LMs finetuned on NER. Our library also provides a web app where users can get model predictions interactively for arbitrary text, which facilitates qualitative model evaluation for non-expert programmers. We show the potential of the library by compiling nine public NER datasets into a unified format and evaluating the cross-domain and cross-lingual performance across the datasets. The results from our initial experiments show that in-domain performance is generally competitive across datasets. However, cross-domain generalization is challenging even with a large pretrained LM, which has nevertheless capacity to learn domain-specific features if fine-tuned on a combined dataset. To facilitate future research, we also release all our LM checkpoints via the Hugging Face model hub.
Estimating Multi-label Accuracy using Labelset Distributions
Park, Laurence A. F., Read, Jesse
A multi-label classifier estimates the binary label state (relevant vs irrelevant) for each of a set of concept labels, for any given instance. Probabilistic multi-label classifiers provide a predictive posterior distribution over all possible labelset combinations of such label states (the powerset of labels) from which we can provide the best estimate, simply by selecting the labelset corresponding to the largest expected accuracy, over that distribution. For example, in maximizing exact match accuracy, we provide the mode of the distribution. But how does this relate to the confidence we may have in such an estimate? Confidence is an important element of real-world applications of multi-label classifiers (as in machine learning in general) and is an important ingredient in explainability and interpretability. However, it is not obvious how to provide confidence in the multi-label context and relating to a particular accuracy metric, and nor is it clear how to provide a confidence which correlates well with the expected accuracy, which would be most valuable in real-world decision making. In this article we estimate the expected accuracy as a surrogate for confidence, for a given accuracy metric. We hypothesise that the expected accuracy can be estimated from the multi-label predictive distribution. We examine seven candidate functions for their ability to estimate expected accuracy from the predictive distribution. We found three of these to correlate to expected accuracy and are robust. Further, we determined that each candidate function can be used separately to estimate Hamming similarity, but a combination of the candidates was best for expected Jaccard index and exact match.
General Place Recognition Survey: Towards the Real-world Autonomy Age
Yin, Peng, Zhao, Shiqi, Cisneros, Ivan, Abuduweili, Abulikemu, Huang, Guoquan, Milford, Micheal, Liu, Changliu, Choset, Howie, Scherer, Sebastian
Place recognition is the fundamental module that can assist Simultaneous Localization and Mapping (SLAM) in loop-closure detection and re-localization for long-term navigation. The place recognition community has made astonishing progress over the last $20$ years, and this has attracted widespread research interest and application in multiple fields such as computer vision and robotics. However, few methods have shown promising place recognition performance in complex real-world scenarios, where long-term and large-scale appearance changes usually result in failures. Additionally, there is a lack of an integrated framework amongst the state-of-the-art methods that can handle all of the challenges in place recognition, which include appearance changes, viewpoint differences, robustness to unknown areas, and efficiency in real-world applications. In this work, we survey the state-of-the-art methods that target long-term localization and discuss future directions and opportunities. We start by investigating the formulation of place recognition in long-term autonomy and the major challenges in real-world environments. We then review the recent works in place recognition for different sensor modalities and current strategies for dealing with various place recognition challenges. Finally, we review the existing datasets for long-term localization and introduce our datasets and evaluation API for different approaches. This paper can be a tutorial for researchers new to the place recognition community and those who care about long-term robotics autonomy. We also provide our opinion on the frequently asked question in robotics: Do robots need accurate localization for long-term autonomy? A summary of this work and our datasets and evaluation API is publicly available to the robotics community at: https://github.com/MetaSLAM/GPRS.