Goto

Collaborating Authors

 Oceania


Koopman Neural Forecaster for Time Series with Temporal Distribution Shifts

arXiv.org Artificial Intelligence

Temporal distributional shifts, with underlying dynamics changing over time, frequently occur in real-world time series, and pose a fundamental challenge for deep neural networks (DNNs). In this paper, we propose a novel deep sequence model based on the Koopman theory for time series forecasting: Koopman Neural Forecaster (KNF) that leverages DNNs to learn the linear Koopman space and the coefficients of chosen measurement functions. KNF imposes appropriate inductive biases for improved robustness against distributional shifts, employing both a global operator to learn shared characteristics, and a local operator to capture changing dynamics, as well as a specially-designed feedback loop to continuously update the learnt operators over time for rapidly varying behaviors. We demonstrate that KNF achieves the superior performance compared to the alternatives, on multiple time series datasets that are shown to suffer from distribution shifts. Temporal distribution shifts frequently occur in real-world time-series applications, from forecasting stock prices to detecting and monitoring sensory measures, to predicting fashion trend based sales. Such distribution shifts over time may due to the data being generated in a highly-dynamic and non-stationary environment, abrupt changes that are difficult to predict, or constantly evolving trends in the underlying data distribution (Gama et al., 2014). Temporal distribution shifts pose a fundamental challenge for time-series forecasting (Kuznetsov & Mohri, 2020). There are two scenarios of distribution shifts. When the distribution shifts only occur between the training and test domains, meta learning and transfer learning approaches (Jin et al., 2021; Oreshkin et al., 2021) have been developed.


Toward Robust Uncertainty Estimation with Random Activation Functions

arXiv.org Artificial Intelligence

In this paper, we focus on ensemble UQ techniques, either Bayesian Recent advances in deep neural networks have demonstrated or non-Bayesian, as this group is less explored compared to remarkable performance in a wide variety of applications, the solely Bayesian techniques. An ensemble model aggregates ranging from recommendation systems and improving user the predictions of multiple individual base-learners (or experience to natural language processing and speech recognition ensemble members), which in our case are neural networks (Abiodun et al. 2018). Nevertheless, blindly relying (NNs), and the empirical variance of their predictions gives on the outcome of these models can have harmful effects, an approximate measure of uncertainty. The idea behind this especially in high-stake domains such as healthcare heuristic is highly intuitive: the more the base-learners disagree and autonomous driving, as models can provide inaccurate on the outcome, the more uncertain they are. Therefore, predictions when queried in out-of-distribution data the goal of ensemble members is to have a great level points (Amodei et al. 2016). Consequently, correctly quantifying of disagreement (variability) in the areas where little or no the uncertainty of models' predictions is an admissible data is available, and to have a high level of agreement in mechanism to distinguish where a model can or cannot regions with abundance of data (Pearce et al. 2018).


Fairness in Language Models Beyond English: Gaps and Challenges

arXiv.org Artificial Intelligence

With language models becoming increasingly ubiquitous, it has become essential to address their inequitable treatment of diverse demographic groups and factors. Most research on evaluating and mitigating fairness harms has been concentrated on English, while multilingual models and non-English languages have received comparatively little attention. This paper presents a survey of fairness in multilingual and non-English contexts, highlighting the shortcomings of current research and the difficulties faced by methods designed for English. We contend that the multitude of diverse cultures and languages across the world makes it infeasible to achieve comprehensive coverage in terms of constructing fairness datasets. Thus, the measurement and mitigation of biases must evolve beyond the current dataset-driven practices that are narrowly focused on specific dimensions and types of biases and, therefore, impossible to scale across languages and cultures.


SMoA: Sparse Mixture of Adapters to Mitigate Multiple Dataset Biases

arXiv.org Artificial Intelligence

Recent studies reveal that various biases exist in different NLP tasks, and over-reliance on biases results in models' poor generalization ability and low adversarial robustness. To mitigate datasets biases, previous works propose lots of debiasing techniques to tackle specific biases, which perform well on respective adversarial sets but fail to mitigate other biases. In this paper, we propose a new debiasing method Sparse Mixture-of-Adapters (SMoA), which can mitigate multiple dataset biases effectively and efficiently. Experiments on Natural Language Inference and Paraphrase Identification tasks demonstrate that SMoA outperforms full-finetuning, adapter tuning baselines, and prior strong debiasing methods. Further analysis indicates the interpretability of SMoA that sub-adapter can capture specific pattern from the training data and specialize to handle specific bias.


Generalization Performance of Empirical Risk Minimization on Over-parameterized Deep ReLU Nets

arXiv.org Artificial Intelligence

In this paper, we study the generalization performance of global minima for implementing empirical risk minimization (ERM) on over-parameterized deep ReLU nets. Using a novel deepening scheme for deep ReLU nets, we rigorously prove that there exist perfect global minima achieving almost optimal generalization error bounds for numerous types of data under mild conditions. Since over-parameterization is crucial to guarantee that the global minima of ERM on deep ReLU nets can be realized by the widely used stochastic gradient descent (SGD) algorithm, our results indeed fill a gap between optimization and generalization.


Children taking the IB WILL be allowed to use ChatGPT to write essays

#artificialintelligence

Controversial AI tool ChatGPT has already been banned in schools across the world over fears it encourages cheating and laziness. But the International Baccalaureate (IB), which offers an alternative to A-levels, is bucking this trend by permitting the use of ChatGPT to write essays. Students undertaking IB programmes will be able to quote passages generated by the chatbot - as long as they do not try to pass it off as their own words. Created by San Francisco-based company OpenAI, the tool has been trained on a massive amount of text so it can generate human-like responses to questions. A university student has already used ChatGPT to write a 2,000-word essay that got a 2:2 grade, although the lecturer called the language used'fishy'.


Children taking the IB WILL be allowed to use AI chatbot ChatGPT to write their essays

Daily Mail - Science & tech

Controversial AI tool ChatGPT has already been banned in schools across the world over fears it encourages cheating and laziness. But the International Baccalaureate (IB), which offers an alternative to A-levels, is bucking this trend by permitting the use of ChatGPT to write essays. Students undertaking IB programmes will be able to quote passages generated by the chatbot - as long as they do not try to pass it off as their own words. Created by San Francisco-based company OpenAI, the tool has been trained on a massive amount of text so it can generate human-like responses to questions. A university student has already used ChatGPT to write a 2,000-word essay that got a 2:2 grade, although the lecturer called the language used'fishy'.


An evaluation of Google Translate for Sanskrit to English translation via sentiment and semantic analysis

arXiv.org Artificial Intelligence

Google Translate has been prominent for language translation; however, limited work has been done in evaluating the quality of translation when compared to human experts. Sanskrit one of the oldest written languages in the world. In 2022, the Sanskrit language was added to the Google Translate engine. Sanskrit is known as the mother of languages such as Hindi and an ancient source of the Indo-European group of languages. Sanskrit is the original language for sacred Hindu texts such as the Bhagavad Gita. In this study, we present a framework that evaluates the Google Translate for Sanskrit using the Bhagavad Gita. We first publish a translation of the Bhagavad Gita in Sanskrit using Google Translate. Our framework then compares Google Translate version of Bhagavad Gita with expert translations using sentiment and semantic analysis via BERT-based language models. Our results indicate that in terms of sentiment and semantic analysis, there is low level of similarity in selected verses of Google Translate when compared to expert translations. In the qualitative evaluation, we find that Google translate is unsuitable for translation of certain Sanskrit words and phrases due to its poetic nature, contextual significance, metaphor and imagery. The mistranslations are not surprising since the Bhagavad Gita is known as a difficult text not only to translate, but also to interpret since it relies on contextual, philosophical and historical information. Our framework lays the foundation for automatic evaluation of other languages by Google Translate


Towards Audit Requirements for AI-based Systems in Mobility Applications

arXiv.org Artificial Intelligence

Various mobility applications like advanced driver assistance systems increasingly utilize artificial intelligence (AI) based functionalities. Typically, deep neural networks (DNNs) are used as these provide the best performance on the challenging perception, prediction or planning tasks that occur in real driving environments. However, current regulations like UNECE R 155 or ISO 26262 do not consider AI-related aspects and are only applied to traditional algorithm-based systems. The non-existence of AI-specific standards or norms prevents the practical application and can harm the trust level of users. Hence, it is important to extend existing standardization for security and safety to consider AI-specific challenges and requirements. To take a step towards a suitable regulation we propose 50 technical requirements or best practices that extend existing regulations and address the concrete needs for DNN-based systems. We show the applicability, usefulness and meaningfulness of the proposed requirements by performing an exemplary audit of a DNN-based traffic sign recognition system using three of the proposed requirements.


Online Black-Box Confidence Estimation of Deep Neural Networks

arXiv.org Artificial Intelligence

Autonomous driving (AD) and advanced driver assistance systems (ADAS) increasingly utilize deep neural networks (DNNs) for improved perception or planning. Nevertheless, DNNs are quite brittle when the data distribution during inference deviates from the data distribution during training. This represents a challenge when deploying in partly unknown environments like in the case of ADAS. At the same time, the standard confidence of DNNs remains high even if the classification reliability decreases. This is problematic since following motion control algorithms consider the apparently confident prediction as reliable even though it might be considerably wrong. To reduce this problem real-time capable confidence estimation is required that better aligns with the actual reliability of the DNN classification. Additionally, the need exists for black-box confidence estimation to enable the homogeneous inclusion of externally developed components to an entire system. In this work we explore this use case and introduce the neighborhood confidence (NHC) which estimates the confidence of an arbitrary DNN for classification. The metric can be used for black-box systems since only the top-1 class output is required and does not need access to the gradients, the training dataset or a hold-out validation dataset. Evaluation on different data distributions, including small in-domain distribution shifts, out-of-domain data or adversarial attacks, shows that the NHC performs better or on par with a comparable method for online white-box confidence estimation in low data regimes which is required for real-time capable AD/ADAS.