Goto

Collaborating Authors

 Oceania


This artist uses AI to show how the world's streets could be more pedestrian-friendly

#artificialintelligence

Artist and musician Zach Katz virtually transforms the streets of the world's major cities in order to show how the space could be made more hospitable to pedestrians. Every day, he posts new creations on his Twitter account and, in the space of a few weeks, has become something of a star among Internet users, city planners and politicians. How do we imagine the downtown zones of the future? The Twitter account @Betterstreetsai explores this question by using DALL-E, the artificial intelligence (AI) that has started a real trend on the Web. Here, fountains, green space, rails or even roads reserved for bikes and cyclists take over the space to dramatically change the urban landscape.


NewsStories: Illustrating articles with visual summaries

arXiv.org Artificial Intelligence

Recent self-supervised approaches have used large-scale image-text datasets to learn powerful representations that transfer to many tasks without finetuning. These methods often assume that there is one-to-one correspondence between its images and their (short) captions. However, many tasks require reasoning about multiple images and long text narratives, such as describing news articles with visual summaries. Thus, we explore a novel setting where the goal is to learn a self-supervised visual-language representation that is robust to varying text length and the number of images. In addition, unlike prior work which assumed captions have a literal relation to the image, we assume images only contain loose illustrative correspondence with the text. To explore this problem, we introduce a large-scale multimodal dataset containing over 31M articles, 22M images and 1M videos. We show that state-of-the-art image-text alignment methods are not robust to longer narratives with multiple images. Finally, we introduce an intuitive baseline that outperforms these methods on zero-shot image-set retrieval by 10% on the GoodNews dataset.


InvisibiliTee: Angle-agnostic Cloaking from Person-Tracking Systems with a Tee

arXiv.org Artificial Intelligence

After a survey for person-tracking system-induced privacy concerns, we propose a black-box adversarial attack method on state-of-the-art human detection models called InvisibiliTee. The method learns printable adversarial patterns for T-shirts that cloak wearers in the physical world in front of person-tracking systems. We design an angle-agnostic learning scheme which utilizes segmentation of the fashion dataset and a geometric warping process so the adversarial patterns generated are effective in fooling person detectors from all camera angles and for unseen black-box detection models. Empirical results in both digital and physical environments show that with the InvisibiliTee on, person-tracking systems' ability to detect the wearer drops significantly.


$\beta$-Divergence-Based Latent Factorization of Tensors model for QoS prediction

arXiv.org Artificial Intelligence

A nonnegative latent factorization of tensors (NLFT) model can well model the temporal pattern hidden in nonnegative quality-of-service (QoS) data for predicting the unobserved ones with high accuracy. However, existing NLFT models' objective function is based on Euclidean distance, which is only a special case of $\beta$-divergence. Hence, can we build a generalized NLFT model via adopting $\beta$-divergence to achieve prediction accuracy gain? To tackle this issue, this paper proposes a $\beta$-divergence-based NLFT model ($\beta$-NLFT). Its ideas are two-fold 1) building a learning objective with $\beta$-divergence to achieve higher prediction accuracy, and 2) implementing self-adaptation of hyper-parameters to improve practicability. Empirical studies on two dynamic QoS datasets demonstrate that compared with state-of-the-art models, the proposed $\beta$-NLFT model achieves the higher prediction accuracy for unobserved QoS data.


Explainable Artificial Intelligence for Assault Sentence Prediction in New Zealand

arXiv.org Artificial Intelligence

The judiciary has historically been conservative in its use of Artificial Intelligence, but recent advances in machine learning have prompted scholars to reconsider such use in tasks like sentence prediction. This paper investigates by experimentation the potential use of explainable artificial intelligence for predicting imprisonment sentences in assault cases in New Zealand's courts. We propose a proof-of-concept explainable model and verify in practice that it is fit for purpose, with predicted sentences accurate to within one year. We further analyse the model to understand the most influential phrases in sentence length prediction. We conclude the paper with an evaluative discussion of the future benefits and risks of different ways of using such an AI model in New Zealand's courts.


Sharp-MAML: Sharpness-Aware Model-Agnostic Meta Learning

arXiv.org Artificial Intelligence

Model-agnostic meta learning (MAML) is currently one of the dominating approaches for few-shot meta-learning. Albeit its effectiveness, the optimization of MAML can be challenging due to the innate bilevel problem structure. Specifically, the loss landscape of MAML is much more complex with possibly more saddle points and local minimizers than its empirical risk minimization counterpart. To address this challenge, we leverage the recently invented sharpness-aware minimization and develop a sharpness-aware MAML approach that we term Sharp-MAML. We empirically demonstrate that Sharp-MAML and its computation-efficient variant can outperform the plain-vanilla MAML baseline (e.g., $+3\%$ accuracy on Mini-Imagenet). We complement the empirical study with the convergence rate analysis and the generalization bound of Sharp-MAML. To the best of our knowledge, this is the first empirical and theoretical study on sharpness-aware minimization in the context of bilevel learning. The code is available at https://github.com/mominabbass/Sharp-MAML.


Pruning Self-attentions into Convolutional Layers in Single Path

arXiv.org Artificial Intelligence

Vision Transformers (ViTs) have achieved impressive performance over various computer vision tasks. However, modelling global correlations with multi-head self-attention (MSA) layers leads to two widely recognized issues: the massive computational resource consumption and the lack of intrinsic inductive bias for modelling local visual patterns. To solve both issues, we devise a simple yet effective method named Single-Path Vision Transformer pruning (SPViT), to efficiently and automatically compress the pre-trained ViTs into compact models with proper locality added. Specifically, we first propose a novel weight-sharing scheme between MSA and convolutional operations, delivering a single-path space to encode all candidate operations. In this way, we cast the operation search problem as finding which subset of parameters to use in each MSA layer, which significantly reduces the computational cost and optimization difficulty, and the convolution kernels can be well initialized using pre-trained MSA parameters. Relying on the single-path space, we further introduce learnable binary gates to encode the operation choices, which are jointly optimized with network parameters to automatically determine the configuration of each layer. We conduct extensive experiments on two representative ViTs showing that our SPViT achieves a new SOTA for pruning on ImageNet-1k. For example, our SPViT can trim 52.0% FLOPs for DeiT-B and get an impressive 0.6% top-1 accuracy gain simultaneously. The source code is available at https://github.com/ziplab/SPViT.


A Theory for Knowledge Transfer in Continual Learning

arXiv.org Artificial Intelligence

Continual learning of a stream of tasks is an active area in deep neural networks. The main challenge investigated has been the phenomenon of catastrophic forgetting or interference of newly acquired knowledge with knowledge from previous tasks. Recent work has investigated forward knowledge transfer to new tasks. Backward transfer for improving knowledge gained during previous tasks has received much less attention. There is in general limited understanding of how knowledge transfer could aid tasks learned continually. We present a theory for knowledge transfer in continual supervised learning, which considers both forward and backward transfer. We aim at understanding their impact for increasingly knowledgeable learners. We derive error bounds for each of these transfer mechanisms. These bounds are agnostic to specific implementations (e.g. deep neural networks). We demonstrate that, for a continual learner that observes related tasks, both forward and backward transfer can contribute to an increasing performance as more tasks are observed.


How AI is helping revitalise indigenous languages - ITU Hub

#artificialintelligence

Thirty-five years ago, New Zealand adopted a law declaring the official status of Te Reo Māori, the language spoken by the country's indigenous Māori people. Decades of repression put the language, also called simply te reo, under serious threat: only one in four Māori spoke it by 1960, with a very low percentage of speakers among children. Since then, the language has started regaining lost ground, enjoying formally equal status with English and being taught widely to New Zealand schoolchildren. Still, reviving it as a living language takes time and persistence. Lately, the nascent te reo renaissance is gaining added momentum with the help of artificial intelligence (AI).


Deep Learning Alone Isn't Getting Us To Human-Like AI

#artificialintelligence

Of course, deep learning has made progress, but on those foundational questions, not so much; on natural language, compositionality and reasoning, which differ from the kinds of pattern recognition on which deep learning excels, these systems remain massively unreliable, exactly as you would expect from systems that rely on statistical correlations, rather than an algebra of abstraction. Minerva, the latest, greatest AI system as of this writing, with billions of "tokens" in its training, still struggles with multiplying 4-digit numbers.