Deep Learning
OLIVAW: Mastering Othello with neither Humans nor a Penny
Norelli, Antonio, Panconesi, Alessandro
We introduce OLIVAW, an AI Othello player adopting the design principles of the famous AlphaGo series. The main motivation behind OLIVAW was to attain exceptional competence in a non-trivial board game, but at a tiny fraction of the cost of its illustrious predecessors. In this paper we show how OLIVAW successfully met this challenge.
SRA-LSTM: Social Relationship Attention LSTM for Human Trajectory Prediction
Peng, Yusheng, Zhang, Gaofeng, Shi, Jun, Xu, Benzhu, Zheng, Liping
Pedestrian trajectory prediction for surveillance video is one of the important research topics in the field of computer vision and a key technology of intelligent surveillance systems. Social relationship among pedestrians is a key factor influencing pedestrian walking patterns but was mostly ignored in the literature. Pedestrians with different social relationships play different roles in the motion decision of target pedestrian. Motivated by this idea, we propose a Social Relationship Attention LSTM (SRA-LSTM) model to predict future trajectories. We design a social relationship encoder to obtain the representation of their social relationship through the relative position between each pair of pedestrians. Afterwards, the social relationship feature and latent movements are adopted to acquire the social relationship attention of this pair of pedestrians. Social interaction modeling is achieved by utilizing social relationship attention to aggregate movement information from neighbor pedestrians. Experimental results on two public walking pedestrian video datasets (ETH and UCY), our model achieves superior performance compared with state-of-the-art methods. Contrast experiments with other attention methods also demonstrate the effectiveness of social relationship attention.
Perun: Secure Multi-Stakeholder Machine Learning Framework with GPU Support
Ozga, Wojciech, Quoc, Do Le, Fetzer, Christof
Confidential multi-stakeholder machine learning (ML) allows multiple parties to perform collaborative data analytics while not revealing their intellectual property, such as ML source code, model, or datasets. State-of-the-art solutions based on homomorphic encryption incur a large performance overhead. Hardware-based solutions, such as trusted execution environments (TEEs), significantly improve the performance in inference computations but still suffer from low performance in training computations, e.g., deep neural networks model training, because of limited availability of protected memory and lack of GPU support. To address this problem, we designed and implemented Perun, a framework for confidential multi-stakeholder machine learning that allows users to make a trade-off between security and performance. Perun executes ML training on hardware accelerators (e.g., GPU) while providing security guarantees using trusted computing technologies, such as trusted platform module and integrity measurement architecture. Less compute-intensive workloads, such as inference, execute only inside TEE, thus at a lower trusted computing base. The evaluation shows that during the ML training on CIFAR-10 and real-world medical datasets, Perun achieved a 161x to 1560x speedup compared to a pure TEE-based approach.
Neural Transformation Learning for Deep Anomaly Detection Beyond Images
Qiu, Chen, Pfrommer, Timo, Kloft, Marius, Mandt, Stephan, Rudolph, Maja
Data transformations (e.g. rotations, reflections, and cropping) play an important role in self-supervised learning. Typically, images are transformed into different views, and neural networks trained on tasks involving these views produce useful feature representations for downstream tasks, including anomaly detection. However, for anomaly detection beyond image data, it is often unclear which transformations to use. Here we present a simple end-to-end procedure for anomaly detection with learnable transformations. The key idea is to embed the transformed data into a semantic space such that the transformed data still resemble their untransformed form, while different transformations are easily distinguishable. Extensive experiments on time series demonstrate that we significantly outperform existing methods on the one-vs.-rest setting but also on the more challenging n-vs.-rest anomaly-detection task. On tabular datasets from the medical and cyber-security domains, our method learns domain-specific transformations and detects anomalies more accurately than previous work.
Contextual Text Embeddings for Twi
Azunre, Paul, Osei, Salomey, Addo, Salomey, Adu-Gyamfi, Lawrence Asamoah, Moore, Stephen, Adabankah, Bernard, Opoku, Bernard, Asare-Nyarko, Clara, Nyarko, Samuel, Amoaba, Cynthia, Appiah, Esther Dansoa, Akwerh, Felix, Lawson, Richard Nii Lante, Budu, Joel, Debrah, Emmanuel, Boateng, Nana, Ofori, Wisdom, Buabeng-Munkoh, Edwin, Adjei, Franklin, Ampomah, Isaac Kojo Essel, Otoo, Joseph, Borkor, Reindorf, Mensah, Standylove Birago, Mensah, Lucien, Marcel, Mark Amoako, Amponsah, Anokye Acheampong, Hayfron-Acquah, James Ben
Transformer-based language models have been changing the modern Natural Language Processing (NLP) landscape for high-resource languages such as English, Chinese, Russian, etc. However, this technology does not yet exist for any Ghanaian language. In this paper, we introduce the first of such models for Twi or Akan, the most widely spoken Ghanaian language. The specific contribution of this research work is the development of several pretrained transformer language models for the Akuapem and Asante dialects of Twi, paving the way for advances in application areas such as Named Entity Recognition (NER), Neural Machine Translation (NMT), Sentiment Analysis (SA) and Part-of-Speech (POS) tagging. Specifically, we introduce four different flavours of ABENA -- A BERT model Now in Akan that is fine-tuned on a set of Akan corpora, and BAKO - BERT with Akan Knowledge only, which is trained from scratch. We open-source the model through the Hugging Face model hub and demonstrate its use via a simple sentiment classification example.
Scoring Graspability based on Grasp Regression for Better Grasp Prediction
Depierre, Amaury, Dellandréa, Emmanuel, Chen, Liming
Grasping objects is one of the most important abilities that a robot needs to master in order to interact with its environment. Current state-of-the-art methods rely on deep neural networks trained to jointly predict a graspability score together with a regression of an offset with respect to grasp reference parameters. However, these two predictions are performed independently, which can lead to a decrease in the actual graspability score when applying the predicted offset. Therefore, in this paper, we extend a state-of-the-art neural network with a scorer that evaluates the graspability of a given position, and introduce a novel loss function which correlates regression of grasp parameters with graspability score. We show that this novel architecture improves performance from 82.13% for a state-of-the-art grasp detection network to 85.74% on Jacquard dataset. When the learned model is transferred onto a real robot, the proposed method correlating graspability and grasp regression achieves a 92.4% rate compared to 88.1% for the baseline trained without the correlation.
Four Deep Learning Papers to Read in April 2021
Welcome to the April edition of the ‚Machine-Learning-Collage' series, where I provide an overview of the different Deep Learning research streams. So what is this series about? Simply put, I draft one-slide visual summaries of one of my favourite recent papers. At the end of the month all of the visual collages are collected in a summary blog post. Thereby, I hope to give you a visual and intuitive deep dive into some of the coolest trends.
Opinion: How AI can protect users in the online world
With more than 74 percent of Gen Z spending their free time online – averaging around 10 hours per day – it's safe to say their online and offline worlds are becoming entwined. With increased social media usage now the norm and all of us living our lives online a little bit more, we must look for ways to mitigate risks, protect our safety and filter out communications that are causing concern. Step forward, Artificial Intelligence (AI) – advanced machine learning technology that plays an important role in modern life and is fundamental in how today's social networks function. With just one click AI tools such as chatbots, algorithms and auto-suggestions impact what you see on your screen and how often you see it, creating a customised feed that has completely changed the way we interact on these platforms. By analysing our behaviours, deep learning tools can determine habits, likes and dislike and only display material they anticipate you will enjoy.
openai/CLIP
CLIP (Contrastive Language-Image Pre-Training) is a neural network trained on a variety of (image, text) pairs. It can be instructed in natural language to predict the most relevant text snippet, given an image, without directly optimizing for the task, similarly to the zero-shot capabilities of GPT-2 and 3. We found CLIP matches the performance of the original ResNet50 on ImageNet "zero-shot" without using any of the original 1.28M labeled examples, overcoming several major challenges in computer vision. First, install PyTorch 1.7.1 and torchvision, as well as small additional dependencies, and then install this repo as a Python package. Returns the model and the TorchVision transform needed by the model, specified by the model name returned by clip.available_models(). The name argument can also be a path to a local checkpoint.
Unstructured Privacy Data Risks: AI Can Help
As per Gartner, 65% of world population's data will be impacted due to privacy regulations by 2023. In fact, it might happen sooner as most countries wish to provide economic nationalism by restricting cross country data transfers and data rationing by global technology businesses. Another Independent trend coupled with the rise of tighter privacy regulations is the volume of unstructured data being collected. Combined, both structured & unstructured data are projected to grow at the rate of 7-12% on an annual basis. Technological advances along with ever falling storage prices have made it quite easy to collect unstructured data from the customers.