Goto

Collaborating Authors

 Deep Learning


Microsoft Proposes GODIVA, A Text-To-Video Machine Learning Framework

#artificialintelligence

A collaboration between Microsoft Research Asia and Duke University has produced a machine learning system capable of generating video solely from a text prompt, without the use of Generative Adversarial Networks (GANs). The project is titled GODIVA (Generating Open-DomaIn Videos from nAtural Descriptions), and builds on some of the approaches used by OpenAI's DALL-E image synthesis system, revealed earlier this year. Early results from GODIVA, with frames from videos created from two prompts. The top two examples were generated from the prompt'Play golf on grass', and the bottom third from the prompt'A baseball game is played'. GODIVA uses the Vector Quantised-Variational AutoEncoder (VQ-VAE) model first introduced by researchers from Google's DeepMind project in 2018, and also an essential component in DALL-E's transformational capabilities. Earlier work: VQ-VAE infers frames from very limited supplied source material.


[N] Wired: It Began As an AI-Fueled Dungeon Game. It Got Much Darker (AI Dungeon + GPT-3)

#artificialintelligence

If real children are never involved in the process, what's the harm? Schoolgirl fantasies are extremely common in written fiction as well as drawn and live action pornography, despite the (fictional) subjects being underage. Should we police those, too? There's a clear line between generated fictitious content and actual victimization, and IMO AI Dungeon doesn't cross it. Obviously, they're under no obligation to intentionally host such content on their platform, but as has been demonstrated by the controversy it is quite difficult to actually police it in a way that both preserves user privacy and doesn't block similar but non-offensive content.


Top 20 Artificial Intelligence Research Labs In The World In 2021

#artificialintelligence

Artificial intelligence is continuously evolving and propagating across every industry. With much of the groundbreaking innovations moving the industry forward, the technology is continuously making headlines every day. AI refers to software or systems that perform intelligent tasks like those of human brains such as learning, reasoning, and judgment. Its applications range from automation and translation systems for natural languages that people use daily, to image recognition systems that help identify faces and letters from images. Today, AI is used in different forms including digital assistants, chatbots, and machine learning, among others.


A Gentle Introduction to Audio Classification With Tensorflow

#artificialintelligence

We have seen a lot of recent advances in deep learning related to vision and language fields, it is intuitive to understand why CNN performs very well on images, with pixel's local correlation, and how sequential models like RNNs or transformers also perform very well on language, with its sequential nature, but what about audio? In this article you will learn how to approach a simple audio classification problem, you will learn some of the common and efficient methods used, and the Tensorflow code to do it. Disclaimer: The code presented here is based on my work developed for the "Rainforest Connection Species Audio Detection" Kaggle competition, but for demonstration purposes, I will use the "Speech Commands" dataset. We usually have audio files in the ".wav" format, they are commonly referred to as waveforms, a waveform is a time series with the signal amplitude at each specific time, if we visualize one of those waveform samples we will get something like this: Intuitively one might consider modeling this data like a regular time series (e.g. stock price forecasting) using some kind of RNN model, in fact, this could be done, but since we are using audio signals, a more appropriate choice is to transform the waveform samples into spectrograms. A spectrogram is an image representation of the waveform signal, it shows its frequency intensity range over time, it can be very useful when we want to evaluate the signal's frequency distribution over time.


PyTorch on Google Cloud: How To train PyTorch models on AI Platform

#artificialintelligence

After creating the AI Platform Notebooks instance, you can start with your experiments. Let's look into the model specifics for the use case. For analyzing sentiments of the movie reviews in IMDB dataset, we will be fine-tuning a pre-trained BERT model from Hugging Face. Fine-tuning involves taking a model that has already been trained for a given task and then tweaking the model for another similar task. Specifically, the tweaking involves replicating all the layers in the pre-trained model including weights and parameters, except the output layer.


Toward a brain-like AI with hyperdimensional computing

#artificialintelligence

The human brain has always been under study for inspiration of computing systems. Although there's a very long way to go until we can achieve a computing system that matches the efficiency of the human brain for cognitive tasks, several brain-inspired computing paradigms are being researched. Convolutional neural networks are a widely used machine learning approach for AI-related applications due to their significant performance relative to rules-based or symbolic approaches. Nonetheless, for many tasks machine learning requires vast amounts of data and training to converge to an acceptable level of performance. A Ph.D. student from Khalifa University, Eman Hasan, is investigating another AI computation methodology called'hyperdimensional computing," which can possibly take AI systems a step closer toward human-like cognition.


Direct Prediction of Steady-State Flow Fields in Meshed Domain with Graph Networks

arXiv.org Artificial Intelligence

We propose a model to directly predict the steady-state flow field for a given geometry setup. The setup is an Eulerian representation of the fluid flow as a meshed domain. We introduce a graph network architecture to process the mesh-space simulation as a graph. The benefit of our model is a strong understanding of the global physical system, while being able to explore the local structure. This is essential to perform direct prediction and is thus superior to other existing methods.


Model-Driven Deep Learning Based Channel Estimation and Feedback for Millimeter-Wave Massive Hybrid MIMO Systems

arXiv.org Artificial Intelligence

This paper proposes a model-driven deep learning (MDDL)-based channel estimation and feedback scheme for wideband millimeter-wave (mmWave) massive hybrid multiple-input multiple-output (MIMO) systems, where the angle-delay domain channels' sparsity is exploited for reducing the overhead. Firstly, we consider the uplink channel estimation for time-division duplexing systems. To reduce the uplink pilot overhead for estimating the high-dimensional channels from a limited number of radio frequency (RF) chains at the base station (BS), we propose to jointly train the phase shift network and the channel estimator as an auto-encoder. Particularly, by exploiting the channels' structured sparsity from an a priori model and learning the integrated trainable parameters from the data samples, the proposed multiple-measurement-vectors learned approximate message passing (MMV-LAMP) network with the devised redundant dictionary can jointly recover multiple subcarriers' channels with significantly enhanced performance. Moreover, we consider the downlink channel estimation and feedback for frequency-division duplexing systems. Similarly, the pilots at the BS and channel estimator at the users can be jointly trained as an encoder and a decoder, respectively. Besides, to further reduce the channel feedback overhead, only the received pilots on part of the subcarriers are fed back to the BS, which can exploit the MMV-LAMP network to reconstruct the spatial-frequency channel matrix. Numerical results show that the proposed MDDL-based channel estimation and feedback scheme outperforms the state-of-the-art approaches.


Game Plan: What AI can do for Football, and What Football can do for AI

Journal of Artificial Intelligence Research

The rapid progress in artificial intelligence (AI) and machine learning has opened unprecedented analytics possibilities in various team and individual sports, including baseball, basketball, and tennis. More recently, AI techniques have been applied to football, due to a huge increase in data collection by professional teams, increased computational power, and advances in machine learning, with the goal of better addressing new scientific challenges involved in the analysis of both individual players' and coordinated teams' behaviors. The research challenges associated with predictive and prescriptive football analytics require new developments and progress at the intersection of statistical learning, game theory, and computer vision. In this paper, we provide an overarching perspective highlighting how the combination of these fields, in particular, forms a unique microcosm for AI research, while offering mutual benefits for professional teams, spectators, and broadcasters in the years to come. We illustrate that this duality makes football analytics a game changer of tremendous value, in terms of not only changing the game of football itself, but also in terms of what this domain can mean for the field of AI. We review the state-of-the-art and exemplify the types of analysis enabled by combining the aforementioned fields, including illustrative examples of counterfactual analysis using predictive models, and the combination of game-theoretic analysis of penalty kicks with statistical learning of player attributes. We conclude by highlighting envisioned downstream impacts, including possibilities for extensions to other sports (real and virtual).


PLSM: A Parallelized Liquid State Machine for Unintentional Action Detection

arXiv.org Artificial Intelligence

Reservoir Computing (RC) offers a viable option to deploy AI algorithms on low-end embedded system platforms. Liquid State Machine (LSM) is a bio-inspired RC model that mimics the cortical microcircuits and uses spiking neural networks (SNN) that can be directly realized on neuromorphic hardware. In this paper, we present a novel Parallelized LSM (PLSM) architecture that incorporates spatio-temporal read-out layer and semantic constraints on model output. To the best of our knowledge, such a formulation has been done for the first time in literature, and it offers a computationally lighter alternative to traditional deep-learning models. Additionally, we also present a comprehensive algorithm for the implementation of parallelizable SNNs and LSMs that are GPU-compatible. We implement the PLSM model to classify unintentional/accidental video clips, using the Oops dataset. From the experimental results on detecting unintentional action in video, it can be observed that our proposed model outperforms a self-supervised model and a fully supervised traditional deep learning model. All the implemented codes can be found at our repository https://github.com/anonymoussentience2020/Parallelized_LSM_for_Unintentional_Action_Recognition.