Goto

Collaborating Authors

 Media


Three Challenges Ahead for Stable Diffusion

#artificialintelligence

Stable Diffusion latent diffusion image synthesis model a couple of weeks ago may be one of the most significant technological disclosures since DeCSS in 1999; it's certainly the biggest event in AI-generated imagery since the 2017 deepfakes code was copied over to GitHub and forked into what would become DeepFaceLab and FaceSwap, as well as the real-time streaming deepfake software DeepFaceLive. At a stroke, user frustration over the content restrictions in DALL-E 2's image synthesis API were swept aside, as it transpired that Stable Diffusion's NSFW filter could be disabled by changing a sole line of code. Porn-centric Stable Diffusion Reddits sprung up almost immediately, and were as quickly cut down, while the developer and user camp divided on Discord into the official and NSFW communities, and Twitter began to fill up with fantastical Stable Diffusion creations. At the moment, each day seems to bring some amazing innovation from the developers who have adopted the system, with plugins and third-party adjuncts being hastily written for Krita, Photoshop, Cinema4D, Blender, and many other application platforms. In the meantime, promptcraft – the now- professional art of'AI whispering', which may end up being the shortest career option since'Filofax binder' – is already becoming commercialized, while early monetization of Stable Diffusion is taking place at the Patreon level, with the certainty of more sophisticated offerings to come, for those unwilling to navigate Conda-based installs of the source code, or the proscriptive NSFW filters of web-based implementations.


AI artist imagine what's outside the frame of famous paintings including Girl with a Pearl Earring

Daily Mail - Science & tech

An AI artist can now provide a glimpse of what the background settings of famous paintings and photos may have looked like. OpenAI, a San Francisco-based company, has created a new tool called'Outpainting' for its text-to-image AI system, DALL-E. Outpainting allows the system to imagine what's outside the frame of famous paintings such as Girl with The Pearl Earring, Mona Lisa and Dogs Playing Poker. As users have shown, it can do this with any kind of image, such as the man on the Quaker Oats logo and the cover of the Beatles album'Abbey Road'. DALL-E relies on artificial neural networks (ANNs), which simulate the way the brain works in order to learn and create an image from text.


Neural-Symbolic Models for Logical Queries on Knowledge Graphs

arXiv.org Artificial Intelligence

Answering complex first-order logic (FOL) queries on knowledge graphs is a fundamental task for multi-hop reasoning. Traditional symbolic methods traverse a complete knowledge graph to extract the answers, which provides good interpretation for each step. Recent neural methods learn geometric embeddings for complex queries. These methods can generalize to incomplete knowledge graphs, but their reasoning process is hard to interpret. In this paper, we propose Graph Neural Network Query Executor (GNN-QE), a neural-symbolic model that enjoys the advantages of both worlds. GNN-QE decomposes a complex FOL query into relation projections and logical operations over fuzzy sets, which provides interpretability for intermediate variables. To reason about the missing links, GNN-QE adapts a graph neural network from knowledge graph completion to execute the relation projections, and models the logical operations with product fuzzy logic. Experiments on 3 datasets show that GNN-QE significantly improves over previous state-of-the-art models in answering FOL queries. Meanwhile, GNN-QE can predict the number of answers without explicit supervision, and provide visualizations for intermediate variables.


Monolingual alignment of word senses and definitions in lexicographical resources

arXiv.org Artificial Intelligence

The focus of this thesis is broadly on the alignment of lexicographical data, particularly dictionaries. In order to tackle some of the challenges in this field, two main tasks of word sense alignment and translation inference are addressed. The first task aims to find an optimal alignment given the sense definitions of a headword in two different monolingual dictionaries. This is a challenging task, especially due to differences in sense granularity, coverage and description in two resources. After describing the characteristics of various lexical semantic resources, we introduce a benchmark containing 17 datasets of 15 languages where monolingual word senses and definitions are manually annotated across different resources by experts. In the creation of the benchmark, lexicographers' knowledge is incorporated through the annotations where a semantic relation, namely exact, narrower, broader, related or none, is selected for each sense pair. This benchmark can be used for evaluation purposes of word-sense alignment systems. The performance of a few alignment techniques based on textual and non-textual semantic similarity detection and semantic relation induction is evaluated using the benchmark. Finally, we extend this work to translation inference where translation pairs are induced to generate bilingual lexicons in an unsupervised way using various approaches based on graph analysis. This task is of particular interest for the creation of lexicographical resources for less-resourced and under-represented languages and also, assists in increasing coverage of the existing resources. From a practical point of view, the techniques and methods that are developed in this thesis are implemented within a tool that can facilitate the alignment task.


An Indoor Localization Dataset and Data Collection Framework with High Precision Position Annotation

arXiv.org Artificial Intelligence

We introduce a novel technique and an associated high resolution dataset that aims to precisely evaluate wireless signal based indoor positioning algorithms. The technique implements an augmented reality (AR) based positioning system that is used to annotate the wireless signal parameter data samples with high precision position data. We track the position of a practical and low cost navigable setup of cameras and a Bluetooth Low Energy (BLE) beacon in an area decorated with AR markers. We maximize the performance of the AR-based localization by using a redundant number of markers. Video streams captured by the cameras are subjected to a series of marker recognition, subset selection and filtering operations to yield highly precise pose estimations. Our results show that we can reduce the positional error of the AR localization system to a rate under 0.05 meters. The position data are then used to annotate the BLE data that are captured simultaneously by the sensors stationed in the environment, hence, constructing a wireless signal data set with the ground truth, which allows a wireless signal based localization system to be evaluated accurately.



Directed Speech Separation for Automatic Speech Recognition of Long Form Conversational Speech

arXiv.org Artificial Intelligence

Many of the recent advances in speech separation are primarily aimed at synthetic mixtures of short audio utterances with high degrees of overlap. Most of these approaches need an additional stitching step to stitch the separated speech chunks for long form audio. Since most of the approaches involve Permutation Invariant training (PIT), the order of separated speech chunks is nondeterministic and leads to difficulty in accurately stitching homogenous speaker chunks for downstream tasks like Automatic Speech Recognition (ASR). Also, most of these models are trained with synthetic mixtures and do not generalize to real conversational data. In this paper, we propose a speaker conditioned separator trained on speaker embeddings extracted directly from the mixed signal using an over-clustering based approach. This model naturally regulates the order of the separated chunks without the need for an additional stitching step. We also introduce a data sampling strategy with real and synthetic mixtures which generalizes well to real conversation speech. With this model and data sampling technique, we show significant improvements in speaker-attributed word error rate (SA-WER) on Hub5 data.


What is reinforcement learning? How AI trains itself – VentureBeat

#artificialintelligence

Reinforcement learning is the process by which a machine learning algorithm, robot, etc. can be programmed to respond to complex, real-time and real- …


Sentiment Analysis using Transformers - Part I - Analytics Vidhya

#artificialintelligence

The dataset has 25000 positive and negative reviews in the training set and 25000 positive and negative reviews in the test set. The image below shows the number of unique reviews and unique sentiment values in the dataset. The movie reviews are classified as having either a positive sentiment or a negative sentiment. The image below takes a peek at four reviews and their target sentiments. As can be seen from the keywords of the first three reviews – hooked, wonderful, unassuming, wonderful – lend the review a positive connotation.


New machine learning method to analyze complex scientific data of proteins

#artificialintelligence

Scientists have developed a method using machine learning to better analyze data from a powerful scientific tool: nuclear magnetic resonance (NMR). One way NMR data can be used is to understand proteins and chemical reactions in the human body. NMR is closely related to magnetic resonance imaging (MRI) for medical diagnosis. NMR spectrometers allow scientists to characterize the structure of molecules, such as proteins, but it can take highly skilled human experts a significant amount of time to analyze that data. This new machine learning method can analyze the data much more quickly and just as accurately.