Asia
Unsupervised Singing Voice Conversion
We present a deep learning method for singing voice conversion. The proposed network is not conditioned on the text or on the notes, and it directly converts the audio of one singer to the voice of another. Training is performed without any form of supervision: no lyrics or any kind of phonetic features, no notes, and no matching samples between singers. The proposed network employs a single CNN encoder for all singers, a single WaveNet decoder, and a classifier that enforces the latent representation to be singer-agnostic. Each singer is represented by one embedding vector, which the decoder is conditioned on. In order to deal with relatively small datasets, we propose a new data augmentation scheme, as well as new training losses and protocols that are based on backtranslation. Our evaluation presents evidence that the conversion produces natural signing voices that are highly recognizable as the target singer.
OCKELM+: Kernel Extreme Learning Machine based One-class Classification using Privileged Information (or KOC+: Kernel Ridge Regression or Least Square SVM with zero bias based One-class Classification using Privileged Information)
Gautam, Chandan, Tiwari, Aruna, Tanveer, M.
Kernel method-based one-class classifier is mainly used for outlier or novelty detection. In this letter, kernel ridge regression (KRR) based one-class classifier (KOC) has been extended for learning using privileged information (LUPI). LUPI-based KOC method is referred to as KOC+. This privileged information is available as a feature with the dataset but only for training (not for testing). KOC+ utilizes the privileged information differently compared to normal feature information by using a so-called correction function. Privileged information helps KOC+ in achieving better generalization performance which is exhibited in this letter by testing the classifiers with and without privileged information. Existing and proposed classifiers are evaluated on the datasets from UCI machine learning repository and also on MNIST dataset. Moreover, experimental results evince the advantage of KOC+ over KOC and support vector machine (SVM) based one-class classifiers.
HAKE: Human Activity Knowledge Engine
Li, Yong-Lu, Xu, Liang, Huang, Xijie, Liu, Xinpeng, Ma, Ze, Chen, Mingyang, Wang, Shiyi, Fang, Hao-Shu, Lu, Cewu
Human activity understanding is crucial for building automatic intelligent system. With the help of deep learning, activity understanding has made huge progress recently. But some challenges such as imbalanced data distribution, action ambiguity, complex visual patterns still remain. To address these and promote the activity understanding, we build a large-scale Human Activity Knowledge Engine (HAKE) based on the human body part states. Upon existing activity datasets, we annotate the part states of all the active persons in all images, thus establish the relationship between instance activity and body part states. Furthermore, we propose a HAKE based part state recognition model with a knowledge extractor named Activity2Vec and a corresponding part state based reasoning network. With HAKE, our method can alleviate the learning difficulty brought by the long-tail data distribution, and bring in interpretability. Now our HAKE has more than 7 M+ part state annotations and is still under construction. We first validate our approach on a part of HAKE in this preliminary paper, where we show 7.2 mAP performance improvement on Human-Object Interaction recognition, and 12.38 mAP improvement on the one-shot subsets.
UR-FUNNY: A Multimodal Language Dataset for Understanding Humor
Hasan, Md Kamrul, Rahman, Wasifur, Zadeh, Amir, Zhong, Jianyuan, Tanveer, Md Iftekhar, Morency, Louis-Philippe, Mohammed, null, Hoque, null
Humor is a unique and creative communicative behavior displayed during social interactions. It is produced in a multimodal manner, through the usage of words (text), gestures (vision) and prosodic cues (acoustic). Understanding humor from these three modalities falls within boundaries of multimodal language; a recent research trend in natural language processing that models natural language as it happens in face-to-face communication. Although humor detection is an established research area in NLP, in a multimodal context it is an understudied area. This paper presents a diverse multimodal dataset, called UR-FUNNY, to open the door to understanding multimodal language used in expressing humor. The dataset and accompanying studies, present a framework in multimodal humor detection for the natural language processing community. UR-FUNNY is publicly available for research.
Self-Paced Probabilistic Principal Component Analysis for Data with Outliers
Zhao, Bowen, Xiao, Xi, Zhang, Wanpeng, Zhang, Bin, Xia, Shutao
Principal Component Analysis (PCA) is a popular tool for dimensionality reduction and feature extraction in data analysis. There is a probabilistic version of PCA, known as Probabilistic PCA (PPCA). However, standard PCA and PPCA are not robust, as they are sensitive to outliers. To alleviate this problem, this paper introduces the Self-Paced Learning mechanism into PPCA, and proposes a novel method called Self-Paced Probabilistic Principal Component Analysis (SP-PPCA). Furthermore, we design the corresponding optimization algorithm based on the alternative search strategy and the expectation-maximization algorithm. SP-PPCA looks for optimal projection vectors and filters out outliers iteratively. Experiments on both synthetic problems and real-world datasets clearly demonstrate that SP-PPCA is able to reduce or eliminate the impact of outliers.
Maximum Correntropy Criterion with Variable Center
Chen, Badong, Wang, Xin, Li, Yingsong, Principe, Jose C.
Correntropy is a local similarity measure defined in kernel space and the maximum correntropy criterion (MCC) has been successfully applied in many areas of signal processing and machine learning in recent years. The kernel function in correntropy is usually restricted to the Gaussian function with center located at zero. However, zero-mean Gaussian function may not be a good choice for many practical applications. In this study, we propose an extended version of correntropy, whose center can locate at any position. Accordingly, we propose a new optimization criterion called maximum correntropy criterion with variable center (MCC-VC). We also propose an efficient approach to optimize the kernel width and center location in MCC-VC. Simulation results of regression with linear in parameters (LIP) models confirm the desirable performance of the new method.
How machine learning is improving manufacturing product quality and supply chain visibility
Bottom Line: Manufacturers' most valuable data is generated on shop floors daily, bringing with it the challenge of analysing it to find prescriptive insights fast โ and an ideal problem for machine learning to solve. Manufacturing is the most data-prolific industry there is, generating on average 1.9 petabytes of data every year according to the McKinsey Global Insititute. Supply chains, sourcing, factory operations, and the phases of compliance and quality management generate the majority of data. The most valuable data of all comes from product inspections that can immediately find exceptionally strong or weak suppliers, quality management and compliance practices in a factory. Manufacturing's massive problem is in getting quality inspection results out fast enough across brands & retailers, other factories, suppliers and vendors to make a difference in future product quality.
The 10 Hottest AI Fintech Startups in Europe Fintech Schweiz Digital Finance News - FintechNewsCH
Artificial intelligence (AI) has become a critical aspect in financial services. Financial institutions around the world are making efforts to adopt AI for task automation, customer services, behavior analysis, as well as fraud finding, and are making large-scale investments in related technologies. The World Economic Forum (WEF) estimates the number to reach US$10 billion by 2020. In financial services, applications for AI technologies exist across nearly the entire spectrum of business, from algorithmic stock trading applications and credit card fraud detection, to auto investment advisors and market research and sentiment analysis. The following 10 AI fintech companies are some of Europe's rising stars to watch very closely: Swiss startup Parashift develops AI-based accounting document management technologies which it offers through a SaaS platform and APIs.
What you may not understand about China's AI scene
Jeff Ding, a researcher at the University of Oxford who studies China's AI development, shared some recent reflections on the most important things he's learned in the past year. They offer a great snapshot into the current state of the industry, so I've summarized them briefly below. The Chinese- and English-speaking AI communities have an asymmetrical understanding of each other. Most Chinese researchers can read English, and nearly all major research developments in the Western world are immediately translated into Chinese, but the reverse is not true. Therefore, the Chinese research community has a much deeper understanding than the English-speaking one of what's happening on both sides of the aisle.
What you may not understand about China's AI scene
Jeff Ding, a researcher at the University of Oxford who studies China's AI development, shared some recent reflections on the most important things he's learned in the past year. They offer a great snapshot into the current state of the industry, so I've summarized them briefly below. The Chinese- and English-speaking AI communities have an asymmetrical understanding of each other. Most Chinese researchers can read English, and nearly all major research developments in the Western world are immediately translated into Chinese, but the reverse is not true. Therefore, the Chinese research community has a much deeper understanding than the English-speaking one of what's happening on both sides of the aisle.