Africa
Joint-training on Symbiosis Networks for Deep Nueral Machine Translation models
Yu, Zhengzhe, Guo, Jiaxin, Wang, Minghan, Wei, Daimeng, Shang, Hengchao, Li, Zongyao, Wu, Zhanglin, Wang, Yuxia, Chen, Yimeng, Su, Chang, Zhang, Min, Lei, Lizhi, tao, shimin, Yang, Hao
Deep encoders have been proven to be effective in improving neural machine translation (NMT) systems, but it reaches the upper bound of translation quality when the number of encoder layers exceeds 18. Worse still, deeper networks consume a lot of memory, making it impossible to train efficiently. In this paper, we present Symbiosis Networks, which include a full network as the Symbiosis Main Network (M-Net) and another shared sub-network with the same structure but less layers as the Symbiotic Sub Network (S-Net). We adopt Symbiosis Networks on Transformer-deep (m-n) architecture and define a particular regularization loss $\mathcal{L}_{\tau}$ between the M-Net and S-Net in NMT. We apply joint-training on the Symbiosis Networks and aim to improve the M-Net performance. Our proposed training strategy improves Transformer-deep (12-6) by 0.61, 0.49 and 0.69 BLEU over the baselines under classic training on WMT'14 EN->DE, DE->EN and EN->FR tasks. Furthermore, our Transformer-deep (12-6) even outperforms classic Transformer-deep (18-6).
Data driven design of optical resonators
Lenaerts, Joeri, Pinson, Hannah, Ginis, Vincent
Optical devices lie at the heart of most of the technology we see around us. When one actually wants to make such an optical device, one can predict its optical behavior using computational simulations of Maxwell's equations. If one then asks what the optimal design would be in order to obtain a certain optical behavior, the only way to go further would be to try out all of the possible designs and compute the electromagnetic spectrum they produce. When there are many design parameters, this brute force approach quickly becomes too computationally expensive. We therefore need other methods to create optimal optical devices. An alternative to the brute force approach is inverse design. In this paradigm, one starts from the desired optical response of a material and then determines the design parameters that are needed to obtain this optical response. There are many algorithms known in the literature that implement this inverse design. Some of the best performing, recent approaches are based on Deep Learning. The central idea is to train a neural network to predict the optical response for given design parameters. Since neural networks are completely differentiable, we can compute gradients of the response with respect to the design parameters. We can use these gradients to update the design parameters and get an optical response closer to the one we want. This allows us to obtain an optimal design much faster compared to the brute force approach. In my thesis, I use Deep Learning for the inverse design of the Fabry-P\'erot resonator. This system can be described fully analytically and is therefore ideal to study.
Identifying Mixtures of Bayesian Network Distributions
Gordon, Spencer L., Mazaheri, Bijan, Rabani, Yuval, Schulman, Leonard J.
A Bayesian Network is a directed acyclic graph (DAG) on a set of $n$ random variables (identified with the vertices); a Bayesian Network Distribution (BND) is a probability distribution on the rv's that is Markovian on the graph. A finite mixture of such models is the projection on these variables of a BND on the larger graph which has an additional "hidden" (or "latent") random variable $U$, ranging in $\{1,\ldots,k\}$, and a directed edge from $U$ to every other vertex. Models of this type are fundamental to research in Causal Inference, where $U$ models a confounding effect. One extremely special case has been of longstanding interest in the theory literature: the empty graph. Such a distribution is simply a mixture of $k$ product distributions. A longstanding problem has been, given the joint distribution of a mixture of $k$ product distributions, to identify each of the product distributions, and their mixture weights. Our results are: (1) We improve the sample complexity (and runtime) for identifying mixtures of $k$ product distributions from $\exp(O(k^2))$ to $\exp(O(k \log k))$. This is almost best possible in view of a known $\exp(\Omega(k))$ lower bound. (2) We give the first algorithm for the case of non-empty graphs. The complexity for a graph of maximum degree $\Delta$ is $\exp(O(k(\Delta^2 + \log k)))$. (The above complexities are approximate and suppress dependence on secondary parameters.)
MultiPath++: Efficient Information Fusion and Trajectory Aggregation for Behavior Prediction
Varadarajan, Balakrishnan, Hefny, Ahmed, Srivastava, Avikalp, Refaat, Khaled S., Nayakanti, Nigamaa, Cornman, Andre, Chen, Kan, Douillard, Bertrand, Lam, Chi Pang, Anguelov, Dragomir, Sapp, Benjamin
Predicting the future behavior of road users is one of the most challenging and important problems in autonomous driving. Applying deep learning to this problem requires fusing heterogeneous world state in the form of rich perception signals and map information, and inferring highly multi-modal distributions over possible futures. In this paper, we present MultiPath++, a future prediction model that achieves state-of-the-art performance on popular benchmarks. MultiPath++ improves the MultiPath architecture by revisiting many design choices. The first key design difference is a departure from dense image-based encoding of the input world state in favor of a sparse encoding of heterogeneous scene elements: MultiPath++ consumes compact and efficient polylines to describe road features, and raw agent state information directly (e.g., position, velocity, acceleration). We propose a context-aware fusion of these elements and develop a reusable multi-context gating fusion component. Second, we reconsider the choice of pre-defined, static anchors, and develop a way to learn latent anchor embeddings end-to-end in the model. Lastly, we explore ensembling and output aggregation techniques -- common in other ML domains -- and find effective variants for our probabilistic multimodal output representation. We perform an extensive ablation on these design choices, and show that our proposed model achieves state-of-the-art performance on the Argoverse Motion Forecasting Competition and the Waymo Open Dataset Motion Prediction Challenge.
Online content moderation: Can AI help clean up social media?
Dec 20 (Thomson Reuters Foundation) -Two days after it was sued by Rohingya refugees from Myanmar over allegations that it did not take action against hate speech, social media company Meta, formerly known as Facebook, announced a new artificial intelligence system to tackle harmful content. Machine learning tools have increasingly become the go-to solution for tech firms to police their platforms, but questions have been raised about their accuracy and their potential threat to freedom of speech. WHY ARE SOCIAL MEDIA FIRMS UNDER FIRE OVER CONTENT MODERATION? The $150 billion Rohingya class-action lawsuit filed this month came at the end of a tumultuous period for social media giants, which have been criticised for failing to effectively tackle hate speech online and increasing polarization. The complaint argues that calls for violence shared on Facebook contributed to real-world violence against the Rohingya community, which suffered a military crackdown in 2017 that refugees said included mass killings and rape.
Global Artificial Intelligence Consulting Service Market 2022 Size, Share, CAGR Status by Sales, Revenue, Global Growth Rate, Modern Trends, Emerging Demands, Industry Analysis, Key Players and Forecast 2027
Pune, Dec. 20, 2021 (GLOBE NEWSWIRE) -- Global Artificial Intelligence Consulting Service Market research report study covers the global and regional market with an in-depth analysis of the overall growth prospects in the market. Artificial Intelligence Consulting Service Market Research Report identifies various key manufacturers of the market. It helps the reader understand the strategies and collaborations that players are focusing on combatting competition in the market. The researchers used advanced primary and secondary research methodologies and tools for preparing this report on the Artificial Intelligence Consulting Service market. In 2021, the global Artificial Intelligence Consulting Service market size will be USD million and it is expected to reach USD million by the end of 2027, with a CAGR of % during 2021-2027.
'AI-driven' Label of Snafu: Will it Replace Record Executives with Technology?
Snafu's investor ABBA is looking for sounds from India. ABBA is a Swedish pop group that has anticipated its first album, Voyage, in 40 years. It will be streaming on the air on November 5. But before the release, the legendary comeback band sprinkled stardust on Snafu Records, a music label headed by an Indian. Snafu has introduced a new approach to search for music talent. Agnetha Fältskog, the ABBA singer, has joined a $6 million funding round for AI-powered record label Snafu records.
Japan and U.S. block advancement in U.N. talks on autonomous weapons
GENEVA – Japan, the United States and other countries have blocked any advancement in U.N. talks toward legally binding measures to ban and regulate the development and use of lethal autonomous weapon systems. The Sixth Review Conference of the Convention on Certain Conventional Weapons ended Friday in Geneva without progress, failing to reflect eight years of work and leaving countries and nongovernmental organizations that have called for legally binding rules expressing disappointment. Also referred to as "killer robots," autonomous weapons are artificial intelligence-powered weapons using facial recognition and algorithms. Once activated, the weapons can select and attack targets without the assistance of a human operator. They pose ethical, legal and security risks.
CausalMTA: Eliminating the User Confounding Bias for Causal Multi-touch Attribution
Yao, Di, Gong, Chang, Zhang, Lei, Chen, Sheng, Bi, Jingping
Multi-touch attribution (MTA), aiming to estimate the contribution of each advertisement touchpoint in conversion journeys, is essential for budget allocation and automatically advertising. Existing methods first train a model to predict the conversion probability of the advertisement journeys with historical data and calculate the attribution of each touchpoint using counterfactual predictions. An assumption of these works is the conversion prediction model is unbiased, i.e., it can give accurate predictions on any randomly assigned journey, including both the factual and counterfactual ones. Nevertheless, this assumption does not always hold as the exposed advertisements are recommended according to user preferences. This confounding bias of users would lead to an out-of-distribution (OOD) problem in the counterfactual prediction and cause concept drift in attribution. In this paper, we define the causal MTA task and propose CausalMTA to eliminate the influence of user preferences. It systemically eliminates the confounding bias from both static and dynamic preferences to learn the conversion prediction model using historical data. We also provide a theoretical analysis to prove CausalMTA can learn an unbiased prediction model with sufficient data. Extensive experiments on both public datasets and the impression data in an e-commerce company show that CausalMTA not only achieves better prediction performance than the state-of-the-art method but also generates meaningful attribution credits across different advertising channels.
Predicting treatment effects from observational studies using machine learning methods: A simulation study
Smith, Bevan I., Chimedza, Charles
Measuring treatment effects in observational studies is challenging because of confounding bias. Confounding occurs when a variable affects both the treatment and the outcome. Traditional methods such as propensity score matching estimate treatment effects by conditioning on the confounders. Recent literature has presented new methods that use machine learning to predict the counterfactuals in observational studies which then allow for estimating treatment effects. These studies however, have been applied to real world data where the true treatment effects have not been known. This study aimed to study the effectiveness of this counterfactual prediction method by simulating two main scenarios: with and without confounding. Each type also included linear and non-linear relationships between input and output data. The key item in the simulations was that we generated known true causal effects. Linear regression, lasso regression and random forest models were used to predict the counterfactuals and treatment effects. These were compared these with the true treatment effect as well as a naive treatment effect. The results show that the most important factor in whether this machine learning method performs well, is the degree of non-linearity in the data. Surprisingly, for both non-confounding \textit{and} confounding, the machine learning models all performed well on the linear dataset. However, when non-linearity was introduced, the models performed very poorly. Therefore under the conditions of this simulation study, the machine learning method performs well under conditions of linearity, even if confounding is present, but at this stage should not be trusted when non-linearity is introduced.