Africa
Differentiable Graph Module (DGM) Graph Convolutional Networks
Kazi, Anees, Cosmo, Luca, Navab, Nassir, Bronstein, Michael
Graph deep learning has recently emerged as a powerful ML concept allowing to generalize successful deep neural architectures to non-Euclidean structured data. Such methods have shown promising results on a broad spectrum of applications ranging from social science, biomedicine, and particle physics to computer vision, graphics, and chemistry. One of the limitations of the majority of the current graph neural network architectures is that they are often restricted to the transductive setting and rely on the assumption that the underlying graph is known and fixed. In many settings, such as those arising in medical and healthcare applications, this assumption is not necessarily true since the graph may be noisy, partially- or even completely unknown, and one is thus interested in inferring it from the data. This is especially important in inductive settings when dealing with nodes not present in the graph at training time. Furthermore, sometimes such a graph itself may convey insights that are even more important than the downstream task. In this paper, we introduce Differentiable Graph Module (DGM), a learnable function predicting the edge probability in the graph relevant for the task, that can be combined with convolutional graph neural network layers and trained in an end-to-end fashion. We provide an extensive evaluation of applications from the domains of healthcare (disease prediction), brain imaging (gender and age prediction), computer graphics (3D point cloud segmentation), and computer vision (zero-shot learning). We show that our model provides a significant improvement over baselines both in transductive and inductive settings and achieves state-of-the-art results.
Better Theory for SGD in the Nonconvex World
Khaled, Ahmed, Richtárik, Peter
Large-scale nonconvex optimization problems are ubiquitous in modern machine learning, and among practitioners interested in solving them, Stochastic Gradient Descent (SGD) reigns supreme. We revisit the analysis of SGD in the nonconvex setting and propose a new variant of the recently introduced expected smoothness assumption which governs the behaviour of the second moment of the stochastic gradient. We show that our assumption is both more general and more reasonable than assumptions made in all prior work. Moreover, our results yield the optimal $\mathcal{O}(\varepsilon^{-4})$ rate for finding a stationary point of nonconvex smooth functions, and recover the optimal $\mathcal{O}(\varepsilon^{-1})$ rate for finding a global solution if the Polyak-{\L}ojasiewicz condition is satisfied. We compare against convergence rates under convexity and prove a theorem on the convergence of SGD under Quadratic Functional Growth and convexity, which might be of independent interest. Moreover, we perform our analysis in a framework which allows for a detailed study of the effects of a wide array of sampling strategies and minibatch sizes for finite-sum optimization problems. We corroborate our theoretical results with experiments on real and synthetic data.
Fair Prediction with Endogenous Behavior
Jung, Christopher, Kannan, Sampath, Lee, Changhwa, Pai, Mallesh M., Roth, Aaron, Vohra, Rakesh
There is increasing regulatory interest in whether machine learning algorithms deployed in consequential domains (e.g. in criminal justice) treat different demographic groups "fairly." However, there are several proposed notions of fairness, typically mutually incompatible. Using criminal justice as an example, we study a model in which society chooses an incarceration rule. Agents of different demographic groups differ in their outside options (e.g. opportunity for legal employment) and decide whether to commit crimes. We show that equalizing type I and type II errors across groups is consistent with the goal of minimizing the overall crime rate; other popular notions of fairness are not.
Bill Gates: AI and gene therapy have the power to save lives
Microsoft founder Bill Gates thinks artificial intelligence and gene therapy are the two technologies with the greatest power to change lives. In a speech Friday at the American Association for the Advancement of Science, Gates said AI can "make sense of complex biological systems," while gene-based tools have the potential to cure AIDS. The potential of AI is only just being realized now, the billionaire philanthropist said, with computational power doubling every three and a half months. Along with improvements in handling data, Gates said it's enabling "the ability to synthesize, analyze, see patterns, gain insights and make predictions across many, many more dimensions than a human can comprehend." Gates said the most exciting part of AI "is how it can help us make sense of complex biological systems and accelerate the discovery of therapeutics to improve health in the poorest countries."
Handling Missing Annotations in Supervised Learning Data
Abdel-Hakim, Alaa E., Deabes, Wael
Data annotation is an essential stage in supervised learning. However, the annotation process is exhaustive and time consuming, specially for large datasets. Activities of Daily Living (ADL) recognition is an example of systems that exploit very large raw sensor data readings. In such systems, sensor readings are collected from activity-monitoring sensors in a 24/7 manner. The size of the generated dataset is so huge that it is almost impossible for a human annotator to give a certain label to every single instance in the dataset. This results in annotation gaps in the input data to the adopting supervised learning system. The performance of the recognition system is negatively affected by these gaps. In this work, we propose and investigate three different paradigms to handle these gaps. In the first paradigm, the gaps are taken out by dropping all unlabeled readings. A single "Unknown" or "Do-Nothing" label is given to the unlabeled readings within the operation of the second paradigm. The last paradigm handles these gaps by giving every one of them a unique label identifying the encapsulating deterministic labels. Also, we propose a semantic preprocessing method of annotation gaps by constructing a hybrid combination of some of these paradigms for further performance improvement. The performance of the proposed three paradigms and their hybrid combination is evaluated using an ADL benchmark dataset containing more than $2.5\times 10^6$ sensor readings that had been collected over more than nine months. The evaluation results emphasize the performance contrast under the operation of each paradigm and support a specific gap handling approach for better performance.
Causal Feature Discovery through Strategic Modification
Bechavod, Yahav, Ligett, Katrina, Wu, Zhiwei Steven, Ziani, Juba
As algorithmic decision-making takes a more and more important role in myriad application domains, incentives emerge to change the inputs presented to these algorithms--people may either invest in truly relevant attributes or strategically lie about their data. Recently, a collection of very interesting papers has explored various models of strategic behavior on the part of the classified individuals in learning settings, and ways to mitigate the harms to accuracy that can arise from falsified features [Dalvi et al., 2004, Brückner et al., 2012, Hardt et al., 2016, Dong et al., 2018]. Additionally, some recent work has focused on the design of learning algorithms that incentivize the classified individuals to make"good" investments in true changes to their variables[Kleinberg and Raghavan, 2019]. The present paper takes a different tack, and explores another potential effect of strategic investment in true changes to variables, in an online learning setting: we claim that interaction between the online learning and the strategic individuals may actually aid the learning algorithm in identifying causal variables. By causal, we mean, informally, variables such that changes in their true value cause changes in the true label and lead agents to improve.
$\pi$VAE: Encoding stochastic process priors with variational autoencoders
Mishra, Swapnil, Flaxman, Seth, Bhatt, Samir
Stochastic processes provide a mathematically elegant way model complex data. In theory, they provide flexible priors over function classes that can encode a wide range of interesting assumptions. In practice, however, efficient inference by optimisation or marginalisation is difficult, a problem further exacerbated with big data and high dimensional input spaces. We propose a novel variational autoencoder (VAE) called the prior encoding variational autoencoder ($\pi$VAE). The $\pi$VAE is finitely exchangeable and Kolmogorov consistent, and thus is a continuous stochastic process. We use $\pi$VAE to learn low dimensional embeddings of function classes. We show that our framework can accurately learn expressive function classes such as Gaussian processes, but also properties of functions to enable statistical inference (such as the integral of a log Gaussian process). For popular tasks, such as spatial interpolation, $\pi$VAE achieves state-of-the-art performance both in terms of accuracy and computational efficiency. Perhaps most usefully, we demonstrate that the low dimensional independently distributed latent space representation learnt provides an elegant and scalable means of performing Bayesian inference for stochastic processes within probabilistic programming languages such as Stan.
Conditional Self-Attention for Query-based Summarization
Xie, Yujia, Zhou, Tianyi, Mao, Yi, Chen, Weizhu
Self-attention mechanisms have achieved great success on a variety of NLP tasks due to its flexibility of capturing dependency between arbitrary positions in a sequence. For problems such as query-based summarization (Qsumm) and knowledge graph reasoning where each input sequence is associated with an extra query, explicitly modeling such conditional contextual dependencies can lead to a more accurate solution, which however cannot be captured by existing self-attention mechanisms. In this paper, we propose \textit{conditional self-attention} (CSA), a neural network module designed for conditional dependency modeling. CSA works by adjusting the pairwise attention between input tokens in a self-attention module with the matching score of the inputs to the given query. Thereby, the contextual dependencies modeled by CSA will be highly relevant to the query. We further studied variants of CSA defined by different types of attention. Experiments on Debatepedia and HotpotQA benchmark datasets show CSA consistently outperforms vanilla Transformer and previous models for the Qsumm problem.
6 Billion People's Personal Biometrics Stolen by China for their Quantum Artificial Intelligence Military Program - THE AI ORGANIZATION
China's Communist Government has extracted over 6 billion peoples biometrics, including facial, voice and personal health data to empower their Quantum Artificial Intelligence program meant for military purposes. This includes almost every American, Canadian, and European persons living today, every person in China, and Less so from groups in Africa, the Middle East, and South America. I initially made the finding public by publishing the discovery in the book AI, Trump, China and the Weaponization of Robotics without providing company names. Later, I included the findings with company names in the updated book Artificial Intelligence Dangers to Humanity. More than 1,000 AI, Robotics and Bio-Metric companies were researched to obtain the results of over 6 billion human beings who have had their bio-metrics stolen or transferred to China.
Deloitte Showcases Latest In AI Technology In Dubai Al Bawaba
Deloitte today launched its Middle East inaugural Experience Analytics event in Dubai at Dubai Studio City. Experience Analytics is a globally recognised Deloitte event and has previously taken place in London, Amsterdam and Berlin. The theme is'Me, Myself, and AI' and brings together a combination of technology showcases and practical sessions that explore a number of topics across Analytics and Artificial Intelligence. It is not about people vs. machines but about how human collaboration and decision-making can be enhanced through the use of machines – this has been coined by Deloitte as the "Age of With". "We believe that we are entering an important phase for society in the Age of With, and in order to make a true impact that matters, we understand how important it is to collaborate and leverage our relationships with our alliances and eco-system partners to build the best solutions for our clients. Some of our global alliance partners for the event are Google, SAP, Informatica and Cloudera," said Rajeev Lalwani, Deloitte's Leader for Strategy, Analytics and M&A in the Middle East.