Africa
Interpretability in Machine Learning
Should we always trust a model that performs well? A model could reject your application for a mortgage or diagnose you with cancer. The consequences of these decisions are serious and, even if they are correct, we would expect an explanation. A human would be able to tell you that your income is too low for a mortgage or that a specific cluster of cells is likely malignant. A model that provided similar explanations would be more useful than one that just provided predictions. By obtaining these explanations, we say we are interpreting a machine learning model.
Origin of the 'Black Beauty' meteorite is revealed
Scientists have revealed more about the origins of the famous'Black Beauty' meteorite, also known as NWA 7034. The researchers used AI to analyse thousands of high-resolution planetary images of the Martian surface from a range of Mars missions. They found Black Beauty was ejected into space when an asteroid impacted the planet's surface and created the six-mile-wide Karratha Crater 5-10 million years ago. Black Beauty, which weighs just 11 ounces (320 grams), led to the creation of a new class of meteorite when it was discovered in 2011 in the Western Sahara Desert. The meteorite was ejected from Mars' Karratha Crater 5-10 million years ago by an asteroid impact Five to ten million years ago an asteroid smashed into Mars.
Artificial Intelligence in Manufacturing Market Size to reach USD 78,744 Million by 2030 Exclusive Report by Acumen Research and Consulting
Acumen Research and Consulting recently published report titled "Artificial Intelligence in Manufacturing Market Size, Share, Analysis Report and Region Forecast, 2022 - 2030" BEIJING, July 11, 2022 (GLOBE NEWSWIRE) -- The Global Artificial Intelligence in Manufacturing Market size accounted for USD 2,963 Million in 2021 and is estimated to reach USD 78,744 Million by 2030. The rising volume of complex data sets is the leading factor boosting the global artificial intelligence in manufacturing market revenue. Our worldwide artificial intelligence in manufacturing industry analysis suggests that the manufacturers require artificial intelligence (AI) in their facilities due to the surging need for enhanced productivity and automation. AI is being used by manufacturers to enhance day-to-day operations, introduce new products, personalize designs, and forecast future financials. According to an MIT survey, about 60% of industry players are already using artificial intelligence.
Meta's AI-based Sphere 'may be the next big break in NLP'
Meta has open-sourced a machine-learning resource that could one day supplant Wikipedia as the world's biggest publicly available knowledge-verification database. Dubbed Sphere, it can be used to perform knowledge-intensive natural language processing, or KI-NLP, we're told. In practical terms, that means it can be used to answer complicated questions using natural language, and find sources for claims. A given example of its use is asking Sphere, "Who is Joรซlle Sambi Nzeba?" Wikipedia doesn't have an entry for her, but Sphere said she was "born in Belgium and grew up partly in Kinshasa (Congo). She currently lives in Brussels. She is a writer and slammer, alongside her activism in a feminist movement," and links to a website where it got that information about her work.
'AI Bumblebees:' These AI Robots Act Like Bees to Pollinate Tomato Plants
Is AI taking over the jobs of bumblebees? Bumblebees are typically used to pollinate plants in glasshouses all over the world. However, they are prohibited in Australia, so pollination must be done manually. Hence, prominent Australian fresh produce company Costa Group is deploying AI to implement robotic pollination in one of its tomato glasses, thanks to its partnership with Israeli firm Arugga AI Farming. The AI-powered robot is named "Polly" and will pollinate truss tomato plants in Costa's tomato glasshouse facilities in Guyra, New South Wales.
Can Machines Learn Morality? The Delphi Experiment
Jiang, Liwei, Hwang, Jena D., Bhagavatula, Chandra, Bras, Ronan Le, Liang, Jenny, Dodge, Jesse, Sakaguchi, Keisuke, Forbes, Maxwell, Borchardt, Jon, Gabriel, Saadia, Tsvetkov, Yulia, Etzioni, Oren, Sap, Maarten, Rini, Regina, Choi, Yejin
As AI systems become increasingly powerful and pervasive, there are growing concerns about machines' morality or a lack thereof. Yet, teaching morality to machines is a formidable task, as morality remains among the most intensely debated questions in humanity, let alone for AI. Existing AI systems deployed to millions of users, however, are already making decisions loaded with moral implications, which poses a seemingly impossible challenge: teaching machines moral sense, while humanity continues to grapple with it. To explore this challenge, we introduce Delphi, an experimental framework based on deep neural networks trained directly to reason about descriptive ethical judgments, e.g., "helping a friend" is generally good, while "helping a friend spread fake news" is not. Empirical results shed novel insights on the promises and limits of machine ethics; Delphi demonstrates strong generalization capabilities in the face of novel ethical situations, while off-the-shelf neural network models exhibit markedly poor judgment including unjust biases, confirming the need for explicitly teaching machines moral sense. Yet, Delphi is not perfect, exhibiting susceptibility to pervasive biases and inconsistencies. Despite that, we demonstrate positive use cases of imperfect Delphi, including using it as a component model within other imperfect AI systems. Importantly, we interpret the operationalization of Delphi in light of prominent ethical theories, which leads us to important future research questions.
Embedded analytics emerges to offer new level of business intelligence
Business analytics is an increasingly powerful tool for organisations, but one that is associated with steep learning curves and significant investments in infrastructure. The idea of using data to drive better decision-making is well established. But the conventional approach โ centred around reporting and analysis tools โ relies on specialist applications and highly trained staff. Often, firms find they have to build teams of data scientists to gather the data and manage the tools, and to build queries. This creates bottlenecks in the flow of information, as business units rely on specialist teams to interrogate the data, and to report back.
Overview of the Shared Task on Fake News Detection in Urdu at FIRE 2021
Amjad, Maaz, Butt, Sabur, Amjad, Hamza Imam, Zhila, Alisa, Sidorov, Grigori, Gelbukh, Alexander
Automatic detection of fake news is a highly important task in the contemporary world. This study reports the 2nd shared task called UrduFake@FIRE2021 on identifying fake news detection in Urdu. The goal of the shared task is to motivate the community to come up with efficient methods for solving this vital problem, particularly for the Urdu language. The task is posed as a binary classification problem to label a given news article as a real or a fake news article. The organizers provide a dataset comprising news in five domains: (i) Health, (ii) Sports, (iii) Showbiz, (iv) Technology, and (v) Business, split into training and testing sets. The training set contains 1300 annotated news articles -- 750 real news, 550 fake news, while the testing set contains 300 news articles -- 200 real, 100 fake news. 34 teams from 7 different countries (China, Egypt, Israel, India, Mexico, Pakistan, and UAE) registered to participate in the UrduFake@FIRE2021 shared task. Out of those, 18 teams submitted their experimental results, and 11 of those submitted their technical reports, which is substantially higher compared to the UrduFake shared task in 2020 when only 6 teams submitted their technical reports. The technical reports submitted by the participants demonstrated different data representation techniques ranging from count-based BoW features to word vector embeddings as well as the use of numerous machine learning algorithms ranging from traditional SVM to various neural network architectures including Transformers such as BERT and RoBERTa. In this year's competition, the best performing system obtained an F1-macro score of 0.679, which is lower than the past year's best result of 0.907 F1-macro. Admittedly, while training sets from the past and the current years overlap to a large extent, the testing set provided this year is completely different.
On Computing Relevant Features for Explaining NBCs
Izza, Yacine, Marques-Silva, Joao
Despite the progress observed with model-agnostic explainable AI (XAI), it is the case that model-agnostic XAI can produce incorrect explanations. One alternative are the so-called formal approaches to XAI, that include PI-explanations. Unfortunately, PI-explanations also exhibit important drawbacks, the most visible of which is arguably their size. The computation of relevant features serves to trade off probabilistic precision for the number of features in an explanation. However, even for very simple classifiers, the complexity of computing sets of relevant features is prohibitive. This paper investigates the computation of relevant sets for Naive Bayes Classifiers (NBCs), and shows that, in practice, these are easy to compute. Furthermore, the experiments confirm that succinct sets of relevant features can be obtained with NBCs.
ELLE: Efficient Lifelong Pre-training for Emerging Data
Qin, Yujia, Zhang, Jiajie, Lin, Yankai, Liu, Zhiyuan, Li, Peng, Sun, Maosong, Zhou, Jie
Current pre-trained language models (PLM) are typically trained with static data, ignoring that in real-world scenarios, streaming data of various sources may continuously grow. This requires PLMs to integrate the information from all the sources in a lifelong manner. Although this goal could be achieved by exhaustive pre-training on all the existing data, such a process is known to be computationally expensive. To this end, we propose ELLE, aiming at efficient lifelong pre-training for emerging data. Specifically, ELLE consists of (1) function preserved model expansion, which flexibly expands an existing PLM's width and depth to improve the efficiency of knowledge acquisition; and (2) pre-trained domain prompts, which disentangle the versatile knowledge learned during pre-training and stimulate the proper knowledge for downstream tasks. We experiment ELLE with streaming data from 5 domains on BERT and GPT. The results show the superiority of ELLE over various lifelong learning baselines in both pre-training efficiency and downstream performances. The codes are publicly available at https://github.com/thunlp/ELLE.