Africa
Multilingual Content Moderation: A Case Study on Reddit
Ye, Meng, Sikka, Karan, Atwell, Katherine, Hassan, Sabit, Divakaran, Ajay, Alikhani, Malihe
Content moderation is the process of flagging content based on pre-defined platform rules. There has been a growing need for AI moderators to safeguard users as well as protect the mental health of human moderators from traumatic content. While prior works have focused on identifying hateful/offensive language, they are not adequate for meeting the challenges of content moderation since 1) moderation decisions are based on violation of rules, which subsumes detection of offensive speech, and 2) such rules often differ across communities which entails an adaptive solution. We propose to study the challenges of content moderation by introducing a multilingual dataset of 1.8 Million Reddit comments spanning 56 subreddits in English, German, Spanish and French. We perform extensive experimental analysis to highlight the underlying challenges and suggest related research problems such as cross-lingual transfer, learning under label noise (human biases), transfer of moderation models, and predicting the violated rule. Our dataset and analysis can help better prepare for the challenges and opportunities of auto moderation.
Towards Federated Learning on Time-Evolving Heterogeneous Data
Guo, Yongxin, Lin, Tao, Tang, Xiaoying
Federated Learning (FL) is a learning paradigm that protects privacy by keeping client data on edge devices. However, optimizing FL in practice can be difficult due to the diversity and heterogeneity of the learning system. Despite recent research efforts to improve the optimization of heterogeneous data, the impact of time-evolving heterogeneous data in real-world scenarios, such as changing client data or intermittent clients joining or leaving during training, has not been studied well. In this work, we propose Continual Federated Learning (CFL), a flexible framework for capturing the time-evolving heterogeneity of FL. CFL can handle complex and realistic scenarios, which are difficult to evaluate in previous FL formulations, by extracting information from past local data sets and approximating local objective functions. We theoretically demonstrate that CFL methods have a faster convergence rate than FedAvg in time-evolving scenarios, with the benefit depending on approximation quality. Through experiments, we show that our numerical findings match the convergence analysis and that CFL methods significantly outperform other state-of-the-art FL baselines.
Working with Long short-term memory models part1(Machine Learning 2023)
Abstract: The release of toxic gases by industries, emissions from vehicles, and an increase in the concentration of harmful gases and particulate matter in the atmosphere are all contributing factors to the deterioration of the quality of the air. Factors such as industries, urbanization, population growth, and the increased use of vehicles contribute to the rapid increase in pollution levels, which can adversely impact human health. This paper presents a model for forecasting the air quality index in Nigeria using the Bi-directional LSTM model. The air pollution data was downloaded from an online database (UCL). The dataset was pre-processed using both pandas tools in python.
ChatGPT Is the New Hook-Up Tool. Walk right up, sit right down, baby letโฆ
As you can see from this interview of Moonlair360, ultimate Content Creator Bro, ChatGPT is the new Tinder. This crazy dude, to whom I happen to be distantly related, has figured out the percentages to respond to every Tinder match with ChatGPT. I use it for news writing. ChatGPT can't replace investigative journalism, but it's perfect for the regurgitated not-so-happy meals put out as news on platforms that shall remain unnamed. They know who they are. I give my new best friend -- ChatGP T-- a topic on a local subject, and watch it spit out content like the coins from the slot machine that one time you hit it big.
Efficient and Flexible Topic Modeling using Pretrained Embeddings and Bag of Sentences
Pre-trained language models have led to a new state-of-the-art in many NLP tasks. However, for topic modeling, statistical generative models such as LDA are still prevalent, which do not easily allow incorporating contextual word vectors. They might yield topics that do not align very well with human judgment. In this work, we propose a novel topic modeling and inference algorithm. We suggest a bag of sentences (BoS) approach using sentences as the unit of analysis. We leverage pre-trained sentence embeddings by combining generative process models with clustering. We derive a fast inference algorithm based on expectation maximization, hard assignments, and an annealing process. Our evaluation shows that our method yields state-of-the art results with relatively little computational demands. Our methods is more flexible compared to prior works leveraging word embeddings, since it provides the possibility to customize topic-document distributions using priors. Code is at \url{https://github.com/JohnTailor/BertSenClu}.
AI in HCI Design and User Experience
The use of AI/ML capabilities for improving HCI/UX work and delivering better UX in solutions is becoming a trend (Abbas et al., 2022; Wu et al., 2019; Nikiforova et al., 2021) and creates many new opportunities for HCI/UX professionals (Holmquist, 2017; Yang et al., 2020). Some even speculate "AI/ML is the new UX" (Yang et al., 2018). Researchers proposed that AI can perform as an assistant, collaborator, researcher, or facilitator (Bertรฃo & Joo, 2021; Main & Grierson, 2020). AI technology will change the role of designers in the design process and generate an opportunity for creative collaboration between AI and designers (McCormack et al., 2020). Also, companies are moving fast to adopt AI for improving customer experience (CX).
Memory-assisted prompt editing to improve GPT-3 after deployment
Madaan, Aman, Tandon, Niket, Clark, Peter, Yang, Yiming
Large LMs such as GPT-3 are powerful, but can commit mistakes that are obvious to humans. For example, GPT-3 would mistakenly interpret "What word is similar to good?" to mean a homophone, while the user intended a synonym. Our goal is to effectively correct such errors via user interactions with the system but without retraining, which will be prohibitively costly. We pair GPT-3 with a growing memory of recorded cases where the model misunderstood the user's intents, along with user feedback for clarification. Such a memory allows our system to produce enhanced prompts for any new query based on the user feedback for error correction on similar cases in the past. On four tasks (two lexical tasks, two advanced ethical reasoning tasks), we show how a (simulated) user can interactively teach a deployed GPT-3, substantially increasing its accuracy over the queries with different kinds of misunderstandings by the GPT-3. Our approach is a step towards the low-cost utility enhancement for very large pre-trained LMs. Code, data, and instructions to implement MEMPROMPT for a new task at https://www.memprompt.com/.
MAILS -- Meta AI Literacy Scale: Development and Testing of an AI Literacy Questionnaire Based on Well-Founded Competency Models and Psychological Change- and Meta-Competencies
Carolus, Astrid, Koch, Martin, Straka, Samantha, Latoschik, Marc Erich, Wienrich, Carolin
The goal of the present paper is to develop and validate a questionnaire to assess AI literacy. In particular, the questionnaire should be deeply grounded in the existing literature on AI literacy, should be modular (i.e., including different facets that can be used independently of each other) to be flexibly applicable in professional life depending on the goals and use cases, and should meet psychological requirements and thus includes further psychological competencies in addition to the typical facets of AIL. We derived 60 items to represent different facets of AI Literacy according to Ng and colleagues conceptualisation of AI literacy and additional 12 items to represent psychological competencies such as problem solving, learning, and emotion regulation in regard to AI. For this purpose, data were collected online from 300 German-speaking adults. The items were tested for factorial structure in confirmatory factor analyses. The result is a measurement instrument that measures AI literacy with the facets Use & apply AI, Understand AI, Detect AI, and AI Ethics and the ability to Create AI as a separate construct, and AI Self-efficacy in learning and problem solving and AI Self-management. This study contributes to the research on AI literacy by providing a measurement instrument relying on profound competency models. In addition, higher-order psychological competencies are included that are particularly important in the context of pervasive change through AI systems.
Lebanon, Slovenia, UAE lead interest in AI Crypto
Lebanon, Slovenia, and the United Arab Emirates (UAE) are the top three countries that are most interested in Artificial Intelligence (AI) crypto, according to CoinGecko's recent report. Countries with major economic problems, like Nigeria, Sri Lanka, and Pakistan, have also ranked higher in the charts -- while the U.S. was placed 33rd, the CoinGecko report stated. The report measured the search popularity of 14 English search terms related to AI crypto between Nov. 30, 2022, and Feb. 16. A 100 indicates maximum popularity, while 50 indicates half -- zero would mean there was not enough data to examine. Lebanon scored 100 on almost all 14 search terms -- collecting 1,200 points and ranking first on the list.
One Startup's Plan to Help Africa Lure Back Its AI Talent
During a trip home to Johannesburg, South Africa, while completing an engineering master's program in Japan, Pelonomi Moiloa attended the largest machine learning community gathering she'd ever seen in Africa, just a few miles from where she grew up. In all, 600 people from 22 nations attended 2017's Deep Learning Indaba, held at the University of Witwatersrand, discussing topics like health care and agriculture solutions custom-made to meet the needs of African people. That week-long gathering made Moiloa feel she could have an impact on the lives of Africans, and it helped convince her to move back to South Africa and look for a way to put her engineering skills to work on her home continent. "The conversations were around making a genuine impact and positive change in African lives on a mass scale, and that was something I really wanted to be a part of," she says. This month, Moiloa will join some organizers of Deep Learning Indaba to launch Lelapa, a commercial and industrial AI research company focused on serving the needs of the 1 billion people in Africa.