Africa
Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection
Gururangan, Suchin, Card, Dallas, Dreier, Sarah K., Gade, Emily K., Wang, Leroy Z., Wang, Zeyu, Zettlemoyer, Luke, Smith, Noah A.
Language models increasingly rely on massive web dumps for diverse text data. However, these sources are rife with undesirable content. As such, resources like Wikipedia, books, and newswire often serve as anchors for automatically selecting web text most suitable for language modeling, a process typically referred to as quality filtering. Using a new dataset of U.S. high school newspaper articles -- written by students from across the country -- we investigate whose language is preferred by the quality filter used for GPT-3. We find that newspapers from larger schools, located in wealthier, educated, and urban ZIP codes are more likely to be classified as high quality. We then demonstrate that the filter's measurement of quality is unaligned with other sensible metrics, such as factuality or literary acclaim. We argue that privileging any corpus as high quality entails a language ideology, and more care is needed to construct training corpora for language models, with better transparency and justification for the inclusion or exclusion of various texts.
x-prize
Since proposed in 1965, Lotfi Zadeh's idea of "fuzzy logic" has penetrated every science and engineering field. Physix is a "fuzzy metric system" to measure people's opinion and emotional response more accurately. AI lacks a logic that that can interpret the subtleties of thought, emotion and communication. The continua provide simple, unbiased metrics of personal experience and belief for any given topic; every opinion and feeling can be seen relative to others'. A shared visualization of perspectives will defuse antagonism and reduce racism leaving a positive impact on our society.
Future developments in AI could make your credit score obsolete
Did you miss a session from the Future of Work Summit? This article was contributed by Frederik Bussler, consultant and analyst. Around one in four American adults are underbanked, meaning they are underserved by traditional finance, and rely on high-fee alternative financial systems. For underbanked Americans, getting a loan or a credit card can range between being either difficult or next to impossible. For those who do have a credit score, it's often not a very high one.
InstaDeep raises $100M to inject enterprise decision-making with AI
AI has the potential to generate meaningful returns for the enterprise. Responding to a 2018 PricewaterhouseCoopers survey, 54% of business executives say that their adoption of AI within the workplace has led to a boost in productivity. A separate 2019 McKinsey report found that 44% of firms using AI achieved a reduction in business costs in departments where AI is implemented. But barriers stand in the way of deployment, including a lack of production-grade data and expensive tools and development processes. Among the top challenges enterprises face in adopting AI is an absence of in-house talent.
Artificial Intelligence, Machine Learning, and the Fight Against World Hunger
According to the World Health Organization (WHO), the world is going hungry. WHO data shows that in 2018, the most recent year for which data is available, 820 million people lacked enough food to eat, an increase of nine million people over the year before. Hunger kills plenty of people worldwide. It also impacts those who survive, causing serious childhood development issues like stunting, where children are too short for their age, and wasting, where they're too thin for their age. The explosion in our planet's population is a major factor in there not being enough food to go around.
AI May Soon Be Able to Read Your Emotions
Artificial intelligence (AI) may soon know more about you than you think. A startup called Hume AI claims to use algorithms to measure emotions from facial, vocal, and verbal expressions. It's one of a growing number of companies that purport to read human emotions using computers. But some experts say that the concept raises privacy issues. "Whoever controls these systems and platforms are going to have a lot of information on individuals," Bob Bilbruck, a tech startup advisor, told Lifewire in an email interview.
Prediction of Neonatal Respiratory Distress in Term Babies at Birth from Digital Stethoscope Recorded Chest Sounds
Grooby, Ethan, Sitaula, Chiranjibi, Tan, Kenneth, Zhou, Lindsay, King, Arrabella, Ramanathan, Ashwin, Malhotra, Atul, Dumont, Guy A., Marzbanrad, Faezeh
Neonatal respiratory distress is a common condition that if left untreated, can lead to short- and long-term complications. This paper investigates the usage of digital stethoscope recorded chest sounds taken within 1min post-delivery, to enable early detection and prediction of neonatal respiratory distress. Fifty-one term newborns were included in this study, 9 of whom developed respiratory distress. For each newborn, 1min anterior and posterior recordings were taken. These recordings were pre-processed to remove noisy segments and obtain high-quality heart and lung sounds. The random undersampling boosting (RUSBoost) classifier was then trained on a variety of features, such as power and vital sign features extracted from the heart and lung sounds. The RUSBoost algorithm produced specificity, sensitivity, and accuracy results of 85.0%, 66.7% and 81.8%, respectively.
Towards Objective Metrics for Procedurally Generated Video Game Levels
Beukman, Michael, James, Steven, Cleghorn, Christopher
With increasing interest in procedural content generation by academia and game developers alike, it is vital that different approaches can be compared fairly. However, evaluating procedurally generated video game levels is often difficult, due to the lack of standardised, game-independent metrics. In this paper, we introduce two simulation-based evaluation metrics that involve analysing the behaviour of an A* agent to measure the diversity and difficulty of generated levels in a general, game-independent manner. Diversity is calculated by comparing action trajectories from different levels using the edit distance, and difficulty is measured as how much exploration and expansion of the A* search tree is necessary before the agent can solve the level. We demonstrate that our diversity metric is more robust to changes in level size and representation than current methods and additionally measures factors that directly affect playability, instead of focusing on visual information. The difficulty metric shows promise, as it correlates with existing estimates of difficulty in one of the tested domains, but it does face some challenges in the other domain. Finally, to promote reproducibility, we publicly release our evaluation framework.
Probability estimation and structured output prediction for learning preferences in last mile delivery
Canoy, Rocsildes, Bucarey, Victor, Molenbruch, Yves, Mulamba, Maxime, Mandi, Jayanta, Guns, Tias
We study the problem of learning the preferences of drivers and planners in the context of last mile delivery. Given a data set containing historical decisions and delivery locations, the goal is to capture the implicit preferences of the decision-makers. We consider two ways to use the historical data: one is through a probability estimation method that learns transition probabilities between stops (or zones). This is a fast and accurate method, recently studied in a VRP setting. Furthermore, we explore the use of machine learning to infer how to best balance multiple objectives such as distance, probability and penalties. Specifically, we cast the learning problem as a structured output prediction problem, where training is done by repeatedly calling the TSP solver. Another important aspect we consider is that for last-mile delivery, every address is a potential client and hence the data is very sparse. Hence, we propose a two-stage approach that first learns preferences at the zone level in order to compute a zone routing; after which a penalty-based TSP computes the stop routing. Results show that the zone transition probability estimation performs well, and that the structured output prediction learning can improve the results further. We hence showcase a successful combination of both probability estimation and machine learning, all the while using standard TSP solvers, both during learning and to compute the final solution; this means the methodology is applicable to other, real-life, TSP variants, or proprietary solvers.