Education
#ectel2019 #mlearn2019 keynote @GeoffStead on #informal learning at scale #languages #AI
Geoff Stead (@geoffstead) takes the stage with a headset, a black shirt and walking like a fit Californian surfer (looking great). As chief product person of the Babbel language corporation, he talks about informal learning at scale and will offer insights. Well over 1 million subscribers (of which I am one - Spanish). Digital scale and reach Team of 10 people can start the magic of the web. How can we ensure Quality?
Government AI Readiness Index 2019 -- Oxford Insights -- Oxford Insights
Artificial intelligence (AI) technologies are forecast to add US$15 trillion to the global economy by 2030. According to the findings of our Index and as might be expected, the governments of countries in the Global North are better placed to take advantage of these gains than those in the Global South. There is a risk, therefore, that countries in the Global South could be left behind by the so-called fourth industrial revolution. Not only will they not reap the potential benefits of AI, but there is also the danger that unequal implementation widens global inequalities. AI has the power to transform the way that governments around the world deliver public services. In turn, this could greatly improve citizens' experiences of government. Governments are already implementing AI in their operations and service delivery, to improve efficiency, save time and money, and deliver better quality public services. In 2017, Oxford Insights created the world's first Government AI Readiness Index, to answer the question: how well placed are national governments to take advantage of the benefits of AI in their operations and delivery of public services? The results sought to capture the current capacity of governments to exploit the innovative potential of AI. The 2019 Government AI Readiness Index, produced with the support of the International Development Research Centre (IDRC), sees a development of our methodology, and an expansion of scope to cover all UN countries (from our previous group of OECD members). It scores the governments of 194 countries and territories according to their preparedness to use AI in the delivery of public services. The overall score is comprised of 11 input metrics, grouped under four high-level clusters: governance; infrastructure and data; skills and education; and government and public services. The data is derived from a variety of resources, ranging from our own desk research into AI strategies, to databases such as the number of registered AI startups on Crunchbase, to indices such as the UN eGovernment Development Index. We divided the countries by region, principally following UN groupings, with the chief exception of the Western European and Others Group, which we separated to allow more in-depth analysis of higher scoring governments.
Data Science and Machine Learning Bootcamp with R
Udemy Free Discount - Data Science and Machine Learning Bootcamp with R, Learn how to use the R programming language for data science and machine learning and data visualization! Data Scientist has been ranked the number one job on Glassdoor and the average salary of a data scientist is over $120,000 in the United States according to Indeed! Data Science is a rewarding career that allows you to solve some of the world's most interesting problems! This course is designed for both complete beginners with no programming experience or experienced developers looking to make the jump to Data Science! This comprehensive course is comparable to other Data Science bootcamps that usually cost thousands of dollars, but now you can learn all that information at a fraction of the cost!
Use an Azure Resource Manager template to create a workspace - Azure Machine Learning
While the template associated with this document creates a new Azure Container Registry, you can also create a new workspace without creating a container registry. If on container registry is present in the workspace, one will be created when you perform an operation that requires a container registry.
Artificial intelligence is paving the way for less invasive surgical training The McGill Tribune
Repeated practice is necessary to achieve mastery, which is no exception for surgical residents who often train directly on patients for four to six years. However, in this hands-on learning environment, even a minor mistake can be serious. To protect against such fatalities, a McGill research team constructed a solution. "The implementation of competency-based surgical education, along with advances in virtual reality, has resulted in the development and utilization of virtual reality-based surgical simulators," Rolando Del Maestro, professor emeritus in neuro-oncology at McGill, said in an interview with The McGill Tribune. The Neurosurgical Stimulation and Artificial Intelligence Learning Centre recently published a study in JAMA Network Open.
Denis Magda on Continuous Deep Learning with Apache Ignite
At the recent ApacheCon North America, Denis Magda spoke on continuous machine learning with Apache Ignite, an in-memory data grid. Ignite simplifies the machine-learning pipeline by performing training and hosting models in the same cluster that stores the data, and can perform "online" training to incrementally improve models when new data is available. Magda, vice-president of product management at GridGain, began by describing some of the pain points of machine learning on large datasets, in particular the latency involved in moving data across the network from its storage location to the processors that perform training. Models also have to be deployed into a production system after they are trained, and retrained periodically after new data is collected. Because Ignite runs code on the same computers that host data, it can train, deploy, and update a machine-learning model without a time-consuming extract-transform-load (ETL) step.
Language models and Automated Essay Scoring
Rodriguez, Pedro Uria, Jafari, Amir, Ormerod, Christopher M.
In this paper, we present a new comparative study on automatic essay scoring (AES). The current state-of-the-art natural language processing (NLP) neural network architectures are used in this work to achieve above human-level accuracy on the publicly available Kaggle AES dataset. We compare two powerful language models, BERT and XLNet, and describe all the layers and network architectures in these models. We elucidate the network architectures of BERT and XLNet using clear notation and diagrams and explain the advantages of transformer architectures over traditional recurrent neural network architectures. Linear algebra notation is used to clarify the functions of transformers and attention mechanisms. We compare the results with more traditional methods, such as bag of words (BOW) and long short term memory (LSTM) networks.
Uncovering Sociological Effect Heterogeneity using Machine Learning
Brand, Jennie E., Xu, Jiahui, Koch, Bernard, Geraldo, Pablo
Individuals do not respond uniformly to treatments, events, or interventions. Sociologists routinely partition samples into subgroups to explore how the effects of treatments vary by covariates like race, gender, and socioeconomic status. In so doing, analysts determine the key subpopulations based on theoretical priors. Data-driven discoveries are also routine, yet the analyses by which sociologists typically go about them are problematic and seldom move us beyond our expectations, and biases, to explore new meaningful subgroups. Emerging machine learning methods allow researchers to explore sources of variation that they may not have previously considered, or envisaged. In this paper, we use causal trees to recursively partition the sample and uncover sources of treatment effect heterogeneity. We use honest estimation, splitting the sample into a training sample to grow the tree and an estimation sample to estimate leaf-specific effects. Assessing a central topic in the social inequality literature, college effects on wages, we compare what we learn from conventional approaches for exploring variation in effects to causal trees. Given our use of observational data, we use leaf-specific matching and sensitivity analyses to address confounding and offer interpretations of effects based on observed and unobserved heterogeneity. We encourage researchers to follow similar practices in their work on variation in sociological effects.
Large e-retailer image dataset for visual search and product classification
Cdiscount, FranceAbstract Recent results of deep convolutional networks in visual recognition challenges open the path to a whole new set of disruptive user experiences such as visual search or recommendation. The list of companies offering this type of service is growing everyday but the adoption rate and the relevancy of results may vary a lot. We believe that the availability of large and diverse datasets is a necessary condition to improve the relevancy of such recommendation systems and facilitate their adoption. For that purpose, we wish to share with the community this dataset of more than 12M images of the 7M products of our online store classified into 5K categories. This original dataset is introduced in this article and several features are described. We also present some aspects of the winning solutions of our image classification challenge that was organized on the Kaggle platform around this set of images. 1 Introduction Recent advances in artificial intelligence and image recognition allow a whole new set of services to improve the Internet shopping experience [8, 20].
No-Regret Learning in Unknown Games with Correlated Payoffs
Sessa, Pier Giuseppe, Bogunovic, Ilija, Kamgarpour, Maryam, Krause, Andreas
We consider the problem of learning to play a repeated multi-agent game with an unknown reward function. Single player online learning algorithms attain strong regret bounds when provided with full information feedback, which unfortunately is unavailable in many real-world scenarios. Bandit feedback alone, i.e., observing outcomes only for the selected action, yields substantially worse performance. In this paper, we consider a natural model where, besides a noisy measurement of the obtained reward, the player can also observe the opponents' actions. This feedback model, together with a regularity assumption on the reward function, allows us to exploit the correlations among different game outcomes by means of Gaussian processes (GPs). We propose a novel confidence-bound based bandit algorithm GP-MW, which utilizes the GP model for the reward function and runs a multiplicative weight (MW) method. We obtain novel kernel-dependent regret bounds that are comparable to the known bounds in the full information setting, while substantially improving upon the existing bandit results. We experimentally demonstrate the effectiveness of GP-MW in random matrix games, as well as real-world problems of traffic routing and movie recommendation. In our experiments, GP-MW consistently outperforms several baselines, while its performance is often comparable to methods that have access to full information feedback.