Goto

Collaborating Authors

 Education


Guide to Encoding Categorical Features Using Scikit-Learn For Machine Learning

#artificialintelligence

One of the most crucial preprocessing steps in any machine learning project is feature encoding. It is the process of turning categorical data in a dataset into numerical data. It is essential that we perform feature encoding because most machine learning models can only interpret numerical data and not data in text form. As usual, I will demonstrate these concepts through a practical case study using the students' performance in exams dataset on Kaggle. You can find the complete notebook up on my GitHub here.


ParsiNLU: A Suite of Language Understanding Challenges for Persian

arXiv.org Artificial Intelligence

Despite the progress made in recent years in addressing natural language understanding (NLU) challenges, the majority of this progress remains to be concentrated on resource-rich languages like English. This work focuses on Persian language, one of the widely spoken languages in the world, and yet there are few NLU datasets available for this rich language. The availability of high-quality evaluation datasets is a necessity for reliable assessment of the progress on different NLU tasks and domains. We introduce ParsiNLU, the first benchmark in Persian language that includes a range of high-level tasks -- Reading Comprehension, Textual Entailment, etc. These datasets are collected in a multitude of ways, often involving manual annotations by native speakers. This results in over 14.5$k$ new instances across 6 distinct NLU tasks. Besides, we present the first results on state-of-the-art monolingual and multi-lingual pre-trained language-models on this benchmark and compare them with human performance, which provides valuable insights into our ability to tackle natural language understanding challenges in Persian. We hope ParsiNLU fosters further research and advances in Persian language understanding.


Progressive Network Grafting for Few-Shot Knowledge Distillation

arXiv.org Artificial Intelligence

Knowledge distillation has demonstrated encouraging performances in deep model compression. Most existing approaches, however, require massive labeled data to accomplish the knowledge transfer, making the model compression a cumbersome and costly process. In this paper, we investigate the practical few-shot knowledge distillation scenario, where we assume only a few samples without human annotations are available for each category. To this end, we introduce a principled dual-stage distillation scheme tailored for few-shot data. In the first step, we graft the student blocks one by one onto the teacher, and learn the parameters of the grafted block intertwined with those of the other teacher blocks. In the second step, the trained student blocks are progressively connected and then together grafted onto the teacher network, allowing the learned student blocks to adapt themselves to each other and eventually replace the teacher network. Experiments demonstrate that our approach, with only a few unlabeled samples, achieves gratifying results on CIFAR10, CIFAR100, and ILSVRC-2012. On CIFAR10 and CIFAR100, our performances are even on par with those of knowledge distillation schemes that utilize the full datasets. The source code is available at https://github.com/zju-vipa/NetGraft.


Generating Adversarial Disturbances for Controller Verification

arXiv.org Machine Learning

We consider the problem of generating maximally adversarial disturbances for a given controller assuming only blackbox access to it. We propose an online learning approach to this problem that adaptively generates disturbances based on control inputs chosen by the controller. The goal of the disturbance generator is to minimize regret versus a benchmark disturbance-generating policy class, i.e., to maximize the cost incurred by the controller as well as possible compared to the best possible disturbance generator in hindsight (chosen from a benchmark policy class). In the setting where the dynamics are linear and the costs are quadratic, we formulate our problem as an online trust region (OTR) problem with memory and present a new online learning algorithm (MOTR) for this problem. We prove that this method competes with the best disturbance generator in hindsight (chosen from a rich class of benchmark policies that includes linear-dynamical disturbance generating policies). We demonstrate our approach on two simulated examples: (i) synthetically generated linear systems, and (ii) generating wind disturbances for the popular PX4 controller in the AirSim simulator. On these examples, we demonstrate that our approach outperforms several baseline approaches, including $H_{\infty}$ disturbance generation and gradient-based methods.


Machine Learning and AI - What Does The Future Hold?

#artificialintelligence

By 2021, one in four forward-thinking enterprises will push AI to new frontiers, such as holographic meetings for remote work and on-demand personalised manufacturing, as per new predictions by Forrester Research. Even today, all of us are subconsciously using Machine Learning in our daily lives. Wish to stay home and yet be social? A nascent domain that's roughly 60 years old, has changed the way humans and machines perform, that's for sure. AI will create 2.3 million jobs in 2020.


UP govt ties up with US university for AI, machine learning courses

#artificialintelligence

Lucknow: The UP government has collaborated with Austin University of the US for courses in artificial intelligence (AI), machine learning for cybersecurity management and data analytics for students of the state. The students will get a certificate or degree from the Austin University, recognized in over 65 countries and the course fees would be shared by the university and the UP government. Micro, small and medium enterprises (MSME) and export promotion minister, Siddharth Nath Singh told TOI that an MoU for collaboration between the technical education department of the state government and the Austin University has been signed on Wednesday. "This is a game changer in the field of higher and technical education as students will get a certificate from one of the best US universities while residing in any UP city. The certificate will make them useful for global brands," Singh said, adding that the North Californian university is producing over 30,000 students from its technical institutes every year and as compared to the progressive states of the country, UP lacked such advanced courses.


Teaching HAII

#artificialintelligence

Human-AI Interaction (HAII) was taught twice in a fully remote fashion over 12 weeks at Williams College to only undergraduates. The course was organized around 11.5 modules (i.e., topics), and each module included: 2x pre-recorded lecture videos, 2-3 readings (1 research paper 2 popular media), a 7 question quiz on that module's materials, a 60 minute synchronous [remote] class meeting with 8 students, followed by discussion forum posts and 2x responses to peers. The course was implemented in Canvas, referred to as "GLOW" at Williams, but I have made available some versions of the materials via Google documents on this website (apologies for any formatting issues!). Context for each course component is below, but details can be found on the Schedule and Syllabus pages. Lectures: Pre-recorded lectures for a module were posted twice per week (Thursdays & Mondays), which were recorded in 15 minute sections and linked together in a YouTube playlist for no more than 50 minutes total, although usually around 30-40 minutes.


The Peggy Smedley Show: The Age of Screeners

#artificialintelligence

Peggy talks about the young, up-and-coming generation she has coined—the Screeners—which is being taught and raised by computers. She spends some time postulating if these Screeners have to be self-taught or if their education will have to come from the likes of corporations that offer training since the universities have made it difficult to go to school. She also discusses: The percentage of Screeners owning smartphones already? If postsecondary institution enrollment is going up or down? Examples of how technology is disrupting the traditional education model. (12.08.20 - #698) IoT, Internet of Things, Peggy Smedley, artificial intelligence, machine learning, big data, digital transformation, cybersecurity, blockchain, 5G cloud, sustainability, future of work, podcast


Machine learning helps to map invasive plant from space

AIHub

Researchers from CSIRO, Charles Darwin University and The University of Western Australia have developed a machine-learning approach that reliably detects invasive gamba grass from high-resolution satellite imagery. Gamba grass is listed as a Weed of National Significance, and is one of five introduced grass species that pose extensive and significant threats to Australia's biodiversity. The perennial grass can grow to four metres in height and forms dense tussocks which can burn as large, hot fires late in the dry season. Mapping where gamba grass occurs is essential to managing it effectively, but northern Australia is so vast and remote that on-the-ground mapping and even airborne detection of the weed is too labour-intensive. So, the researchers turned to high-quality satellite imagery and developed a technique that could help detect and prioritise gamba grass for management.


DeepMind offer four scholarships at Oxford University for under-represented students

Oxford Comp Sci

One new DPhil scholarship will also be funded, for a student undertaking either a DPhil in Engineering Science, a DPhil in Computer Science, or in the Autonomous Intelligent Machines and Systems EPSRC Centre of Doctoral Training (CDT). These scholarships are open to UK applicants as identified by the same criteria listed above. Overseas applications to the PhD are open to those from any of the categories listed (above) across the Masters, both overseas and home.