Goto

Collaborating Authors

 Education


Fairness implications of encoding protected categorical attributes

arXiv.org Machine Learning

Protected attributes are often presented as categorical features that need to be encoded before feeding them into a machine learning algorithm. Encoding these attributes is paramount as they determine the way the algorithm will learn from the data. Categorical feature encoding has a direct impact on the model performance and fairness. In this work, we compare the accuracy and fairness implications of the two most well-known encoders: one-hot encoding and target encoding. We distinguish between two types of induced bias that can arise while using these encodings and can lead to unfair models. The first type, irreducible bias, is due to direct group category discrimination and a second type, reducible bias, is due to large variance in less statistically represented groups. We take a deeper look into how regularization methods for target encoding can improve the induced bias while encoding categorical features. Furthermore, we tackle the problem of intersectional fairness that arises when mixing two protected categorical features leading to higher cardinality. This practice is a powerful feature engineering technique used for boosting model performance. We study its implications on fairness as it can increase both types of induced bias


SafeAPT: Safe Simulation-to-Real Robot Learning using Diverse Policies Learned in Simulation

arXiv.org Artificial Intelligence

The framework of Simulation-to-real learning, i.e, learning policies in simulation and transferring those policies to the real world is one of the most promising approaches towards data-efficient learning in robotics. However, due to the inevitable reality gap between the simulation and the real world, a policy learned in the simulation may not always generate a safe behaviour on the real robot. As a result, during adaptation of the policy in the real world, the robot may damage itself or cause harm to its surroundings. In this work, we introduce a novel learning algorithm called SafeAPT that leverages a diverse repertoire of policies evolved in the simulation and transfers the most promising safe policy to the real robot through episodic interaction. To achieve this, SafeAPT iteratively learns a probabilistic reward model as well as a safety model using real-world observations combined with simulated experiences as priors. Then, it performs Bayesian optimization on the repertoire with the reward model while maintaining the specified safety constraint using the safety model. SafeAPT allows a robot to adapt to a wide range of goals safely with the same repertoire of policies evolved in the simulation. We compare SafeAPT with several baselines, both in simulated and real robotic experiments and show that SafeAPT finds high-performance policies within a few minutes in the real world while minimizing safety violations during the interactions.


A Survey on Visual Transfer Learning using Knowledge Graphs

arXiv.org Artificial Intelligence

Recent approaches of computer vision utilize deep learning methods as they perform quite well if training and testing domains follow the same underlying data distribution. However, it has been shown that minor variations in the images that occur when using these methods in the real world can lead to unpredictable errors. Transfer learning is the area of machine learning that tries to prevent these errors. Especially, approaches that augment image data using auxiliary knowledge encoded in language embeddings or knowledge graphs (KGs) have achieved promising results in recent years. This survey focuses on visual transfer learning approaches using KGs. KGs can represent auxiliary knowledge either in an underlying graph-structured schema or in a vector-based knowledge graph embedding. Intending to enable the reader to solve visual transfer learning problems with the help of specific KG-DL configurations we start with a description of relevant modeling structures of a KG of various expressions, such as directed labeled graphs, hypergraphs, and hyper-relational graphs. We explain the notion of feature extractor, while specifically referring to visual and semantic features. We provide a broad overview of knowledge graph embedding methods and describe several joint training objectives suitable to combine them with high dimensional visual embeddings. The main section introduces four different categories on how a KG can be combined with a DL pipeline: 1) Knowledge Graph as a Reviewer; 2) Knowledge Graph as a Trainee; 3) Knowledge Graph as a Trainer; and 4) Knowledge Graph as a Peer. To help researchers find evaluation benchmarks, we provide an overview of generic KGs and a set of image processing datasets and benchmarks including various types of auxiliary knowledge. Last, we summarize related surveys and give an outlook about challenges and open issues for future research.


Reasoning Like Program Executors

arXiv.org Artificial Intelligence

Reasoning over natural language is a long-standing goal for the research community. However, studies have shown that existing language models are inadequate in reasoning. To address the issue, we present POET, a new pre-training paradigm. Through pre-training language models with programs and their execution results, POET empowers language models to harvest the reasoning knowledge possessed in program executors via a data-driven approach. POET is conceptually simple and can be instantiated by different kinds of programs. In this paper, we show three empirically powerful instances, i.e., POET-Math, POET-Logic, and POET-SQL. Experimental results on six benchmarks demonstrate that POET can significantly boost model performance on natural language reasoning, such as numerical reasoning, logical reasoning, and multi-hop reasoning. Taking the DROP benchmark as a representative example, POET improves the F1 metric of BART from 69.2% to 80.6%. Furthermore, POET shines in giant language models, pushing the F1 metric of T5-11B to 87.6% and achieving a new state-of-the-art performance on DROP. POET opens a new gate on reasoning-enhancement pre-training and we hope our analysis would shed light on the future research of reasoning like program executors.


Career roadmap: Machine learning engineer

#artificialintelligence

It stands to reason, then, that machine learning engineers are in good place as far as career outlook. These professionals are sophisticated programmers who develop machines and systems that can learn and apply knowledge without specific direction, according to Study.com, an online education platform. The focus of machine learning engineers goes beyond specifically programming machines to perform specific tasks, Study.com They create programs that allow machines to take actions without being specifically directed to perform the tasks. Such an engineer might work on the development of a self-driving vehicle, for example, or program services in such a way that they can attempt to identify a specific individual's interests.


Physical systems perform machine-learning computations

#artificialintelligence

You may not be able to teach an old dog new tricks, but Cornell researchers have found a way to train physical systems, ranging from computer speakers and lasers to simple electronic circuits, to perform machine-learning computations, such as identifying handwritten numbers and spoken vowel sounds. Cornell researchers have successfully trained (from left to right) a computer speaker, a simple electronic circuit and a laser to perform machine-learning computations. The experiment is no mere stunt or parlor trick. By turning these physical systems into the same kind of neural networks that drive services like Google Translate and online searches, the researchers have demonstrated an early but viable alternative to conventional electronic processors โ€“ one with the potential to be orders of magnitude faster and more energy efficient than the power-gobbling chips in data centers and server farms that support many artificial-intelligence applications. "Many different physical systems have enough complexity in them that they can perform a large range of computations," said Peter McMahon, assistant professor of applied and engineering physics in the College of Engineering, who led the project.


Defending Human Rights in the Age of Artificial Intelligence

#artificialintelligence

Whether you've used social media, a navigation app or a picture filter, chances are that Artificial Intelligence (AI) has impacted you. It's not just you โ€” AI is impacting human rights worldwide, and this course will inform and educate you on how your rights are affected by AI, and how you can be empowered to guard these rights. UNESCO and UNITAR jointly launched a new, short online learning course on AI and Human Rights for youths aged 16 to 24. Experts break down complex concepts about AI into straight forward activities built around our daily technology interactions. The course focuses on how freedom of expression, right to privacy and the right to equality are impacted using AI.


How to become a data scientist: A guide to the education, skills, and necessary experience - Fortune

#artificialintelligence

It may come as a surprise that the title of โ€œdata scientistโ€ is relatively newโ€”in fact, it was coined in 2008 by two data analytics professionals at LinkedIn and Facebook. Today, we know it as a fast-growing field, but the term and career really only took shape after the arrival of big tech and the [โ€ฆ]


45-Days Data Science Bootcamp: Build 45 Real Life Projects

#artificialintelligence

Data science plays an important role in virtually all aspects of business operations and strategies. For example, it provides information about customers that helps companies create stronger marketing campaigns and targeted advertising to increase product sales. It aids in managing financial risks, detecting fraudulent transactions, and preventing equipment breakdowns in manufacturing plants and other industrial settings. It helps block cyber-attacks and other security threats in IT systems. We'll cover everything you need to know for the full data science and machine learning tech stack required at the world's top companies.


University of North Carolina, Chapel Hill: Grant will expand University Libraries' use of machine learning to identify historically racist laws

#artificialintelligence

Since 2019, experts at the University of North Carolina at Chapel Hill's University Libraries have investigated the use of machine learning to identify racist laws from North Carolina's past. Now a grant of $400,000 from The Andrew W. Mellon Foundation will allow them to extend that work to two more states. The grant will also fund research and teaching fellowships for scholars interested in using the project's outputs and techniques. On the Books: Jim Crow and Algorithms of Resistance began with a question from a North Carolina social studies teacher: Was there a comprehensive list of all the Jim Crow laws that had ever been passed in the state? Finding little beyond scholar and activist Pauli Murray's 1951 book "States' laws on race and color," a team of librarians, technologists and data experts set out to fill the gap.