Genre
Fairness Constraints: Mechanisms for Fair Classification
Zafar, Muhammad Bilal, Valera, Isabel, Rodriguez, Manuel Gomez, Gummadi, Krishna P.
Algorithmic decision making systems are ubiquitous across a wide variety of online as well as offline services. These systems rely on complex learning methods and vast amounts of data to optimize the service functionality, satisfaction of the end user and profitability. However, there is a growing concern that these automated decisions can lead, even in the absence of intent, to a lack of fairness, i.e., their outcomes can disproportionately hurt (or, benefit) particular groups of people sharing one or more sensitive attributes (e.g., race, sex). In this paper, we introduce a flexible mechanism to design fair classifiers by leveraging a novel intuitive measure of decision boundary (un)fairness. We instantiate this mechanism with two well-known classifiers, logistic regression and support vector machines, and show on real-world data that our mechanism allows for a fine-grained control on the degree of fairness, often at a small cost in terms of accuracy. A Python implementation of our mechanism is available at fate-computing.mpi-sws.org
Classification and regression using an outer approximation projection-gradient method
Barlaud, Michel, Belhajali, Wafa, Combettes, Patrick L., Fillatre, Lionel
This paper deals with sparse feature selection and grouping for classification and regression. The classification or regression problems under consideration consists in minimizing a convex empirical risk function subject to an $\ell^1$ constraint, a pairwise $\ell^\infty$ constraint, or a pairwise $\ell^1$ constraint. Existing work, such as the Lasso formulation, has focused mainly on Lagrangian penalty approximations, which often require ad hoc or computationally expensive procedures to determine the penalization parameter. We depart from this approach and address the constrained problem directly via a splitting method. The structure of the method is that of the classical gradient-projection algorithm, which alternates a gradient step on the objective and a projection step onto the lower level set modeling the constraint. The novelty of our approach is that the projection step is implemented via an outer approximation scheme in which the constraint set is approximated by a sequence of simple convex sets consisting of the intersection of two half-spaces. Convergence of the iterates generated by the algorithm is established for a general smooth convex minimization problem with inequality constraints. Experiments on both synthetic and biological data show that our method outperforms penalty methods.
"Influence Sketching": Finding Influential Samples In Large-Scale Regressions
Wojnowicz, Mike, Cruz, Ben, Zhao, Xuan, Wallace, Brian, Wolff, Matt, Luan, Jay, Crable, Caleb
There is an especially strong need in modern large-scale data analysis to prioritize samples for manual inspection. For example, the inspection could target important mislabeled samples or key vulnerabilities exploitable by an adversarial attack. In order to solve the "needle in the haystack" problem of which samples to inspect, we develop a new scalable version of Cook's distance, a classical statistical technique for identifying samples which unusually strongly impact the fit of a regression model (and its downstream predictions). In order to scale this technique up to very large and high-dimensional datasets, we introduce a new algorithm which we call "influence sketching." Influence sketching embeds random projections within the influence computation; in particular, the influence score is calculated using the randomly projected pseudo-dataset from the post-convergence Generalized Linear Model (GLM). We validate that influence sketching can reliably and successfully discover influential samples by applying the technique to a malware detection dataset of over 2 million executable files, each represented with almost 100,000 features. For example, we find that randomly deleting approximately 10% of training samples reduces predictive accuracy only slightly from 99.47% to 99.45%, whereas deleting the same number of samples with high influence sketch scores reduces predictive accuracy all the way down to 90.24%. Moreover, we find that influential samples are especially likely to be mislabeled. In the case study, we manually inspect the most influential samples, and find that influence sketching pointed us to new, previously unidentified pieces of malware.
On the Existence of Kernel Function for Kernel-Trick of k-Means
This paper corrects the proof of the Theorem 2 from the Gower's paper \cite[page 5]{Gower:1982}. The correction is needed in order to establish the existence of the kernel function used commonly in the kernel trick e.g. for $k$-means clustering algorithm, on the grounds of distance matrix. The scope of correction is explained in section 2.
1 in 4 Brits think robots would be better poloticians
Robots have helped increase revenue for numerous firms and a new survey has revealed that people feel the machines could do the same for the wealth of a nation. Approximately one in four UK citizens believe robots would make better decisions than elected human officials when it comes to boosting the economy. The research also uncovered that 66 percent of people foresee AI powered machines working in government positions by 2037. Approximately one in four UK citizens believe robots would make better decisions than human officials when it comes to boosting the economy. A robot run world is a fear among many people across the globe, as some worry about the future of their jobs. But according to a survey conducted by OpenText, a firm that provides management solutions, people living in the UK believe it could revolutionize their country.
How deep learning is transforming healthcare
Deep learning has been used to transform artificial intelligence (AI) development, whether it is from beating players in games like Go or poker to improving self-driving AI. But perhaps the most important changes for most of us is how AI advances and machine learning are affecting healthcare. In January, a medical startup won FDA approval for an AI-assisted cardiac imaging system called Arterys, and AI is playing vital roles in other health fields such as fighting cancer and aging. NVIDIA boasts that with deep learning, "AI can help doctors make faster, more accurate diagnoses. It can predict the risk of a disease in time to prevent it."
What Is The Best Way To Learn Machine Learning Without Taking Any Online Courses?
What is the best way to start learning machine learning and deep learning without taking any online courses? Let me first start off by saying that there is no single "best way" to learn machine learning, and you should find a system that works well for you. Some people prefer the structure of courses, others like reading books at their own pace, and some want to dive right into code. I started with Andrew Ng's Machine Learning Coursera course in 2012, knowing almost zero linear algebra and nothing about statistics or machine learning. Note that although the class covered neural networks, it was not a course on Deep Learning.
Getting Up Close and Personal with Algorithms
We hear the term "machine learning" a lot these days, usually in the context of predictive analysis and artificial intelligence. Machine learning is, more or less, a way for computers to learn things without being specifically programmed. But how does that actually happen? The answer is, in one word, algorithms. Algorithms are sets of rules that a computer is able to follow.
Big Data Analytics with SAS
The Fourth Industrial Revolution is upon us, even with the Third is still in progress. Big Data, Machine Learning and Artificial Intelligence are three of the driving forces behind it. While the term'Industrial Revolution' has always applied mainly to manufacturing, it now also involves service industries such as banking and insurance, who are investing heavily in Big Data to help them model credit risk, fraud, marketing success and other key data. Meanwhile manufacturing, retail, telco, pharma and many other sectors constantly need people skilled in building, analysing, monitoring and maintaining data models to gain strategic intelligence that helps them inform and adapt their key business processes. A leader in the world of Data Analytics is the SAS Institute, whose flagship product is SAS (Statistical Analysis System).
1 in 4 believe robots would make better politicians - Computer Business Review
Should Number 10 be worried about the impending AI revolution? The impending robot revolution has certainly got people talking – from the workplace to the car and home, robots and AI has really grabbed the attention of the public. However, setting aside Terminator-esque visions of the future, what do consumers really think of the impending AI revolution? Enterprise information management firm, OpenText, went and surveyed 2,000 UK consumers to find answers to that very question. Initial findings from the survey mirrored many other reports and surveys, with consumers expecting AI to impact the human workforce and their daily lives in general.