Goto

Collaborating Authors

 Government


Predictive Biases in Natural Language Processing Models: A Conceptual Framework and Overview

arXiv.org Artificial Intelligence

An increasing number of works in natural language processing have addressed the effect of bias on the predicted outcomes, introducing mitigation techniques that act on different parts of the standard NLP pipeline (data and models). However, these works have been conducted in isolation, without a unifying framework to organize efforts within the field. This leads to repetitive approaches, and puts an undue focus on the effects of bias, rather than on their origins. Research focused on bias symptoms rather than the underlying origins could limit the development of effective countermeasures. In this paper, we propose a unifying conceptualization: the predictive bias framework for NLP . We summarize the NLP literature and propose a general mathematical definition of predictive bias in NLP along with a conceptual framework, differentiating four main origins of biases: label bias, selection bias, model overamplification, and semantic bias . We discuss how past work has countered each bias origin. Our framework serves to guide an introductory overview of predictive bias in NLP, integrating existing work into a single structure and opening avenues for future research.


Modelling Bahdanau Attention using Election methods aided by Q-Learning

arXiv.org Machine Learning

Neural Machine Translation has lately gained a lot of "attention" with the advent of more and more sophisticated but drastically improved models. Attention mechanism has proved to be a boon in this direction by providing weights to the input words, making it easy for the decoder to identify words representing the present context. But by and by, as newer attention models with more complexity came into development, they involved large computation, making inference slow. In this paper, we have modelled the attention network using techniques resonating with social choice theory. Along with that, the attention mechanism, being a Markov Decision Process, has been represented by reinforcement learning techniques. Thus, we propose to use an election method ( k -Borda), fine-tuned using Q-learning, as a replacement for attention networks. The inference time for this network is less than a standard Bahdanau translator, and the results of the translation are comparable. This not only experimentally verifies the claims stated above but also helped provide a faster inference.


Optimal Experimental Design for Staggered Rollouts

arXiv.org Machine Learning

Experimentation has become an increasingly prevalent tool for guiding policy choices, firm decisions, and product innovation. A common hurdle in designing experiments is the lack of statistical power. In this paper, we study optimal multi-period experimental design under the constraint that the treatment cannot be easily removed once implemented; for example, a government or firm might implement treatment in different geographies at different times, where the treatment cannot be easily removed due to practical constraints. The design problem is to select which units to treat at which time, intending to test hypotheses about the effect of the treatment. When the potential outcome is a linear function of a unit effect, a time effect, and observed discrete covariates, we provide an analytically feasible solution to the design problem where the variance of the estimator for the treatment effect is at most 1+O(1/N^2) times the variance of the optimal design, where N is the number of units. This solution assigns units in a staggered treatment adoption pattern, where the proportion treated is a linear function of time. In the general setting where outcomes depend on latent covariates, we show that historical data can be utilized in the optimal design. We propose a data-driven local search algorithm with the minimax decision criterion to assign units to treatment times. We demonstrate that our approach improves upon benchmark experimental designs through synthetic experiments on real-world data sets from several domains, including healthcare, finance, and retail. Finally, we consider the case where the treatment effect changes with the time of treatment, showing that the optimal design treats a smaller fraction of units at the beginning and a greater share at the end.


An artificial intelligence company backed by Microsoft is helping Israel surveil Palestinians

#artificialintelligence

An Israeli startup invested in heavily by American companies, including Microsoft, produces facial recognition software used to conduct biometric surveillance on Palestinians, investigations by NBC and Haaretz revealed. In June, Microsoft -- which has touted its framework for ethical use of facial recognition -- joined a group investment of $78 million to AnyVision, an international tech company based in Israel. One of AnyVision's flagship products is Better Tomorrow, a program that allows the tracking of objects and people on live video feeds, even tracking between independent camera feeds. AnyVision's facial recognition software is at the heart of a military mass surveillance project in the West Bank, according to the NBC and Haaretz reporting. An Israeli Defense Forces statement in February acknowledged the addition of facial recognition verification technology to at least 27 checkpoints between Israel and the West Bank to "upgrade the crossings" and, in an effort to "deter terror attacks," rapidly installed a network of over 1,700 cameras across the occupied territories.


Artificial Intelligence Comes to -- Wait for It -- the Post Office -- The Financial Revolutionist

#artificialintelligence

When you think of innovative organizations, the U.S. Postal Service is probably not at the top of your list. In fact, the post office has sort of become a metaphor for an old, inefficient way of doing things -- think of the term "snail mail." Do people even buy stamps anymore? And when was the last time you, dear reader, stepped foot in a post office? Well, many people still do.


Artificial intelligence warning: AI deemed 'too dangerous' released into the world

#artificialintelligence

Such misuses would require the public to become more critical about the text they consume, which could have been generated by artificial intelligence, they said. The researcher wrote: "These findings, combined with earlier results on synthetic imagery, audio, and video, imply that technologies are reducing the cost of generating fake content and waging disinformation campaigns. "The public at large will need to become more skeptical of text they find online, just as the'deep fakes' phenomenon calls for more skepticism about images."


Eye in the sky: Japan seeks AI-guided surveillance for patrol planes

#artificialintelligence

Japan will begin research on using artificial intelligence to bolster surveillance by naval patrol aircraft, as a changing national security environment forces the Self-Defense Forces to take on wider roles despite a personnel shortage. The AI would help ascertain whether a target spotted by conventional radar is an enemy vessel or some other threat. Machine learning through previous data would be used to develop the ability to identify a vessel from images that are difficult for the human eye to decipher. Currently, radar data converted to black-and-white images are scrutinized by experienced SDF personnel. The Defense Ministry will use a budget of about 900 million yen ($8.25 million) for development in the fiscal year starting in April, with the goal of outfitting Maritime SDF patrol planes with the technology as early as fiscal 2024.


Water, water everywhere!

#artificialintelligence

Several cities in Quebec including Gatineau, Montreal, and Rimouski, as well as Windsor, London and Thunder Bay, ON and Halifax, NS, have been participating in the project. Renato explains, that the reasons pipes break include frost, aging, as well as soil corrosion. However, a very important, yet often overlooked problem, is pressure build-up in the system -- with too much variation, pressure weakens pipes. Renato is developing a method to model pressure, which is not currently used in modeling predictions. He adds that this type of modelling can be of tremendous value to municipalities.


How Is the Medical Industry Using Deep Learning?

#artificialintelligence

The healthcare industry continues to be a major driver of the U.S. economy, with more than $3.5 trillion spent on healthcare in 2018 alone. Researchers believe that the industry will contribute more than $5.6 trillion to the economy by 2025. Much of this revenue comes from the medical research field, which is responsible for improving drug research, disease diagnosis and treatment protocols. Major research companies are collaborating with software development services to integrate deep learning technology into their investigations. Deep learning promises to transform the way that doctors review medical tests and make diagnoses, helping them identify diseases and start treatment quicker.


NIOSH Competition Looks for the Best AI Safety Solutions - EHS Daily Advisor

#artificialintelligence

The National Institute for Occupational Safety and Health (NIOSH) announced a competition for programmers to develop artificial intelligence (AI) capable of analyzing safety reports and assigning occupational safety and health classification codes. Submissions are due by November 21. When a worker is injured or becomes ill on the job, a person records free-form narrative text explaining how the injury or illness occurred. Currently, human evaluators read the reports and assign codes classifying injuries and illnesses. Reports can contain large amounts of information, and assigning codes is a costly and time-consuming task subject to human error.