Media
Hallmarks of Human-Machine Collaboration: A framework for assessment in the DARPA Communicating with Computers Program
Kozierok, Robyn, Aberdeen, John, Clark, Cheryl, Garay, Christopher, Goodman, Bradley, Korves, Tonia, Hirschman, Lynette, McDermott, Patricia L., Peterson, Matthew W.
There is a growing desire to create computer systems that can communicate effectively to collaborate with humans on complex, open-ended activities. Assessing these systems presents significant challenges. We describe a framework for evaluating systems engaged in open-ended complex scenarios where evaluators do not have the luxury of comparing performance to a single right answer. This framework has been used to evaluate human-machine creative collaborations across story and music generation, interactive block building, and exploration of molecular mechanisms in cancer. These activities are fundamentally different from the more constrained tasks performed by most contemporary personal assistants as they are generally open-ended, with no single correct solution, and often no obvious completion criteria. We identified the Key Properties that must be exhibited by successful systems. From there we identified "Hallmarks" of success -- capabilities and features that evaluators can observe that would be indicative of progress toward achieving a Key Property. In addition to being a framework for assessment, the Key Properties and Hallmarks are intended to serve as goals in guiding research direction.
Interrogating the Black Box: Transparency through Information-Seeking Dialogues
Tubella, Andrea Aler, Theodorou, Andreas, Nieves, Juan Carlos
This paper is preoccupied with the following question: given a (possibly opaque) learning system, how can we understand whether its behaviour adheres to governance constraints? The answer can be quite simple: we just need to "ask" the system about it. We propose to construct an investigator agent to query a learning agent -- the suspect agent -- to investigate its adherence to a given ethical policy in the context of an information-seeking dialogue, modeled in formal argumentation settings. This formal dialogue framework is the main contribution of this paper. Through it, we break down compliance checking mechanisms into three modular components, each of which can be tailored to various needs in a vast amount of ways: an investigator agent, a suspect agent, and an acceptance protocol determining whether the responses of the suspect agent comply with the policy. This acceptance protocol presents a fundamentally different approach to aggregation: rather than using quantitative methods to deal with the non-determinism of a learning system, we leverage the use of argumentation semantics to investigate the notion of properties holding consistently. Overall, we argue that the introduced formal dialogue framework opens many avenues both in the area of compliance checking and in the analysis of properties of opaque systems.
Train your classifier first: Cascade Neural Networks Training from upper layers to lower layers
Zhang, Shucong, Do, Cong-Thanh, Doddipatla, Rama, Loweimi, Erfan, Bell, Peter, Renals, Steve
Although the lower layers of a deep neural network learn features which are transferable across datasets, these layers are not transferable within the same dataset. That is, in general, freezing the trained feature extractor (the lower layers) and retraining the classifier (the upper layers) on the same dataset leads to worse performance. In this paper, for the first time, we show that the frozen classifier is transferable within the same dataset. We develop a novel top-down training method which can be viewed as an algorithm for searching for high-quality classifiers. We tested this method on automatic speech recognition (ASR) tasks and language modelling tasks. The proposed method consistently improves recurrent neural network ASR models on Wall Street Journal, self-attention ASR models on Switchboard, and AWD-LSTM language models on WikiText-2.
M2FN: Multi-step Modality Fusion for Advertisement Image Assessment
Park, Kyung-Wha, Ha, Jung-Woo, Lee, JungHoon, Kwon, Sunyoung, Kim, Kyung-Min, Zhang, Byoung-Tak
Assessing advertisements, specifically on the basis of user preferences and ad quality, is crucial to the marketing industry. Although recent studies have attempted to use deep neural networks for this purpose, these studies have not utilized image-related auxiliary attributes, which include embedded text frequently found in ad images. We, therefore, investigated the influence of these attributes on ad image preferences. First, we analyzed large-scale real-world ad log data and, based on our findings, proposed a novel multi-step modality fusion network (M2FN) that determines advertising images likely to appeal to user preferences. Our method utilizes auxiliary attributes through multiple steps in the network, which include conditional batch normalization-based low-level fusion and attention-based high-level fusion. We verified M2FN on the AVA dataset, which is widely used for aesthetic image assessment, and then demonstrated that M2FN can achieve state-of-the-art performance in preference prediction using a real-world ad dataset with rich auxiliary attributes.
Korea's First Space Blockbuster Just Premiered on Netflix. It's a Blast.
After a year with no new major blockbusters, Jo Sung-hee's Space Sweepers arrives as a breath of fresh air. It's not a perfect movie, nor a particularly innovative one, but the science-fiction adventure--touted as the first Korean space blockbuster--is certainly fun, with colorful performances and impressive CGI, and a worthy substitute for a new Star Wars or Marvel movie. However, its presence in a year of absences isn't the only thing that makes it noteworthy. Unlike nearly all of the movies from those two dominant franchises, Space Sweepers is led by people of color. The main characters are a crew of Koreans, and the film is one of the rare space operas that doesn't posit that English has somehow become a universal language.
5 annoying Alexa and Amazon Echo settings you can change
Amazon announced that their voice assistant Alexa can now sign business agreements with health providers under the Health Insurance Portability and Accountability Act, or HIPAA. Third-party health developers can meet the rules that govern how sensitive health information is shared, and major health providers and companies have launched a number of voice programs to help users manage chronic conditions. If you have an Amazon Echo at home, you need to dive into the privacy settings. There are a few important things to lockdown. Don't forget to turn off voice purchasing if you never use that feature, or at least set up a PIN.
iPhone users will soon be able to change their default music app with Siri
After years of being forced to use Apple services on the iPhone by default, those restrictions are finally easing up a bit. As noticed by MacRumors earlier today, the iOS 14.5 beta appears to let you set third-party music services as default with Siri. This means you can ask Siri to play a particular song or album and it'll go straight to Spotify or YouTube Music. Currently, Siri only searches and plays things from whatever music you have in the Apple Music app, be it your own collection of songs or the Apple Music subscription catalog. It sounds like after iOS 14.5 is installed, Siri will ask you what music service you want to use when you ask it to play a song.
When a story is breaking, AI can help consumers identify fake news
Warnings about misinformation are now regularly posted on Twitter, Facebook, and other social media platforms, but not all of these cautions are created equal. New research from Rensselaer Polytechnic Institute shows that artificial intelligence can help form accurate news assessments--but only when a news story is first emerging. These findings were recently published in Computers in Human Behavior Reports by an interdisciplinary team of Rensselaer researchers. They found that AI-driven interventions are generally ineffective when used to flag issues with stories on frequently covered topics about which people have established beliefs, such as climate change and vaccinations. However, when a topic is so new that people have not had time to form an opinion, tailored AI-generated advice can lead readers to make better judgments regarding the legitimacy of news articles.
How True is GPT-2? An Empirical Analysis of Intersectional Occupational Biases
Kirk, Hannah, Jun, Yennie, Iqbal, Haider, Benussi, Elias, Volpin, Filippo, Dreyer, Frederic A., Shtedritski, Aleksandar, Asano, Yuki M.
The capabilities of natural language models trained on large-scale data have increased immensely over the past few years. Downstream applications are at risk of inheriting biases contained in these models, with potential negative consequences especially for marginalized groups. In this paper, we analyze the occupational biases of a popular generative language model, GPT-2, intersecting gender with five protected categories: religion, sexuality, ethnicity, political affiliation, and name origin. Using a novel data collection pipeline we collect 396k sentence completions of GPT-2 and find: (i) The machine-predicted jobs are less diverse and more stereotypical for women than for men, especially for intersections; (ii) Fitting 262 logistic models shows intersectional interactions to be highly relevant for occupational associations; (iii) For a given job, GPT-2 reflects the societal skew of gender and ethnicity in the US, and in some cases, pulls the distribution towards gender parity, raising the normative question of what language models _should_ learn.
Learning Synthetic Environments for Reinforcement Learning with Evolution Strategies
Ferreira, Fabio, Nierhoff, Thomas, Hutter, Frank
This work explores learning agent-agnostic synthetic environments (SEs) for Reinforcement Learning. SEs act as a proxy for target environments and allow agents to be trained more efficiently than when directly trained on the target environment. We formulate this as a bi-level optimization problem and represent an SE as a neural network. By using Natural Evolution Strategies and a population of SE parameter vectors, we train agents in the inner loop on evolving SEs while in the outer loop we use the performance on the target task as a score for meta-updating the SE population. We show empirically that our method is capable of learning SEs for two discrete-action-space tasks (CartPole-v0 and Acrobot-v1) that allow us to train agents more robustly and with up to 60% fewer steps. Not only do we show in experiments with 4000 evaluations that the SEs are robust against hyperparameter changes such as the learning rate, batch sizes and network sizes, we also show that SEs trained with DDQN agents transfer in limited ways to a discrete-action-space version of TD3 and very well to Dueling DDQN.