Media
MIT's Automatic Data-Driven Media Bias Measurement Method Achieves Human-Level Results
Today more than ever, people are voicing concerns regarding biases in news media. Especially in the political arena, there are accusations of favouritism or disfavour in reporting, often expressed through the emphasizing or ignoring of certain political actors, policies, events, or topics. Is it possible to develop objective and transparent data-driven methods to identify such biases, rather than relying on subjective human judgements? MIT researchers Samantha D'Alonzo and Max Tegmark say "yes," and have proposed an automated method for measuring media bias. The proposed data-driven approach produces results that are in close accordance with human-judgement classifications on left-right and establishment biases.
Ex-Google exec describes 4 top dangers of artificial intelligence
California's Senate last week advanced a bill that would force Amazon (AMZN) to reveal details behind the productivity-tracking algorithm used in its warehouses; meanwhile, Facebook (FB) this week faced criticism over a Wall Street Journal report finding it knows its Instagram feed makes some teenage girls feel worse about themselves. These developments make up a backlash not necessarily against big tech, so much as its algorithms, which use artificial intelligence (AI) to adapt performance for individual users or employees. In a new interview, AI expert Kai-Fu Lee -- who worked as an executive at Google (GOOG, GOOGL), Apple (AAPL), and Microsoft (MSFT) -- explained the top four dangers of burgeoning AI technology: externalities, personal data risks, inability to explain consequential choices, and warfare. "The single largest danger is autonomous weapons," he says. "That's when AI can be trained to kill, and more specifically trained to assassinate," adds Lee, the co-author of a new book entitled "AI 2041: Ten Visions for Our Future."
MIT: Measuring Media Bias in Major News Outlets With Machine Learning
A study from MIT has used machine learning techniques to identify biased phrasing across around 100 of the largest and most influential news outlets in the US and beyond, including 83 of the most influential print news publications. It's a research effort that shows the way towards automated systems that could potentially auto-classify the political character of a publication, and give readers a deeper insight into the ethical stance of an outlet on topics that they may feel passionately about. The work centers on the way topics are addressed with particular phrasing, such as undocumented immigrant illegal Immigrant, fetus unborn baby, demonstrators anarchists. The project used Natural Language Processing (NLP) techniques to extract and classify such instances of'charged' language (on the assumption that apparently more'neutral' terms also represent a political stance) into a broad mapping that reveals left and right-leaning bias across over three million articles from around 100 news outlets, resulting in a navigable bias landscape of the publications in question. The paper comes from Samantha D'Alonzo and Max Tegmark at MIT's Department of Physics, and observes that a number of recent initiatives around'fact checking', in the wake of numerous'fake news' scandals, can be interpreted as disingenuous and serving the causes of particular interests.
Disney is remaking the classic sci-fi movie 'Flight of the Navigator'
Disney is about to lean more on sci-fi nostalgia to reel in viewers. Deadline reports Disney is remaking its 1986 classic Flight of the Navigator for the streaming service. Details of the reboot are scarce, but it would feature a female lead and see Bryce Dallas Howard (who directed two The Mandalorian episodes) both direct and produce the title. It's safe to say the basic premise, of a child who bonds with an alien spaceship, won't change much for this adaptation. The project is a shrewd move for Disney.
Challenges in Detoxifying Language Models
Welbl, Johannes, Glaese, Amelia, Uesato, Jonathan, Dathathri, Sumanth, Mellor, John, Hendricks, Lisa Anne, Anderson, Kirsty, Kohli, Pushmeet, Coppin, Ben, Huang, Po-Sen
Large language models (LM) generate remarkably fluent text and can be efficiently adapted across NLP tasks. Measuring and guaranteeing the quality of generated text in terms of safety is imperative for deploying LMs in the real world; to this end, prior work often relies on automatic evaluation of LM toxicity. We critically discuss this approach, evaluate several toxicity mitigation strategies with respect to both automatic and human evaluation, and analyze consequences of toxicity mitigation in terms of model bias and LM quality. We demonstrate that while basic intervention strategies can effectively optimize previously established automatic metrics on the RealToxicityPrompts dataset, this comes at the cost of reduced LM coverage for both texts about, and dialects of, marginalized groups. Additionally, we find that human raters often disagree with high automatic toxicity scores after strong toxicity reduction interventions -- highlighting further the nuances involved in careful evaluation of LM toxicity.
Cross-Register Projection for Headline Part of Speech Tagging
Benton, Adrian, Li, Hanyang, Malioutov, Igor
Part of speech (POS) tagging is a familiar NLP task. State of the art taggers routinely achieve token-level accuracies of over 97% on news body text, evidence that the problem is well understood. However, the register of English news headlines, "headlinese", is very different from the register of long-form text, causing POS tagging models to underperform on headlines. In this work, we automatically annotate news headlines with POS tags by projecting predicted tags from corresponding sentences in news bodies. We train a multi-domain POS tagger on both long-form and headline text and show that joint training on both registers improves over training on just one or naively concatenating training sets. We evaluate on a newly-annotated corpus of over 5,248 English news headlines from the Google sentence compression corpus, and show that our model yields a 23% relative error reduction per token and 19% per headline. In addition, we demonstrate that better headline POS tags can improve the performance of a syntax-based open information extraction system. We make POSH, the POS-tagged Headline corpus, available to encourage research in improved NLP models for news headlines.
Co-Embedding: Discovering Communities on Bipartite Graphs through Projection
Candel, Gaëlle, Naccache, David
Many datasets take the form of a bipartite graph where two types of nodes are connected by relationships, like the movies watched by a user or the tags associated with a file. The partitioning of the bipartite graph could be used to fasten recommender systems, or reduce the information retrieval system's index size, by identifying groups of items with similar properties. This type of graph is often processed by algorithms using the Vector Space Model representation, where a binary vector represents an item with 0 and 1. The main problem with this representation is the dimension relatedness, like words' synonymity, which is not considered. This article proposes a co-clustering algorithm using items projection, allowing the measurement of features similarity. We evaluated our algorithm on a cluster retrieval task. Over various datasets, our algorithm produced well balanced clusters with coherent items in, leading to high retrieval scores on this task.
BacHMMachine: An Interpretable and Scalable Model for Algorithmic Harmonization for Four-part Baroque Chorales
Zhu, Yunyao, Hahn, Stephen, Mak, Simon, Jiang, Yue, Rudin, Cynthia
Algorithmic harmonization - the automated harmonization of a musical piece given its melodic line - is a challenging problem that has garnered much interest from both music theorists and computer scientists. One genre of particular interest is the four-part Baroque chorales of J.S. Bach. Methods for algorithmic chorale harmonization typically adopt a black-box, "data-driven" approach: they do not explicitly integrate principles from music theory but rely on a complex learning model trained with a large amount of chorale data. We propose instead a new harmonization model, called BacHMMachine, which employs a "theory-driven" framework guided by music composition principles, along with a "data-driven" model for learning compositional features within this framework. As its name suggests, BacHMMachine uses a novel Hidden Markov Model based on key and chord transitions, providing a probabilistic framework for learning key modulations and chordal progressions from a given melodic line. This allows for the generation of creative, yet musically coherent chorale harmonizations; integrating compositional principles allows for a much simpler model that results in vast decreases in computational burden and greater interpretability compared to state-of-the-art algorithmic harmonization methods, at no penalty to quality of harmonization or musicality. We demonstrate this improvement via comprehensive experiments and Turing tests comparing BacHMMachine to existing methods.
Move over James Bond! World's first hands-free JETPACK prototype is unveiled
From James Bond to The Jetsons, jetpacks have been a staple feature in blockbuster movies for years. Now, the technology is slowly but surely becoming a reality, with one company unveiling what it claims is the world's first hands-free jetpack prototype. Maverick Aviation has developed a device called the Maverick Jetpack, which it claims will travel at speeds of up to 30mph and could be ready by 2022. Unlike most existing jetpacks, which require intense training to get the hang of, the Maverick Jetpack has an in-built autopilot system and is intuitive to control, according to the team. The developers believe the device could be used to enter structures that are difficult to access in the near future, including wind turbines and construction sites.