Media
AI can contain gender bias, leading to potential disadvantages for women, expert says
Justine Bateman told Fox News Digital the increased use of artificial intelligence makes her sad because she feels it takes away from genuine human connection. Some experts have raised concerns that artificial intelligence (AI) could have its own gender gap if more women aren't involved in its development and dataset analysis. "It's not just AI, but I would say engineering as a whole," Dr. Georgianna Shea, chief technologist at the Foundation for Defense of Democracies' Center on Cyber and Technology Innovation (CCTI), told Fox News Digital. "Whenever there's any type of engineering process for anything, you don't want to end up with bias-based engineers." Adding to the debate, Melinda French Gates, co-chair of the Bill & Melinda Gates Foundation, recently said in an interview that she was concerned there was a lack of women working in the field of artificial intelligence, which she said made her nervous about potential biases in platforms.
Of Spiky SVDs and Music Recommendation
Afchar, Darius, Hennequin, Romain, Guigue, Vincent
The truncated singular value decomposition is a widely used methodology in music recommendation for direct similar-item retrieval or embedding musical items for downstream tasks. This paper investigates a curious effect that we show naturally occurring on many recommendation datasets: spiking formations in the embedding space. We first propose a metric to quantify this spiking organization's strength, then mathematically prove its origin tied to underlying communities of items of varying internal popularity. With this new-found theoretical understanding, we finally open the topic with an industrial use case of estimating how music embeddings' top-k similar items will change over time under the addition of data.
SMILE: Evaluation and Domain Adaptation for Social Media Language Understanding
Bashlovkina, Vasilisa, Matthews, Riley, Kuang, Zhaobin, Baumgartner, Simon, Bendersky, Michael
We study the ability of transformer-based language models (LMs) to understand social media language. Social media (SM) language is distinct from standard written language, yet existing benchmarks fall short of capturing LM performance in this socially, economically, and politically important domain. We quantify the degree to which social media language differs from conventional language and conclude that the difference is significant both in terms of token distribution and rate of linguistic shift. Next, we introduce a new benchmark for Social MedIa Language Evaluation (SMILE) that covers four SM platforms and eleven tasks. Finally, we show that learning a tokenizer and pretraining on a mix of social media and conventional language yields an LM that outperforms the best similar-sized alternative by 4.2 points on the overall SMILE score.
Stay on topic with Classifier-Free Guidance
Sanchez, Guillaume, Fan, Honglu, Spangher, Alexander, Levi, Elad, Ammanamanchi, Pawan Sasanka, Biderman, Stella
Classifier-Free Guidance (CFG) [37] has recently emerged in text-to-image generation as a lightweight technique to encourage prompt-adherence in generations. In this work, we demonstrate that CFG can be used broadly as an inference-time technique in pure language modeling. We show that CFG (1) improves the performance of Pythia, GPT-2 and LLaMA-family models across an array of tasks: Q&A, reasoning, code generation, and machine translation, achieving SOTA on LAMBADA with LLaMA-7B over PaLM-540B; (2) brings improvements equivalent to a model with twice the parameter-count; (3) can stack alongside other inference-time methods like Chain-of-Thought and Self-Consistency, yielding further improvements in difficult tasks; (4) can be used to increase the faithfulness and coherence of assistants in challenging form-driven and content-driven prompts: in a human evaluation we show a 75% preference for GPT4All using CFG over baseline.
Representer Point Selection for Explaining Regularized High-dimensional Models
Tsai, Che-Ping, Zhang, Jiong, Chien, Eli, Yu, Hsiang-Fu, Hsieh, Cho-Jui, Ravikumar, Pradeep
We introduce a novel class of sample-based explanations we term high-dimensional representers, that can be used to explain the predictions of a regularized high-dimensional model in terms of importance weights for each of the training samples. Our workhorse is a novel representer theorem for general regularized high-dimensional models, which decomposes the model prediction in terms of contributions from each of the training samples: with positive (negative) values corresponding to positive (negative) impact training samples to the model's prediction. We derive consequences for the canonical instances of $\ell_1$ regularized sparse models, and nuclear norm regularized low-rank models. As a case study, we further investigate the application of low-rank models in the context of collaborative filtering, where we instantiate high-dimensional representers for specific popular classes of models. Finally, we study the empirical performance of our proposed methods on three real-world binary classification datasets and two recommender system datasets. We also showcase the utility of high-dimensional representers in explaining model recommendations.
Conversational Question Answering on Heterogeneous Sources
Christmann, Philipp, Roy, Rishiraj Saha, Weikum, Gerhard
Conversational question answering (ConvQA) tackles sequential information needs where contexts in follow-up questions are left implicit. Current ConvQA systems operate over homogeneous sources of information: either a knowledge base (KB), or a text corpus, or a collection of tables. This paper addresses the novel issue of jointly tapping into all of these together, this way boosting answer coverage and confidence. We present CONVINSE, an end-to-end pipeline for ConvQA over heterogeneous sources, operating in three stages: i) learning an explicit structured representation of an incoming question and its conversational context, ii) harnessing this frame-like representation to uniformly capture relevant evidences from KB, text, and tables, and iii) running a fusion-in-decoder model to generate the answer. We construct and release the first benchmark, ConvMix, for ConvQA over heterogeneous sources, comprising 3000 real-user conversations with 16000 questions, along with entity annotations, completed question utterances, and question paraphrases. Experiments demonstrate the viability and advantages of our method, compared to state-of-the-art baselines.
'Blade Runner 2033: Labyrinth' is a new game set between the two movies
Annapurna Interactive is developing a game based on the iconic science fiction film Blade Runner. The game's set between the events of Blade Runner and Blade Runner 2049, so you can get some closure as to what Deckard was doing before meeting up with Ryan Gosling in an abandoned casino or whatever. Blade Runner 2033: Labyrinth follows a Blade Runner -- the name on their ID is blanked out in the trailer -- as they explore a mysterious location called the "land of the dead." You can't tell much from the trailer, but we see footage of what looks like an early version of the memory-crafting technology seen in Blade Runner 2049. Annapurna says this game is actually canon and it takes place just one year after the events of the original film, which would put it directly in the crosshairs of some big events alluded to in the sequel. It's always good to see more Blade Runner in gaming, especially after the criminally underrated and recently remastered 1997 adventure title.
'Storyteller' is the latest hot indie game coming to Netflix
Storyteller is a game about writing, word puzzles and the twisted tales we tell ourselves just to get through the day, and it'll be playable on Android and iOS via Netflix on September 26th. Storyteller is published by Annapurna Interactive and it landed on Switch and PC on March 23rd -- after spending more than a decade in development. Solo creator Daniel Benmergui announced Storyteller in 2011, and a prototype of the game actually won the Nuovo award for innovation at the Game Developers Conference in 2012. After that, life happened and Benmergui stopped working on Storyteller for a few years, but he eventually picked it back up and found a publishing partner in Annapurna. When Storyteller lands on iOS and Android in September, it'll come with free DLC that offers new stories for players to weave.
Election watchdog issues urgent warning over AI interference: 'race against the clock'
Tech Policy Center director Kara Fredrick explains how individuals and companies can mitigate the spread of misinformation by A.I. on'The Faulkner Focus.' British election regulators have urged politicians to pass new laws to limit spending on artificial intelligence (AI) as well as new requirements to identify AI-generated content. "The next U.K. general election is a ripe target for electronic disinformation given we are in the infancy of the AI age," Alan Mendoza, co-founder and executive director of the Henry Jackson Society, told Fox News Digital. "Many of the possible problems that may emerge have not even been considered." "As a result, we face a race against the clock to introduce appropriate protections, or run the nightmare risk of bad actors influencing campaigns and destroying public trust in our democratic process," he added.
The Download: AI disinformation, and lab-grown meat
The news: Disinformation generated by AI may be more convincing than disinformation written by humans, according to a new study. It found that people were 3% less likely to spot false tweets that had been generated by AI than real-life examples collected from Twitter. But the way in which GPT-3 orders information could have something to do with it, as AI-generated text tends to be more structured and condensed in comparison to how humans write. Why it matters: AI models can generate incorrect text that appears convincing, which could be used to generate false narratives quickly and cheaply for conspiracy theorists and disinformation campaigns. In theory, this could be spread further and faster than online disinformation networks manned by humans.