Goto

Collaborating Authors

 Media


The Download: escalating pandemic risks, and ask us anything on Reddit

MIT Technology Review

AI hype is built on high test scores. In the past few years, multiple researchers claim to have shown that large language models can pass cognitive tests designed for humans, from working through problems step by step, to guessing what other people are thinking. These kinds of results are feeding a hype machine predicting that these machines will soon come for white-collar jobs. But there's a problem: There's little agreement on what those results really mean. The Sun turned blue 200 years ago and no one knew why--until now. There's nothing better than a surreal TV crossover (Arrested Development Law & Order: SVU, anyone?)


Zuckerberg approved Meta's use of 'pirated' books to train AI models, authors claim

The Guardian

Citing internal Meta communications, the filing claims that the social network company's chief executive backed the use of the LibGen dataset, a vast online archive of books, despite warnings within the company's AI executive team that it is a dataset "we know to be pirated". The internal message says that using a database containing pirated material could weaken the Facebook and Instagram owner's negotiations with regulators, according to the filing. "Media coverage suggesting we have used a dataset we know to be pirated, such as LibGen, may undermine our negotiating position with regulators." The authors sued Meta in 2023, arguing that the social media company misused their books to train Llama, the large language model that powers its chatbots. The Library Genesis, or LibGen, dataset is a "shadow library" that originated in Russia and claims to contain millions of novels, nonfiction books and science magazine articles.


Update your iPhone NOW: Apple releases urgent iOS 18.2.1 update with important bug fixes - here's how to install it on your smartphone

Daily Mail - Science & tech

They are some of the world's most popular smartphones. And if you are an iPhone user, be sure to update your device today. Apple has released iOS 18.2.1 for the iPhone and recommends downloading it immediately. According to the tech giant, this update'provides important bug fixes and is recommended for all users'. This is the first big software update of 2025 and comes alongside the new iPadOS 18.2.1 for Apple's tablets.


Low rank matrix completion and realization of graphs: results and problems

arXiv.org Artificial Intelligence

The Netflix problem (from machine learning) asks the following. Given a ratings matrix in which each entry $(i,j)$ represents the rating of movie $j$ by customer $i$, if customer $i$ has watched movie $j$, and is otherwise missing, we would like to predict the remaining entries in order to make good recommendations to customers on what to watch next. The remaining entries are predicted so as to minimize the {\it rank} of the completed matrix. In this survey we study a more general problem, in which instead of knowing specific matrix elements, we know linear relations on such elements. We describe applications of these results to embeddings of graphs in surfaces (more precisely, embeddings with rotation systems, and embeddings modulo 2).


Hermit Kingdom Through the Lens of Multiple Perspectives: A Case Study of LLM Hallucination on North Korea

arXiv.org Artificial Intelligence

Hallucination in large language models (LLMs) remains a significant challenge for their safe deployment, particularly due to its potential to spread misinformation. Most existing solutions address this challenge by focusing on aligning the models with credible sources or by improving how models communicate their confidence (or lack thereof) in their outputs. While these measures may be effective in most contexts, they may fall short in scenarios requiring more nuanced approaches, especially in situations where access to accurate data is limited or determining credible sources is challenging. In this study, we take North Korea - a country characterised by an extreme lack of reliable sources and the prevalence of sensationalist falsehoods - as a case study. We explore and evaluate how some of the best-performing multilingual LLMs and specific language-based models generate information about North Korea in three languages spoken in countries with significant geo-political interests: English (United States, United Kingdom), Korean (South Korea), and Mandarin Chinese (China). Our findings reveal significant differences, suggesting that the choice of model and language can lead to vastly different understandings of North Korea, which has important implications given the global security challenges the country poses.


Text2Playlist: Generating Personalized Playlists from Text on Deezer

arXiv.org Artificial Intelligence

The streaming service Deezer heavily relies on the search to help users navigate through its extensive music catalog. Nonetheless, it is primarily designed to find specific items and does not lead directly to a smooth listening experience. We present Text2Playlist, a stand-alone tool that addresses these limitations. Text2Playlist leverages generative AI, music information retrieval and recommendation systems to generate query-specific and personalized playlists, successfully deployed at scale.


Controlling Large Language Models Through Concept Activation Vectors

arXiv.org Artificial Intelligence

As large language models (LLMs) are widely deployed across various domains, the ability to control their generated outputs has become more critical. This control involves aligning LLMs outputs with human values and ethical principles or customizing LLMs on specific topics or styles for individual users. Existing controlled generation methods either require significant computational resources and extensive trial-and-error or provide coarse-grained control. In this paper, we propose Generation with Concept Activation Vector (GCAV), a lightweight model control framework that ensures accurate control without requiring resource-extensive fine-tuning. Specifically, GCAV first trains a concept activation vector for specified concepts to be controlled, such as toxicity. During inference, GCAV steers the concept vector in LLMs, for example, by removing the toxicity concept vector from the activation layers. Control experiments from different perspectives, including toxicity reduction, sentiment control, linguistic style, and topic control, demonstrate that our framework achieves state-of-the-art performance with granular control, allowing for fine-grained adjustments of both the steering layers and the steering magnitudes for individual samples.


This crowdsourcing app is a lifeline for Californians tracking wildfires

Popular Science

Tens of thousands of Californians are turning to a crowdsourced, nonprofit app called Watch Duty for critical, up-to-the-moment disaster updates as deadly fires continue to rage through the state. The app, which uses a mixture of official government and volunteer data to track wildfires, surpassed OpenAI's ChatGPT and Meta's Threads as the most downloaded app on the Apple App Store on Wednesday. Social media users have encouraged residents in affected areas to download the app in order to track the fire's rapid movements and stay aware of possible evacuation orders. Apps like Watch Duty, which have seen a surge in interest in recent years, may become even more important as climate change-related natural disasters intensify in scope and scale. It gives you updates on fires nearby, evacuation notices, and even will show you where an evacuation center is if you need to evacuate!


Optimizing Estonian TV Subtitles with Semi-supervised Learning and LLMs

arXiv.org Artificial Intelligence

For instance, Both iterative pseudo-labeling and LLM-based recent studies (Mykhalevych and Preply, 2024; post-editing have been an active area of research Kim et al., 2023) have revealed that 50% of Americans in the context of verbatim automatic speech and 85% of the Netflix users overall frequently recognition (ASR). Pseudo-labeling based semisupervised watch TV and streaming video content learning in ASR has been studied since with subtitles. Studies show that subtitles can enhance at least (Zavaliagkos et al., 1998) and has been understanding and memory retention. A lot later investigated in several works, e.g. by Veselỳ of viewers choose to enjoy their content quietly et al. (2013); Xu et al. (2020).


De-centering the (Traditional) User: Multistakeholder Evaluation of Recommender Systems

arXiv.org Artificial Intelligence

Expanding the frame of evaluation to include other parties, as well as the ecosystem in which the system is deployed, leads us to a multistakeholder view of recommender system evaluation as defined in [2]: "A multistakeholder evaluation is one in which the quality of recommendations is assessed across multiple groups of stakeholders." In this article, we provide (i) an overview of the types of recommendation stakeholders that can be considered in conducting such evaluations, (ii) a discussion of the considerations and values that enter into developing measures that capture outcomes of interest for a diversity of stakeholders, (iii) an outline of a methodology for developing and applying multistakeholder evaluation, and (iv) three examples of different multistakeholder scenarios including derivations of evaluation metrics for different stakeholder groups in these different scenarios. The variety of possible stakeholders we identified that are part of the general recommendation ecosystem is suggested in Figure 1 and defined here, using the terminology from [1, 2]: Recommendation consumers are the traditional recommender system users to whom recommendations are delivered and to which typical forms of recommender system evaluation are oriented. Item providers form the general class of individuals or entities who create or otherwise stand behind the items being recommended.