Media
We Have Our First Great Summer Movie Disappointment of 2025
The taglines on M3GAN 2.0 posters read like text messages from an overconfident tween: "HEY, QUEENS." "MISS ME?" "I'M STILL THAT B." (Another that apparently exists, though I haven't seen in the wild, hilariously reads: "THIS BITCH.") Next to them, the titular robot who looks like an uncanny-valley Olsen twin peers from above circular sunglasses. This character that, per her 2023 film debut, will kill you and your little dog, too, is now being marketed with big child-star energy. While she always had more to offer than malice (her late-movie dance break went viral from its trailer alone), this moment marks a clear pivot on M3GAN's Mary Janes.
M3gan 2.0 review – hit-and-miss sequel replaces horror with action comedy
As the very first image of devil doll sequel M3gan 2.0 emerges on screen, of a desert with the words "somewhere on the Turkish-Iranian border" popping up like it's a Bond movie, you'd be forgiven for double-checking if you're in the right cinema. The original, a grabby artificial intelligence (AI) riff on Child's Play and Annabelle, was a brisk, by-the-numbers domestic horror, released on the first weekend of 2023, a slot usually given to the very worst genre films. M3gan was smarter than most, often sly and frequently funny and introducing what's now become a rarity, an almost instant non-IP pop culture icon, whose virality exploded the film into a surprise smash (raking in over 180m from a 12m budget). Like the films it was inspired by, a franchise was inevitable although where we're taken in M3gan 2.0 was far less of a given. For the follow-up, writer-director Gerard Johnstone has swerved from horror to action while retaining and tweaking the comedy with a release date that's been upgraded to summer blockbuster territory. It doesn't always work – a two-hour runtime that's a little too long, world-saving stakes that are a little too big, funny lines that are a little too not funny – but it's a mostly watchable second-tier event movie that, in a world of inconsequential sequels that fail to justify their existence, will do.
'The Dead Have Never Been This Talkative': The Rise of AI Resurrection
On June 18, AI image-generation company Midjourney released a tool that lets users create short video clips using their own images as a template. Days later, Reddit cofounder Alexis Ohanian posted on X about how he used the tech to animate a photo of his late mother, which shows him as a child wrapped in her embrace. In the artificial video, she laughs and smiles before rocking him in her arms. "Damn, I wasn't ready for how this would feel," he wrote. "This is how she hugged me.
Optimising Language Models for Downstream Tasks: A Post-Training Perspective
Language models (LMs) have demonstrated remarkable capabilities in NLP, yet adapting them efficiently and robustly to specific tasks remains challenging. As their scale and complexity grow, fine-tuning LMs on labelled data often underutilizes available unlabelled data, leads to overfitting on small task-specific sets, and imposes significant computational costs. These limitations hamper their application to the open-ended landscape of real-world language tasks. This thesis proposes a series of methods to better adapt LMs to downstream applications. First, we explore strategies for extracting task-relevant knowledge from unlabelled data, introducing a novel continued pre-training technique that outperforms state-of-the-art semi-supervised approaches. Next, we present a parameter-efficient fine-tuning method that substantially reduces memory and compute costs while maintaining competitive performance. We also introduce improved supervised fine-tuning methods that enable LMs to better follow instructions, especially when labelled data is scarce, enhancing their performance across a range of NLP tasks, including open-ended generation. Finally, we develop new evaluation methods and benchmarks, such as multi-hop spatial reasoning tasks, to assess LM capabilities and adaptation more comprehensively. Through extensive empirical studies across diverse NLP tasks, our results demonstrate that these approaches substantially improve LM robustness, efficiency, and generalization, making them more adaptable to a broad range of applications. These advances mark a significant step towards more robust and efficient LMs, bringing us closer to the goal of artificial general intelligence.
Structuralist Approach to AI Literary Criticism: Leveraging Greimas Semiotic Square for Large Language Models
Dong, Fangzhou, Zeng, Yifan, Sang, Yingpeng, Shen, Hong
Large Language Models (LLMs) excel in understanding and generating text but struggle with providing professional literary criticism for works with profound thoughts and complex narratives. This paper proposes GLASS (Greimas Literary Analysis via Semiotic Square), a structured analytical framework based on Greimas Semiotic Square (GSS), to enhance LLMs' ability to conduct in-depth literary analysis. GLASS facilitates the rapid dissection of narrative structures and deep meanings in narrative works. We propose the first dataset for GSS-based literary criticism, featuring detailed analyses of 48 works. Then we propose quantitative metrics for GSS-based literary criticism using the LLM-as-a-judge paradigm. Our framework's results, compared with expert criticism across multiple works and LLMs, show high performance. Finally, we applied GLASS to 39 classic works, producing original and high-quality analyses that address existing research gaps. This research provides an AI-based tool for literary research and education, offering insights into the cognitive mechanisms underlying literary engagement.
Improving Diffusion-Based Image Editing Faithfulness via Guidance and Scheduling
Text-guided diffusion models have become essential for high-quality image synthesis, enabling dynamic image editing. In image editing, two crucial aspects are editability, which determines the extent of modification, and faithfulness, which reflects how well unaltered elements are preserved. However, achieving optimal results is challenging because of the inherent trade-off between editability and faithfulness. To address this, we propose Faithfulness Guidance and Scheduling (FGS), which enhances faithfulness with minimal impact on editability. FGS incorporates faithfulness guidance to strengthen the preservation of input image information and introduces a scheduling strategy to resolve misalignment between editability and faithfulness. Experimental results demonstrate that FGS achieves superior faithfulness while maintaining editability. Moreover, its compatibility with various editing methods enables precise, high-quality image edits across diverse tasks.
Cat and Mouse -- Can Fake Text Generation Outpace Detector Systems?
McGlinchey, Andrea, Barclay, Peter J
Large language models (LLMs) can produce convincing'fake text' in domains such as academic writing, product reviews, and political news. Many approaches have been investigated for the detection of artificially generated text. While this may seem to presage an endless'arms race', we note that newer LLMs use ever more parameters, training data, and energy, while relatively simple classifiers demonstrate a good level of detection accuracy with modest resources. To approach the question of whether the models ability to beat the detectors may therefore reach a plateau, we examine the ability of statistical classifiers to identify'fake text' in the style of classical detective fiction. Over a 0.5 version increase, we found that Gemini showed an increased ability to generate deceptive text, while GPT did not. This suggests that reliable detection of fake text may remain feasible even for ever-larger models, though new model architectures may improve their deceptiveness.
Aligning Spoken Dialogue Models from User Interactions
Wu, Anne, Mazaré, Laurent, Zeghidour, Neil, Défossez, Alexandre
We propose a novel preference alignment framework for improving spoken dialogue models on real-time conversations from user interactions. Current preference learning methods primarily focus on text-based language models, and are not directly suited to the complexities of real-time speech interactions, with richer dynamics (e.g. interruption, interjection) and no explicit segmentation between speaker turns.We create a large-scale dataset of more than 150,000 preference pairs from raw multi-turn speech conversations, annotated with AI feedback, to cover preferences over both linguistic content and temporal context variations. We leverage offline alignment methods to finetune a full-duplex autoregressive speech-to-speech model. Extensive experiments demonstrate that feedback on generic conversations can be consistently effective in improving spoken dialogue models to produce more factual, safer and more contextually aligned interactions. We deploy the finetuned model and conduct holistic human evaluations to assess the impact beyond single-turn conversations. Our findings shed light on the importance of a well-calibrated balance among various dynamics, crucial for natural real-time speech dialogue systems.
skLEP: A Slovak General Language Understanding Benchmark
Šuppa, Marek, Ridzik, Andrej, Hládek, Daniel, Javůrek, Tomáš, Ondrejová, Viktória, Sásiková, Kristína, Tamajka, Martin, Šimko, Marián
In this work, we introduce skLEP, the first comprehensive benchmark specifically designed for evaluating Slovak natural language understanding (NLU) models. We have compiled skLEP to encompass nine diverse tasks that span token-level, sentence-pair, and document-level challenges, thereby offering a thorough assessment of model capabilities. To create this benchmark, we curated new, original datasets tailored for Slovak and meticulously translated established English NLU resources. Within this paper, we also present the first systematic and extensive evaluation of a wide array of Slovak-specific, multilingual, and English pre-trained language models using the skLEP tasks. Finally, we also release the complete benchmark data, an open-source toolkit facilitating both fine-tuning and evaluation of models, and a public leaderboard at https://github.com/slovak-nlp/sklep in the hopes of fostering reproducibility and drive future research in Slovak NLU.
A Hierarchical Deep Learning Approach for Minority Instrument Detection
Sechet, Dylan, Bugiotti, Francesca, Kowalski, Matthieu, d'Hérouville, Edouard, Langiewicz, Filip
Identifying instrument activities within audio excerpts is vital in music information retrieval, with significant implications for music cataloging and discovery. Prior deep learning endeavors in musical instrument recognition have predominantly emphasized instrument classes with ample data availability. Recent studies have demonstrated the applicability of hierarchical classification in detecting instrument activities in orchestral music, even with limited fine-grained annotations at the instrument level. Based on the Hornbostel-Sachs classification, such a hierarchical classification system is evaluated using the MedleyDB dataset, renowned for its diversity and richness concerning various instruments and music genres. This work presents various strategies to integrate hierarchical structures into models and tests a new class of models for hierarchical music prediction. This study showcases more reliable coarse-level instrument detection by bridging the gap between detailed instrument identification and group-level recognition, paving the way for further advancements in this domain.