Personal
What makes a quantum computer good?
What makes a quantum computer good? Claims that one quantum computer is better than another rest on terms like quantum advantage or quantum supremacy, fault-tolerance or qubits with better coherence - what does it all mean? Eleven years ago, I was just getting a start on my PhD in theoretical physics, and to be honest with you I never thought about quantum computers, or writing about them, at all. Meanwhile, staff were hard at work putting together the world's first " Quantum computer buyer's guide " (we've always been ahead of the curve). Looking through it reveals what a different time it was - John Martinis at University of California, Santa Barbara got a shout out for working on an array of only nine qubits, and just last week he was awarded the Nobel Prize in Physics .
How to Make STEM Funny--and Go Viral Doing It
If you stayed awake in science class as a kid, the payoff comes when you get a good laugh out of Freya McGhee's jokes. Stop me if you've heard this one before. An aspiring chemist goes to college, realizes she's not good at chemistry, and bombs her dissertation. She takes a class in standup comedy and decides the best way to talk about STEM is to make jokes at its expense. Based in London, the comedian had a strong interest in science as a kid, but after attending the University of Brighton to study chemistry, she realized that she liked learning science more than she liked applying it. Her thesis dissertation--"Synthesis of Iron Nitroxide radical species using radical derivatized ligands and its use as a single-molecule magnet"--flopped.
Taxonomy of User Needs and Actions
Shelby, Renee, Diaz, Fernando, Prabhakaran, Vinodkumar
The growing ubiquity of conversational AI highlights the need for frameworks that capture not only users' instrumental goals but also the situated, adaptive, and social practices through which they achieve them. Existing taxonomies of conversational behavior either overgeneralize, remain domain-specific, or reduce interactions to narrow dialogue functions. To address this gap, we introduce the Taxonomy of User Needs and Actions (TUNA), an empirically grounded framework developed through iterative qualitative analysis of 1193 human-AI conversations, supplemented by theoretical review and validation across diverse contexts. TUNA organizes user actions into a three-level hierarchy encompassing behaviors associated with information seeking, synthesis, procedural guidance, content creation, social interaction, and meta-conversation. By centering user agency and appropriation practices, TUNA enables multi-scale evaluation, supports policy harmonization across products, and provides a backbone for layering domain-specific taxonomies. This work contributes a systematic vocabulary for describing AI use, advancing both scholarly understanding and practical design of safer, more responsive, and more accountable conversational systems.
Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models
Nihal, Ragib Amin, Wen, Rui, Nakadai, Kazuhiro, Sakuma, Jun
Large language models (LLMs) remain vulnerable to multi-turn jailbreaking attacks that exploit conversational context to bypass safety constraints gradually. These attacks target different harm categories (like malware generation, harassment, or fraud) through distinct conversational approaches (educational discussions, personal experiences, hypothetical scenarios). Existing multi-turn jailbreaking methods often rely on heuristic or ad hoc exploration strategies, providing limited insight into underlying model weaknesses. The relationship between conversation patterns and model vulnerabilities across harm categories remains poorly understood. We propose Pattern Enhanced Chain of Attack (PE-CoA), a framework of five conversation patterns to construct effective multi-turn jailbreaks through natural dialogue. Evaluating PE-CoA on twelve LLMs spanning ten harm categories, we achieve state-of-the-art performance, uncovering pattern-specific vulnerabilities and LLM behavioral characteristics: models exhibit distinct weakness profiles where robustness to one conversational pattern does not generalize to others, and model families share similar failure modes. These findings highlight limitations of safety training and indicate the need for pattern-aware defenses. Code available on: https://github.com/Ragib-Amin-Nihal/PE-CoA
Benchmarking Chinese Commonsense Reasoning with a Multi-hop Reasoning Perspective
You, Wangjie, Wang, Xusheng, Wang, Xing, Jiao, Wenxiang, Feng, Chao, Li, Juntao, Zhang, Min
While Large Language Models (LLMs) have demonstrated advanced reasoning capabilities, their comprehensive evaluation in general Chinese-language contexts remains understudied. To bridge this gap, we propose Chinese Commonsense Multi-hop Reasoning (CCMOR), a novel benchmark designed to evaluate LLMs' ability to integrate Chinese-specific factual knowledge with multi-step logical reasoning. Specifically, we first construct a domain-balanced seed set from existing QA datasets, then develop an LLM-powered pipeline to generate multi-hop questions anchored on factual unit chains. To ensure the quality of resulting dataset, we implement a human-in-the-loop verification system, where domain experts systematically validate and refine the generated questions. Using CCMOR, we evaluate state-of-the-art LLMs, demonstrating persistent limitations in LLMs' ability to process long-tail knowledge and execute knowledge-intensive reasoning. Notably, retrieval-augmented generation substantially mitigates these knowledge gaps, yielding significant performance gains.
Text2Stories: Evaluating the Alignment Between Stakeholder Interviews and Generated User Stories
Dente, Francesco, Dalpiaz, Fabiano, Papotti, Paolo
Large language models (LLMs) can be employed for automating the generation of software requirements from natural language inputs such as the transcripts of elicitation interviews. However, evaluating whether those derived requirements faithfully reflect the stakeholders' needs remains a largely manual task. We introduce Text2Stories, a task and metrics for text-to-story alignment that allow quantifying the extent to which requirements (in the form of user stories) match the actual needs expressed by the elicitation session participants. Given an interview transcript and a set of user stories, our metric quantifies (i) correctness: the proportion of stories supported by the transcript, and (ii) completeness: the proportion of transcript supported by at least one story. We segment the transcript into text chunks and instantiate the alignment as a matching problem between chunks and stories. Experiments over four datasets show that an LLM-based matcher achieves 0.86 macro-F1 on held-out annotations, while embedding models alone remain behind but enable effective blocking. Finally, we show how our metrics enable the comparison across sets of stories (e.g., human vs. generated), positioning Text2Stories as a scalable, source-faithful complement to existing user-story quality criteria.
Move over, Alan Turing: meet the working-class hero of Bletchley Park you didn't see in the movies
Tommy Flowers: nothing like the machine he proposed had ever been contemplated. Tommy Flowers: nothing like the machine he proposed had ever been contemplated. Move over, Alan Turing: meet the working-class hero of Bletchley Park you didn't see in the movies The Oxbridge-educated boffin is feted as the codebreaking genius who helped Britain win the war. But should a little-known Post Office engineer named Tommy Flowers be seen as the real father of computing? T his is a story you know, right? It's early in the war and western Europe has fallen. Only the Channel stands between Britain and the fascist yoke; only Atlantic shipping lanes offer hope of the population continuing to be fed, clothed and armed. But hunting "wolf packs" of Nazi U-boats pick off merchant shipping at will, coordinated by radio instructions the Brits can intercept but can't read, thanks to the fiendish Enigma encryption machine.
Aftermath of RSF drone attack which killed dozens in Sudan's el-Fasher
Aftermath of RSF drone attack which killed dozens in Sudan's el-Fasher NewsFeed Aftermath of RSF drone attack which killed dozens in Sudan's el-Fasher Video shows the aftermath of drone and artillery strikes on a shelter in the besieged city of el-Fasher in Sudan's North Darfur state, which killed at least 60 people. The attack was carried out by the paramilitary Rapid Support Forces (RSF), according to a Sudanese medical advocacy group. Al Jazeera reporters follow Palestinians' return to northern Gaza Who is Nobel Peace Prize winner Maria Corina Machado?
Osaka Expo androids to be moved to Kyoto
Android robots shown at the Osaka Expo in a pavilion produced by University of Osaka professor Hiroshi Ishiguro will be relocated to Kyoto Prefecture. OSAKA - Seven android robots shown at the 2025 World Exposition in Osaka in a pavilion produced by University of Osaka professor Hiroshi Ishiguro will be relocated to Kyoto Prefecture after the end of the event on Monday. In addition, the Dutch pavilion will be moved to Awaji Island, Hyogo Prefecture. People involved in the use of expo assets after the event hope that they will be loved as tourist attractions in their new places. The prefectural government of Kyoto was chosen as the new owner of the androids in an open tender held by the expo organizer, the Japan Association for the 2025 World Exposition, in September. The robots will be shown to the public at a research facility in the Keihanna Science City research district straddling the Kyoto municipalities of Seika and Kizugawa.
Japan group to launch AI service for saury size predictions
Saury catches from August to the end of September this year totaled about 28,500 tons -- a 2.4-fold increase from the same period last year. The Japan Fisheries Information Service Center will start a service next fishing season that shows expected fishing grounds for saury by size class based on analysis using artificial intelligence technology. The Tokyo-based group of fisheries organizations provides information on fishing and ocean conditions. Since 2020, the group provides its predictions of likely saury fishing spots using AI, based on seawater temperature changes and past fishing records. The accuracy of the predictions has improved year after year.