Goto

Collaborating Authors

 lévy


What China's Vibe-Coding Capital Can Tell Us About the A.I. Boom

The New Yorker

What China's Vibe-Coding Capital Can Tell Us About the A.I. Boom Tech workers have flocked to Liangzhu, a small town outside Hangzhou, looking to escape China's corporate culture and build their future with A.I. In mid-May, days after Donald Trump visited Beijing, a Chinese tech founder named Levy invited me to an office party for his A.I. fortune-telling company. It's called FateTell, and it makes an app that, in the vein of traditional Chinese astrology, reads your, a chart that tracks the elemental forces associated with your birth date and time. According to his invitation, my governing element, fire, would counteract the water element that Levy carries in oversupply.) So I travelled from Beijing to FateTell's headquarters, situated on the edge of Hangzhou--one of China's tech hubs, home to Alibaba and the insurgent Chinese A.I. company DeepSeek --in a town called Liangzhu. Liangzhu means "beautiful waterland," and it more than earns the name--rivers, lakes, and marshes thread through the low hills at the town's edge. The town is hip, idyllic, contrarian, and increasingly known as an escape for workers burnt out by the grind of China's booming tech industry.


"You Had to Be There" review: Martin Short, Eugene Levy, and other comedy legends reveal their shared showbiz start

Mashable

Say More Look Up Versus Mashable's Best: E-readers, robovacs, laptops, earbuds, smart home and more Mashable Selects Creator Playbook In My Bag Trending Now Back to School Good Connection: Uplifting stories for a digital age Switch Off Mashable Voices All Series From SCTV' to Saturday Night Live, one production of Godspell changed comedy forever. Kristy Puchko is the Entertainment Editor at Mashable. Based in New York City, she's an established film critic and entertainment reporter who has traveled the world on assignment, covered a variety of film festivals, co-hosted movie-focused podcasts, and interviewed a wide array of performers and filmmakers. Victor Garber, Martin Short, Gilda Radner, Eugene Levy, and others perform in the 1972 production of Godspell in Toronto. Even the most devoted comedy nerd may not realize how many truly iconic comedies of past and present might never have existed if it weren't for a single theatrical production.


'Brutal': thousands of T-shirt designs stolen and listed on Temu, Sydney label claims

The Guardian

A comparison of a T-shirt as sold by Lonely Kids Club and Temu. The business owner says AI may have been used to scrape the content of his site and replicate the designs. A comparison of a T-shirt as sold by Lonely Kids Club and Temu. The business owner says AI may have been used to scrape the content of his site and replicate the designs. 'Brutal': thousands of T-shirt designs stolen and listed on Temu, Sydney label claims After 15 years running a small business making T-shirts, Warwick Levy says it was "brutal" when he first discovered an identical design for sale on Temu.




How Well Does First-Token Entropy Approximate Word Entropy as a Psycholinguistic Predictor?

arXiv.org Artificial Intelligence

Contextual entropy is a psycholinguistic measure capturing the anticipated difficulty of processing a word just before it is encountered. Recent studies have tested for entropy-related effects as a potential complement to well-known effects from surprisal. For convenience, entropy is typically estimated based on a language model's probability distribution over a word's first subword token. However, this approximation results in underestimation and potential distortion of true word entropy. To address this, we generate Monte Carlo (MC) estimates of word entropy that allow words to span a variable number of tokens. Regression experiments on reading times show divergent results between first-token and MC word entropy, suggesting a need for caution in using first-token approximations of contextual entropy.


Is this the raciest conference invite ever?

New Scientist

Feedback is New Scientist's popular sideways look at the latest science and technology news. You can submit items you believe may amuse readers to Feedback by emailing feedback@newscientist.com Recently, Feedback was delighted to peruse the raciest conference invitation we have ever received. We get a lot of conference invites from organisers labouring under the delusion we are doing something akin to science journalism, and they are mostly a little prosaic: what's new in G-protein signalling, more findings about the biology of molluscs, that kind of thing. Here is the opening line: "From its groundbreaking inception in London to its spectacular evolution in the vibrant heart of China, the Love and Sex with Robots Conference is gearing up for its most thrilling chapters yet: its landmark 12th International edition, scheduled for June 2026."


ReFeed: Multi-dimensional Summarization Refinement with Reflective Reasoning on Feedback

arXiv.org Artificial Intelligence

Summarization refinement faces challenges when extending to multi-dimension. In this paper, we introduce ReFeed, a powerful summarization refinement pipeline that enhances multiple dimensions through reflective reasoning on feedback. To achieve this, we release SumFeed-CoT, a large-scale Long-CoT-based dataset optimized for training a lightweight model with reflective reasoning. Our experiments reveal how the number of dimensions, feedback exposure, and reasoning policy influence refinement performance, highlighting reflective reasoning and simultaneously addressing multiple feedback is crucial to mitigate trade-off between dimensions. Furthermore, ReFeed is robust to noisy feedback and feedback order. Lastly, our finding emphasizes that creating data with a proper goal and guideline constitutes a fundamental pillar of effective reasoning. The dataset and model will be released.


Structure based SAT dataset for analysing GNN generalisation

arXiv.org Artificial Intelligence

Satisfiability (SAT) solvers based on techniques such as conflict driven clause learning (CDCL) have produced excellent performance on both synthetic and real world industrial problems. While these CDCL solvers only operate on a per-problem basis, graph neural network (GNN) based solvers bring new benefits to the field by allowing practitioners to exploit knowledge gained from solved problems to expedite solving of new SAT problems. However, one specific area that is often studied in the context of CDCL solvers, but largely overlooked in GNN solvers, is the relationship between graph theoretic measure of structure in SAT problems and the generalisation ability of GNN solvers. To bridge the gap between structural graph properties (e.g., modularity, self-similarity) and the generalisability (or lack thereof) of GNN based SAT solvers, we present StructureSAT: a curated dataset, along with code to further generate novel examples, containing a diverse set of SAT problems from well known problem domains. Furthermore, we utilise a novel splitting method that focuses on deconstructing the families into more detailed hierarchies based on their structural properties. With the new dataset, we aim to help explain problematic generalisation in existing GNN SAT solvers by exploiting knowledge of structural graph properties. We conclude with multiple future directions that can help researchers in GNN based SAT solving develop more effective and generalisable SAT solvers.


Large Language Models Are Human-Like Internally

arXiv.org Artificial Intelligence

Recent cognitive modeling studies have reported that larger language models (LMs) exhibit a poorer fit to human reading behavior, leading to claims of their cognitive implausibility. In this paper, we revisit this argument through the lens of mechanistic interpretability and argue that prior conclusions were skewed by an exclusive focus on the final layers of LMs. Our analysis reveals that next-word probabilities derived from internal layers of larger LMs align with human sentence processing data as well as, or better than, those from smaller LMs. This alignment holds consistently across behavioral (self-paced reading times, gaze durations, MAZE task processing times) and neurophysiological (N400 brain potentials) measures, challenging earlier mixed results and suggesting that the cognitive plausibility of larger LMs has been underestimated. Furthermore, we first identify an intriguing relationship between LM layers and human measures: earlier layers correspond more closely with fast gaze durations, while later layers better align with relatively slower signals such as N400 potentials and MAZE processing times. Our work opens new avenues for interdisciplinary research at the intersection of mechanistic interpretability and cognitive modeling.