arnold
Sainsbury's store pauses AI scanning after false shoplifting accusation
Matt Arnold was incorrectly identified of being a shoplifter at Sainsbury's East Dulwich superstore. Matt Arnold was incorrectly identified of being a shoplifter at Sainsbury's East Dulwich superstore. Sainsbury's store pauses AI scanning after false shoplifting accusation Supermarket chain says'human error', not its Facewatch technology, to blame for ejecting a customer Sainsbury's has paused the use of AI face scanning in one of its stores after a customer was wrongly identified as a shoplifter and ejected from the shop. "I was embarrassed, mortified even, and felt quite humiliated and powerless," Matt Arnold, 46, said of his ordeal. The comedy promoter was buying supplies in the store in East Dulwich, in south-east London, for a standup event at Dulwich Hamlet football club when, after scanning his items and a Nectar card, he was approached by two managers who told him he could not be served owing to an earlier incident.
Sainsbury's pauses AI cameras after shopper ousted
Sainsbury's pauses AI cameras after shopper ousted Sainsbury's has suspended its use of live facial recognition technology at one of its London branches after a customer was wrongly challenged as a shoplifter and asked to leave. Matt Arnold, 46, was buying items at the East Dulwich store when he was stopped by management at a self-service till. Describing the experience as a terrifying glimpse of the future, he has warned that retail surveillance tech risks humiliating customers and creating a system where staff feel compelled to follow automated alerts without applying human logic. The supermarket said it had apologised to Arnold and claimed the mistake was caused by human error, but critics have called for the technology to be scrapped. The incident happened on 6 August as Arnold - a comedy promoter getting supplies for a stand-up event hosted at the next door Dulwich Hamlet Football Club - was waiting for a staff member to approve an alcohol purchase.
The real storm chasers of the Great Plains
More information Adding us as a Preferred Source in Google by using this link indicates that you would like to see more of our content in Google News results. Storm chasers took this photo of a rotating wall cloud in Clovis, New Mexico, in May 2023. Breakthroughs, discoveries, and DIY tips sent six days a week. Flying cows, SUVs soaring through the air like toys, quaint towns that are virtually wiped off the map. Hollywood certainly makes the very real world of chasing tornadoes appear exciting on the big screen.
ChessArena: A Chess Testbed for Evaluating Strategic Reasoning Capabilities of Large Language Models
Liu, Jincheng, He, Sijun, Wu, Jingjing, Wang, Xiangsen, Chen, Yang, Kuang, Zhaoqi, Bao, Siqi, Yao, Yuan
Recent large language models (LLMs) have shown strong reasoning capabilities. However, a critical question remains: do these models possess genuine reasoning skills particularly complex strategic reasoning or are they primarily excelling at sophisticated pattern recognition within their training data? To address this question, this paper presents a chess testbed, ChessArena, to evaluate the strategic reasoning capabilities of LLMs. Chess requires complex strategic reasoning capabilities including long-term planning, strict rule comprehension, and multi-turn conversation memorization. Specifically, ChessArena is a competitive framework where LLMs play against each other, under four different play modes. The testbed is equipped with a ranking algorithm and a leaderboard. The testbed can also evaluate fine-grained capabilities including basic understanding, move selection, and puzzle solving. Over 13 LLMs with different modes are evaluated in ChessArena, playing over 800 games. The results reveal significant shortcomings in current LLMs: no model can beat Maia-1100 (a chess engine at human amateur level), while some even failed to defeat a random player that selects moves arbitrarily. We also present a strong baseline to the testbed: our fine-tuned Qwen3-8B substantially improved performance, approaching much larger state-of-the-art reasoning models.
EEFSUVA: A New Mathematical Olympiad Benchmark
Khatibi, Nicole N, Radamovich, Daniil A., Brenner, Michael P.
Recent breakthroughs have spurred claims that large language models (LLMs) match gold medal Olympiad to graduate level proficiency on mathematics benchmarks. In this work, we examine these claims in detail and assess the extent to which current benchmarks capture genuine LLM mathematical reasoning. The composition of these benchmarks, primarily drawing from the International Mathematics Olympiad (IMO) and related competitions, may overstate models reasoning ability due to potential data contamination and a narrow focus on familiar problem types. To enable a more holistic assessment of mathematical understanding, we introduce EEFSUVA, a novel benchmark curated from under circulated regional and national Olympiads of Eastern Europe and the countries from the former Soviet Union. These contests feature problems of comparable difficulty to the IMO and are renowned for demanding nonstandard problem-solving techniques, yet their problems are far less prevalent in online corpora. Preliminary results suggest that even state-of-the-art LLMs exhibit a notable performance decline on EEFSUVA relative to other Olympiad-style benchmarks. These findings also suggest the potential importance of broader evaluation datasets for a fuller assessment of mathematical reasoning and for guiding future model development.
Researchers Are Already Leaving Meta's New Superintelligence Lab
At least three artificial intelligence researchers have resigned from Meta's new superintelligence lab, just two months after CEO Mark Zuckerberg first announced the initiative. Two of the staffers have returned to OpenAI, where they both previously worked, after less than one-month stints at Meta, WIRED has confirmed. Ethan Knight worked at the ChatGPT maker earlier in his career but joined Meta from Elon Musk's xAI. A third researcher, Rishabh Agarwal, announced publicly on Monday he was leaving Meta's lab as well. He joined the tech giant in April to work on generative AI projects before switching to a role at Meta Superintelligence Labs (MSL), according to his LinkedIn profile.
Arnold: a generalist muscle transformer policy
Chiappa, Alberto Silvio, An, Boshi, Simos, Merkourios, Li, Chengkun, Mathis, Alexander
Controlling high-dimensional and nonlinear musculoskeletal models of the human body is a foundational scientific challenge. Recent machine learning breakthroughs have heralded policies that master individual skills like reaching, object manipulation and locomotion in musculoskeletal systems with many degrees of freedom. However, these agents are merely "specialists", achieving high performance for a single skill. In this work, we develop Arnold, a generalist policy that masters multiple tasks and embodiments. Arnold combines behavior cloning and fine-tuning with PPO to achieve expert or super-expert performance in 14 challenging control tasks from dexterous object manipulation to locomotion. A key innovation is Arnold's sensorimotor vocabulary, a compositional representation of the semantics of heterogeneous sensory modalities, objectives, and actuators. Arnold leverages this vocabulary via a transformer architecture to deal with the variable observation and action spaces of each task. This framework supports efficient multi-task, multi-embodiment learning and facilitates rapid adaptation to novel tasks. Finally, we analyze Arnold to provide insights into biological motor control, corroborating recent findings on the limited transferability of muscle synergies across tasks.