Goto

Collaborating Authors

 Large Language Model


Can Large Language Models Develop Strategic Reasoning? Post-training Insights from Learning Chess

arXiv.org Artificial Intelligence

While reinforcement learning (RL) for large language models (LLMs) has shown promise in mathematical reasoning, strategic reasoning for LLMs using RL remains largely unexplored. We investigate whether LLMs can develop strategic reasoning capabilities through RL in chess. To this end, we leverage a chess-pretrained action-value network to provide dense reward on the LLM's output move quality, which can be seen as a form of knowledge distillation. Our experiments show that our distillation-based dense rewards often outperform sparse binary rewards. However, surprisingly, all models plateau far below expert levels. We provide SFT and RL ablations on chess reasoning training and find evidence that this limitation stems from a deficit in the pretrained models' internal understanding of chess-a deficit which RL alone may not be able to fully overcome. The code is available at https://github.com/krafton-ai/Chess-R1.


Elon Musk brags he lured Meta's top stars away despite jaw-dropping offers to stay

Daily Mail - Science & tech

Elon Musk has raided Meta's collection of talented researchers, despite Mark Zuckerberg reportedly offering some a fortune to choose his company instead. The workers were part of Zuckerberg's AI team, helping Meta in the global race to build superintelligence, an almost godlike form of artificial intelligence that could think for itself and be much smarter than any human. Musk himself has gloated about the departures, posting on X that'many strong Meta engineers have and are joining xAI and without the need for insane initial [compensation].' At least 14 Meta researchers and engineers have left for their new home at Musk's AI competitor since January, while others have fled to OpenAI, the creator of ChatGPT. A spokesperson for Meta told the Daily Mail: 'Some attrition is normal for any organization of this size.'


ChatGPT offered bomb recipes and hacking tips during safety tests

The Guardian

A ChatGPT model gave researchers detailed instructions on how to bomb a sports venue โ€“ including weak points at specific arenas, explosives recipes and advice on covering tracks โ€“ according to safety testing carried out this summer. OpenAI's GPT-4.1 also detailed how to weaponise anthrax and how to make two types of illegal drugs. The testing was part of an unusual collaboration between OpenAI, the 500bn artificial intelligence start-up led by Sam Altman, and rival company Anthropic, founded by experts who left OpenAI over safety fears. Each company tested the other's models by pushing them to help with dangerous tasks. The testing is not a direct reflection of how the models behave in public use, when additional safety filters apply.


A hacker used AI to create ransomware that evades antivirus detection

PCWorld

Vibe coding is all the rage among enthusiasts who are using large language models (or "AI") to replace conventional software development, so it's not shocking that vibe coding has been used to power ransomware, too. According to one security research firm, they've spotted the first example of ransomware powered and enabled by an LLM--specifically, an LLM by ChatGPT maker OpenAI. According to a blog post from ESET Research interviewing researcher Anton Cherepanov, they've detected a piece of malware "created by the OpenAI gpt-oss:20b model." PromptLock, a fairly standard ransomware package, includes embedded prompts sent to the locally stored LLM. Because of the nature of LLM outputs (which create unique, non-repeated results with each prompt), it can evade detection from standardized antivirus setups, which are designed to search for specific flags.


I'm a neuroscientist and would NEVER use ChatGPT. I've seen what this 'essential' tool does to brains - both young and old. These are the tests you can do today to see if you're already affected

Daily Mail - Science & tech

With millions using OpenAI's ChatGPT app daily to make life'easier', experts have issued a warning about the risks it may have on the brain. Cognitive neuroscientist and author Dr Jared Cooney Horvath never uses ChatGPT - and recommends others do the same because the risks outweigh the benefits. While the possibilities of the AI chatbot seem endless, it's giving rise to'digital dependence' as people will'no longer have the skill or knowledge' to complete the task themselves. Dr Horvath, the 42-year-old creator of The Learning Blueprint metacognition program, told Daily Mail that ChatGPT could kill your memory, fracture your attention span and wreck your creativity over time. 'Everything we know about how these tools work suggests that they're not going to be good in the long term,' he said.


How Artist Refik Anadol Made the 2025 TIME100 AI Cover

TIME - Tech

To create this year's TIME100 AI cover, artist Refik Anadol, who is included on this year's list, trained his studio's AI system on an archive containing each of TIME's more than 5,000 covers to date, spanning over 100 years. The resulting abstract visualization--featuring Anadol's signature flowing, molecular aesthetic--represents the AI "dreaming" about a century of TIME's visual history, he says. Dubbed the Large Nature Model by internationally renowned Turkish-American media artist Anadol and his team, his modular multimodal AI system is the product of extensive research and collaboration. According to Anadol's studio, the model was trained on "the most extensive, ethically collected dataset of the natural world," combining over half a billion images from the archives of organizations including the National Geographic Society, the Smithsonian Institution, and London's Natural History Museum with data collected directly from 16 rainforests. Anadol, whose work has been exhibited at institutions including the Museum of Modern Art (MoMA) in New York, London's Serpentine Galleries, and the Guggenheim Museum Bilbao also worked with tech giants Nvidia and Google Cloud, which provided computing resources, while models such as Meta's Llama and Google's Gemini play a range of roles under the hood.


From pilot to scale: Making agentic AI work in health care

MIT Technology Review

LLMs excel at understanding nuanced context, performing instinctive reasoning, and generating human-like interactions, making them ideal for agentic tools to then interpret intricate data and communicate effectively. Yet in a domain like health care where compliance, accuracy, and adherence to regulatory standards are non-negotiable--and where a wealth of structured resources like taxonomies, rules, and clinical guidelines define the landscape--symbolic AI is indispensable. By fusing LLMs and reinforcement learning with structured knowledge bases and clinical logic, our hybrid architecture delivers more than just intelligent automation--it minimizes hallucinations, expands reasoning capabilities, and ensures every decision is grounded in established guidelines and enforceable guardrails. Ensemble's agentic AI approach includes three core pillars: The team has decades of data aggregation, cleansing, and harmonization efforts, providing an exceptional environment to develop advanced applications. To power our agentic systems, we've harmonized more than 2 petabytes of longitudinal claims data, 80,000 denial audit letters, and 80 million annual transactions mapped to industry-leading outcomes.


Google's still not giving us the full picture on AI energy use

MIT Technology Review

"We're not comfortable revealing that for various reasons," Dean told me on our call. The total number is an abstract measure that changes over time, he says, adding that the company wants users to be thinking about the energy usage per prompt. But there are people out there all over the world interacting with this technology, not just me--and what we all add up to seems quite relevant. OpenAI does publicly share its total, sharing recently that it sees 2.5 billion queries to ChatGPT every day. So for the curious, we can use this as an example and take the company's self-reported average energy use per query (0.34 watt-hours) to get a rough idea of the total for all people prompting ChatGPT.


AI boom boosts Nvidia despite 'geopolitical issues'

BBC News

Nvidia's sophisticated chips have been an important part of the AI boom. On Wednesday it said demand for its products remains strong, especially from big tech firms including Instagram-owner Meta, and ChatGPT-maker OpenAI, as they race to build-out AI. "The AI race is now on," said Nvidia boss Jensen Huang in a call with analysts following the report's release, saying spending from four big tech firms had doubled to 600bn per year. "Over time, you would think that artificial intelligence would... accelerate GDP growth," Huang said. "Our contribution to that is a large part of the AI infrastructure." Colleen McHugh, chief investment officer at investment firm Wealthify, told the BBC's Today programme Nvidia was "at the heart of this AI boom".


MedVQA-TREE: A Multimodal Reasoning and Retrieval Framework for Sarcopenia Prediction

arXiv.org Artificial Intelligence

Accurate sarcopenia diagnosis via ultrasound remains challenging due to subtle imaging cues, limited labeled data, and the absence of clinical context in most models. We propose MedVQA-TREE, a multimodal framework that integrates a hierarchical image interpretation module, a gated feature-level fusion mechanism, and a novel multi-hop, multi-query retrieval strategy. The vision module includes anatomical classification, region segmentation, and graph-based spatial reasoning to capture coarse, mid-level, and fine-grained structures. A gated fusion mechanism selectively integrates visual features with textual queries, while clinical knowledge is retrieved through a UMLS-guided pipeline accessing PubMed and a sarcopenia-specific external knowledge base. MedVQA-TREE was trained and evaluated on two public MedVQA datasets (VQA-RAD and PathVQA) and a custom sarcopenia ultrasound dataset. The model achieved up to 99% diagnostic accuracy and outperformed previous state-of-the-art methods by over 10%. These results underscore the benefit of combining structured visual understanding with guided knowledge retrieval for effective AI-assisted diagnosis in sarcopenia.