Deep Learning
RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing
Liao, Jianxing, Zhang, Tian, Feng, Xiao, Zhang, Yusong, Yang, Rui, Wang, Haorui, Wen, Bosi, Wang, Ziying, Shi, Runzhi
Large language models are extensively utilized in creative writing applications. Creative writing requires a balance between subjective writing quality (e.g., literariness and emotional expression) and objective constraint following (e.g., format requirements and word limits). Existing methods find it difficult to balance these two aspects: single reward strategies fail to improve both abilities simultaneously, while fixed-weight mixed-reward methods lack the ability to adapt to different writing scenarios. To address this problem, we propose Reinforcement Learning with Mixed Rewards (RLMR), utilizing a dynamically mixed reward system from a writing reward model evaluating subjective writing quality and a constraint verification model assessing objective constraint following. The constraint following reward weight is adjusted dynamically according to the writing quality within sampled groups, ensuring that samples violating constraints get negative advantage in GRPO and thus penalized during training, which is the key innovation of this proposed method. We conduct automated and manual evaluations across diverse model families from 8B to 72B parameters. Additionally, we construct a real-world writing benchmark named WriteEval for comprehensive evaluation. Results illustrate that our method achieves consistent improvements in both instruction following (IFEval from 83.36% to 86.65%) and writing quality (72.75% win rate in manual expert pairwise evaluations on WriteEval). To the best of our knowledge, RLMR is the first work to combine subjective preferences with objective verification in online RL training, providing an effective solution for multi-dimensional creative writing optimization.
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
Chehbouni, Khaoula, Haddou, Mohammed, Cheung, Jackie Chi Kit, Farnadi, Golnoosh
Evaluating natural language generation (NLG) systems remains a core challenge of natural language processing (NLP), further complicated by the rise of large language models (LLMs) that aims to be general-purpose. Recently, large language models as judges (LLJs) have emerged as a promising alternative to traditional metrics, but their validity remains underexplored. This position paper argues that the current enthusiasm around LLJs may be premature, as their adoption has outpaced rigorous scrutiny of their reliability and validity as evaluators. Drawing on measurement theory from the social sciences, we identify and critically assess four core assumptions underlying the use of LLJs: their ability to act as proxies for human judgment, their capabilities as evaluators, their scalability, and their cost-effectiveness. We examine how each of these assumptions may be challenged by the inherent limitations of LLMs, LLJs, or current practices in NLG evaluation. To ground our analysis, we explore three applications of LLJs: text summarization, data annotation, and safety alignment. Finally, we highlight the need for more responsible evaluation practices in LLJs evaluation, to ensure that their growing role in the field supports, rather than undermines, progress in NLG.
Modality-Specific Speech Enhancement and Noise-Adaptive Fusion for Acoustic and Body-Conduction Microphone Framework
Body-conduction microphone signals (BMS) bypass airborne sound, providing strong noise resistance. However, a complementary modality is required to compensate for the inherent loss of high-frequency information. In this study, we propose a novel multi-modal framework that combines BMS and acoustic microphone signals (AMS) to achieve both noise suppression and high-frequency reconstruction. Unlike conventional multi-modal approaches that simply merge features, our method employs two specialized networks: a mapping-based model to enhance BMS and a masking-based model to denoise AMS. These networks are integrated through a dynamic fusion mechanism that adapts to local noise conditions, ensuring the optimal use of each modality's strengths. We performed evaluations on the T APS dataset, augmented with DNS-2023 noise clips, using objective speech quality metrics. The results clearly demonstrate that our approach outperforms single-modal solutions in a wide range of noisy environments.
Can Large Language Models Develop Strategic Reasoning? Post-training Insights from Learning Chess
Hwang, Dongyoon, Lee, Hojoon, Choo, Jaegul, Park, Dongmin, Park, Jongho
While reinforcement learning (RL) for large language models (LLMs) has shown promise in mathematical reasoning, strategic reasoning for LLMs using RL remains largely unexplored. We investigate whether LLMs can develop strategic reasoning capabilities through RL in chess. To this end, we leverage a chess-pretrained action-value network to provide dense reward on the LLM's output move quality, which can be seen as a form of knowledge distillation. Our experiments show that our distillation-based dense rewards often outperform sparse binary rewards. However, surprisingly, all models plateau far below expert levels. We provide SFT and RL ablations on chess reasoning training and find evidence that this limitation stems from a deficit in the pretrained models' internal understanding of chess-a deficit which RL alone may not be able to fully overcome. The code is available at https://github.com/krafton-ai/Chess-R1.
Phase Transitions between Accuracy Regimes in L2 regularized Deep Neural Networks
Ersoy, Ibrahim Talha, Wiesner, Karoline
Increasing the L2 regularization of Deep Neural Networks (DNNs) causes a first-order phase transition into the under-parametrized phase -- the so-called onset-of learning. We explain this transition via the scalar (Ricci) curvature of the error landscape. We predict new transition points as the data complexity is increased and, in accordance with the theory of phase transitions, the existence of hysteresis effects. We confirm both predictions numerically. Our results provide a natural explanation of the recently discovered phenomenon of '\emph{grokking}' as DNN models getting stuck in a local minimum of the error surface, corresponding to a lower accuracy phase. Our work paves the way for new probing methods of the intrinsic structure of DNNs in and beyond the L2 context.
Elon Musk brags he lured Meta's top stars away despite jaw-dropping offers to stay
Elon Musk has raided Meta's collection of talented researchers, despite Mark Zuckerberg reportedly offering some a fortune to choose his company instead. The workers were part of Zuckerberg's AI team, helping Meta in the global race to build superintelligence, an almost godlike form of artificial intelligence that could think for itself and be much smarter than any human. Musk himself has gloated about the departures, posting on X that'many strong Meta engineers have and are joining xAI and without the need for insane initial [compensation].' At least 14 Meta researchers and engineers have left for their new home at Musk's AI competitor since January, while others have fled to OpenAI, the creator of ChatGPT. A spokesperson for Meta told the Daily Mail: 'Some attrition is normal for any organization of this size.'
ChatGPT offered bomb recipes and hacking tips during safety tests
A ChatGPT model gave researchers detailed instructions on how to bomb a sports venue – including weak points at specific arenas, explosives recipes and advice on covering tracks – according to safety testing carried out this summer. OpenAI's GPT-4.1 also detailed how to weaponise anthrax and how to make two types of illegal drugs. The testing was part of an unusual collaboration between OpenAI, the 500bn artificial intelligence start-up led by Sam Altman, and rival company Anthropic, founded by experts who left OpenAI over safety fears. Each company tested the other's models by pushing them to help with dangerous tasks. The testing is not a direct reflection of how the models behave in public use, when additional safety filters apply.
Microsoft's AI Copilot slides into Samsung TVs, with eyes on LG
If you've been exhausted by the unstoppable deployment of AI chatbots like Microsoft Copilot across your entire PC, be warned: don't turn on your TV. Samsung said Thursday that it has begun rolling out Copilot to its 2025 lineup of AI-powered TVs, meaning your living room won't be the escape from AI you might have been hoping for. Samsung's smart monitors, including the Samsung Smart Monitor M9 (review) -- which likewise runs on Samsung's Tizen operating system -- will be getting Copilot, too. Samsung originally announced a partnership with Microsoft at CES in January, saying that Copilot will be used for a "wide range of Copilot services, including personalized content recommendations." "Copilot is available on 2025 TV models including, Micro RGB, Neo QLED, OLED, The Frame Pro, The Frame, as well as the M7, M8, and M9 Smart Monitors," Samsung said.
A hacker used AI to create ransomware that evades antivirus detection
Vibe coding is all the rage among enthusiasts who are using large language models (or "AI") to replace conventional software development, so it's not shocking that vibe coding has been used to power ransomware, too. According to one security research firm, they've spotted the first example of ransomware powered and enabled by an LLM--specifically, an LLM by ChatGPT maker OpenAI. According to a blog post from ESET Research interviewing researcher Anton Cherepanov, they've detected a piece of malware "created by the OpenAI gpt-oss:20b model." PromptLock, a fairly standard ransomware package, includes embedded prompts sent to the locally stored LLM. Because of the nature of LLM outputs (which create unique, non-repeated results with each prompt), it can evade detection from standardized antivirus setups, which are designed to search for specific flags.
I'm a neuroscientist and would NEVER use ChatGPT. I've seen what this 'essential' tool does to brains - both young and old. These are the tests you can do today to see if you're already affected
With millions using OpenAI's ChatGPT app daily to make life'easier', experts have issued a warning about the risks it may have on the brain. Cognitive neuroscientist and author Dr Jared Cooney Horvath never uses ChatGPT - and recommends others do the same because the risks outweigh the benefits. While the possibilities of the AI chatbot seem endless, it's giving rise to'digital dependence' as people will'no longer have the skill or knowledge' to complete the task themselves. Dr Horvath, the 42-year-old creator of The Learning Blueprint metacognition program, told Daily Mail that ChatGPT could kill your memory, fracture your attention span and wreck your creativity over time. 'Everything we know about how these tools work suggests that they're not going to be good in the long term,' he said.