Large Language Model
Lenovo's new Chromebook has exclusive AI powers
The Chromebook team has been pushing Google Gemini AI powers, particularly on the more capable Chromebook Plus label, for a year and change. The new Lenovo Chromebook Plus 14, with a MediaTek processor boasting 50 TOPS, is a good example. Google used the upcoming Lenovo design to showcase Gemini's latest tricks at a press event last week. But I have to confess that I found the laptop itself, particularly its value proposition, more immediately gripping than the latest attempts to sell me a subscription that'll write my emails for me. The Chromebook Plus 14 combines a solid and lightweight build, a 14-inch OLED screen (touchscreen optional), generous memory (12GB or 16GB), and that MediaTek processor.
OpenAI takes down mentions of Jony Ive's io amid trademark row
OpenAI has taken down online content related to its recent deal with Sir Jony Ive's hardware startup, io, after a trademark complaint. The artificial intelligence company has removed promotional materials including a video where Ive โ the former Apple designer behind the iPhone โ and OpenAI's chief executive, Sam Altman, discuss the 6.4bn ( 4.8bn) transaction. However, the nine-minute film can still be viewed on YouTube. OpenAI, the developer of ChatGPT, was forced to act after receiving a legal complaint from iyO, a startup that makes artificial intelligence-backed earbuds. OpenAI said it had taken down a page on its website announcing the company's acquisition of io, which will involve Ive's company taking on creative and design leadership across the combined businesses.
'We were all pretty privileged': Allison Williams on Girls, nepo babies and toxic momfluencers
If you had wandered the set of the film M3gan 2.0 last year, chances are you would have stumbled into M3gan, the terrifying humanoid doll, staring lifelessly while she waited to be called for her next scene. Sometimes she would stand in the corner of the soundstage, says Allison Williams with a nervy laugh. "The dilemma is: do you turn her around so she's facing the wall, or do you let her face the room? In the sequel to the sci-fi horror M3gan, Williams resumes her role as Gemma, a roboticist who has become a crusader against rampant and reckless AI development after her creation โ developed for her orphaned niece โ became murderous. Acting opposite M3gan was unsettling, says Williams, speaking over a video call from a hotel room in New York. Sometimes she was played by the 15-year-old dancer Amie Donald, but often she was a robotic doll, animated by a small team. "When she's been working for a while, her eyelids can get sticky," says Williams. M3gan's handlers would paint lubricant on to her eyeballs with a brush and Williams would have to catch herself: "She's not flinching and for a second you're like: 'Ugh.' Then you remember: this is not a live thing." Still best known for her first role as Marnie in Lena Dunham's landmark TV series Girls, Williams has gravitated towards comedy-tinged horror in recent years. Her first post-Girls film role was in the Oscar-winning dark comedy horror Get Out. It and M3gan were relatively low-budget projects that became cultural phenomena โ Get Out for its commentary on racial politics, M3gan for what it says about the dangers of AI (as well as the uncanniness of M3gan herself). Williams has long been interested in AI โ she knows Sam Altman, the co-founder and CEO of OpenAI, which created ChatGPT, who put her in touch with robotics experts when she was researching the role of Gemma. The film raises questions not only about the danger of rogue AI, but about the ethical concerns โincluding how we should feel about the "rights" of devices. "It's easy to imbue anything that has AI in it with humanity.
FlowRAM: Grounding Flow Matching Policy with Region-Aware Mamba Framework for Robotic Manipulation
Wang, Sen, Wang, Le, Zhou, Sanping, Tian, Jingyi, Li, Jiayi, Sun, Haowen, Tang, Wei
Robotic manipulation in high-precision tasks is essential for numerous industrial and real-world applications where accuracy and speed are required. Yet current diffusion-based policy learning methods generally suffer from low computational efficiency due to the iterative denoising process during inference. Moreover, these methods do not fully explore the potential of generative models for enhancing information exploration in 3D environments. In response, we propose FlowRAM, a novel framework that leverages generative models to achieve region-aware perception, enabling efficient multimodal information processing. Specifically, we devise a Dynamic Radius Schedule, which allows adaptive perception, facilitating transitions from global scene comprehension to fine-grained geometric details. Furthermore, we integrate state space models to integrate multimodal information, while preserving linear computational complexity. In addition, we employ conditional flow matching to learn action poses by regressing deterministic vector fields, simplifying the learning process while maintaining performance. We verify the effectiveness of the FlowRAM in the RLBench, an established manipulation benchmark, and achieve state-of-the-art performance. The results demonstrate that FlowRAM achieves a remarkable improvement, particularly in high-precision tasks, where it outperforms previous methods by 12.0% in average success rate. Additionally, FlowRAM is able to generate physically plausible actions for a variety of real-world tasks in less than 4 time steps, significantly increasing inference speed.
History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation
Habibpour, Mobin, Afghah, Fatemeh
Object Goal Navigation (ObjectNav) challenges robots to find objects in unseen environments, demanding sophisticated reasoning. While Vision-Language Models (VLMs) show potential, current ObjectNav methods often employ them superficially, primarily using vision-language embeddings for object-scene similarity checks rather than leveraging deeper reasoning. This limits contextual understanding and leads to practical issues like repetitive navigation behaviors. This paper introduces a novel zero-shot ObjectNav framework that pioneers the use of dynamic, history-aware prompting to more deeply integrate VLM reasoning into frontier-based exploration. Our core innovation lies in providing the VLM with action history context, enabling it to generate semantic guidance scores for navigation actions while actively avoiding decision loops. We also introduce a VLM-assisted waypoint generation mechanism for refining the final approach to detected objects. Evaluated on the HM3D dataset within Habitat, our approach achieves a 46% Success Rate (SR) and 24.8% Success weighted by Path Length (SPL). These results are comparable to state-of-the-art zero-shot methods, demonstrating the significant potential of our history-augmented VLM prompting strategy for more robust and context-aware robotic navigation.
Cross-Modal Obfuscation for Jailbreak Attacks on Large Vision-Language Models
Jiang, Lei, Zhang, Zixun, Wang, Zizhou, Sun, Xiaobing, Li, Zhen, Zhen, Liangli, Xu, Xiaohua
Large Vision-Language Models (LVLMs) demonstrate exceptional performance across multimodal tasks, yet remain vulnerable to jailbreak attacks that bypass built-in safety mechanisms to elicit restricted content generation. Existing black-box jailbreak methods primarily rely on adversarial textual prompts or image perturbations, yet these approaches are highly detectable by standard content filtering systems and exhibit low query and computational efficiency. In this work, we present Cross-modal Adversarial Multimodal Obfuscation (CAMO), a novel black-box jailbreak attack framework that decomposes malicious prompts into semantically benign visual and textual fragments. By leveraging LVLMs' cross-modal reasoning abilities, CAMO covertly reconstructs harmful instructions through multi-step reasoning, evading conventional detection mechanisms. Our approach supports adjustable reasoning complexity and requires significantly fewer queries than prior attacks, enabling both stealth and efficiency. Comprehensive evaluations conducted on leading LVLMs validate CAMO's effectiveness, showcasing robust performance and strong cross-model transferability. These results underscore significant vulnerabilities in current built-in safety mechanisms, emphasizing an urgent need for advanced, alignment-aware security and safety solutions in vision-language systems.
Entropy-Driven Pre-Tokenization for Byte-Pair Encoding
Hu, Yifan, Liang, Frank, Zhao, Dachuan, Geuter, Jonathan, Reddy, Varshini, Schmidt, Craig W., Tanner, Chris
Byte-Pair Encoding (BPE) has become a widely adopted subword tokenization method in modern language models due to its simplicity and strong empirical performance across downstream tasks. However, applying BPE to unsegmented languages such as Chinese presents significant challenges, as its frequency-driven merge operation is agnostic to linguistic boundaries. To address this, we propose two entropy-informed pre-tokenization strategies that guide BPE segmentation using unsupervised information-theoretic cues. The first approach uses pointwise mutual information and left/right entropy to identify coherent character spans, while the second leverages predictive entropy derived from a pretrained GPT-2 model to detect boundary uncertainty. We evaluate both methods on a subset of the PKU dataset and demonstrate substantial improvements in segmentation precision, recall, and F1 score compared to standard BPE. Our results suggest that entropy-guided pre-tokenization not only enhances alignment with gold-standard linguistic units but also offers a promising direction for improving tokenization quality in low-resource and multilingual settings.
VectorEdits: A Dataset and Benchmark for Instruction-Based Editing of Vector Graphics
Kuchaล, Josef, Kadlฤรญk, Marek, Spiegel, Michal, ล tefรกnik, Michal
We introduce a large-scale dataset for instruction-guided vector image editing, consisting of over 270,000 pairs of SVG images paired with natural language edit instructions. Our dataset enables training and evaluation of models that modify vector graphics based on textual commands. We describe the data collection process, including image pairing via CLIP similarity and instruction generation with vision-language models. Initial experiments with state-of-the-art large language models reveal that current methods struggle to produce accurate and valid edits, underscoring the challenge of this task. To foster research in natural language-driven vector graphic generation and editing, we make our resources created within this work publicly available.
PR-Attack: Coordinated Prompt-RAG Attacks on Retrieval-Augmented Generation in Large Language Models via Bilevel Optimization
Jiao, Yang, Wang, Xiaodong, Yang, Kai
Large Language Models (LLMs) have demonstrated remarkable performance across a wide range of applications, e.g., medical question-answering, mathematical sciences, and code generation. However, they also exhibit inherent limitations, such as outdated knowledge and susceptibility to hallucinations. Retrieval-Augmented Generation (RAG) has emerged as a promising paradigm to address these issues, but it also introduces new vulnerabilities. Recent efforts have focused on the security of RAG-based LLMs, yet existing attack methods face three critical challenges: (1) their effectiveness declines sharply when only a limited number of poisoned texts can be injected into the knowledge database, (2) they lack sufficient stealth, as the attacks are often detectable by anomaly detection systems, which compromises their effectiveness, and (3) they rely on heuristic approaches to generate poisoned texts, lacking formal optimization frameworks and theoretic guarantees, which limits their effectiveness and applicability. To address these issues, we propose coordinated Prompt-RAG attack (PR-attack), a novel optimization-driven attack that introduces a small number of poisoned texts into the knowledge database while embedding a backdoor trigger within the prompt. When activated, the trigger causes the LLM to generate pre-designed responses to targeted queries, while maintaining normal behavior in other contexts. This ensures both high effectiveness and stealth. We formulate the attack generation process as a bilevel optimization problem leveraging a principled optimization framework to develop optimal poisoned texts and triggers. Extensive experiments across diverse LLMs and datasets demonstrate the effectiveness of PR-Attack, achieving a high attack success rate even with a limited number of poisoned texts and significantly improved stealth compared to existing methods.
Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening
He, Andre, Fried, Daniel, Welleck, Sean
Reinforcement learning is emerging as a primary driver for improving language model reasoning capabilities. A fundamental question is whether current reinforcement learning algorithms -- such as Group Relative Policy Optimization (GRPO), the de facto standard algorithm used to improve language model reasoning -- merely sharpen the base model's distribution around problems it can already solve. We investigate this question in the context of formal theorem proving, which has access to a perfect verifier. We identify a degenerate rank bias in GRPO in which highly probable trajectories are reinforced and rare ones are neglected. This results in distribution sharpening: the model can solve some problems with fewer samples, but underperforms simply sampling more solutions from the original model. To overcome GRPO's rank bias we introduce unlikeliness reward, a simple method for explicitly up-weighting rare but correct solutions. We show that unlikeliness reward mitigates rank bias and improves pass@$N$ across a large range of $N$ in both synthetic and real theorem proving settings. We also uncover an unexpected link between rank bias and a seemingly mundane hyperparameter -- the number of updates per batch -- that leads to a second, complementary mitigation. We combine our insights into a revised GRPO training recipe for formal theorem proving, yielding an open pipeline that achieves competitive performance to DeepSeek-Prover-V1.5-RL on the miniF2F-test benchmark. We release our implementation at https://github.com/AndreHe02/rewarding-unlikely-release