Large Language Model
Gear News This Week: The Repairable Fairphone 6 Arrives and Samsung's Galaxy Unpacked Is Up Next
The sixth generation of Fairphone arrived this week, featuring a modular design built to last from ethically sourced components in a climate-conscious way. It has been a couple of years since its predecessor, the Fairphone 5, and the Fairphone 6 is refreshingly smaller and lighter. It boasts a 6.3-inch OLED screen with a 120-Hz adaptive refresh rate, a Qualcomm Snapdragon 7s Gen 3 processor, and a 4,415 mAh battery that Fairphone says is good for up to two days. You also get a 50-megapixel main camera with a 13-MP ultrawide lens and a 32-MP selfie camera. Fairphone says the new device is made with more than 50 percent fair and recycled materials, including cobalt sourced through the Fair Cobalt Alliance, fair gold, silver, and tungsten, and recycled aluminum and rare earth metals.
OpenAI's Unreleased AGI Paper Could Complicate Microsoft Negotiations
A small clause inside OpenAI's contract with Microsoft, once considered a distant hypothetical, has now become a flashpoint in one of the biggest partnerships in tech. The clause states that if OpenAI's board ever declares it has developed artificial general intelligence (AGI), it would limit Microsoft's contracted access to the startup's future technologies. Microsoft, which has invested more than 13 billion in OpenAI, is now reportedly pushing for the removal of the clause and is considering walking away from the deal entirely, according to the Financial Times. Late last year, tensions around AGI's suddenly pivotal role in the Microsoft deal spilled into a debate within OpenAI over an internal research paper, according to multiple sources familiar with the matter. Titled "Five Levels of General AI Capabilities," the paper outlines a framework for classifying progressive stages of AI technology.
Is Using ChatGPT to Write Your Essay Bad for Your Brain? New MIT Study Explained.
TIME reporter Andrew Chow discussed the findings of a new study about how ChatGPT affects critical thinking with Nataliya Kosymyna. Kosymyna was part of a team of researchers at MIT's Media Lab who set out to determine whether ChatGPT and large language models (LLMs) are eroding critical thinking, and the study returned some concerning results. The study divided 54 subjects into three groups, and asked them to write several essays using OpenAI's ChatGPT, Google's search engine, and nothing at all, respectively. Researchers used an EEG to record the writers' brain activity. What they found was that of the three groups, the ChatGPT users had the lowest brain engagement and consistently underperformed at neural, linguistic and behavioral levels.
Google's emissions up 51% as AI electricity demand derails efforts to go green
Google's carbon emissions have soared by 51% since 2019 as artificial intelligence hampers the tech company's efforts to go green. While the corporation has invested in renewable energy and carbon removal technology, it has failed to curb its scope 3 emissions, which are those further down the supply chain, and are in large part influenced by a growth in datacentre capacity required to power artificial intelligence. The company reported a 27% increase in year-on-year electricity consumption as it struggles to decarbonise as quickly as its energy needs increase. Datacentres play a crucial role in training and operating the models that underpin AI models such as Google's Gemini and OpenAI's GPT-4, which powers the ChatGPT chatbot. The International Energy Agency estimates that datacentres' total electricity consumption could double from 2022 levels to 1,000TWh (terawatt hours) in 2026, approximately Japan's level of electricity demand.
Inside a plan to use AI to amplify doubts about the dangers of pollutants
An industry-backed researcher who has forged a career sowing doubt about the dangers of pollutants is attempting to use artificial intelligence (AI) to amplify his perspective. Louis Anthony "Tony" Cox Jr, a Denver-based risk analyst and former Trump adviser who once reportedly claimed there is no proof that cleaning air saves lives, is developing an AI application to scan academic research for what he sees as the false conflation of correlation with causation. Cox has described the project as an attempt to weed "propaganda" out of epidemiological research and perform "critical thinking at scale" in emails to industry researchers, which were obtained via Freedom of Information Act requests by the Energy and Policy Institute, a non-profit advocacy group, and exclusively reviewed by the Guardian. He has long leveled accusations of flimsiness at research linking exposure to chemical compounds with health dangers, including on behalf of polluting interests such as cigarette manufacturer Philip Morris and the American Petroleum Institute โ a fossil fuel lobbying group he has even allowed to "copy edit" his findings. Both the tobacco and oil industries have a history of weaponizing scientific uncertainty, experts say, with some arguing that similar tactics drive the Trump administration's current deregulatory efforts. The president's May "gold standard" science order, for instance, empowered his appointees to "correct scientific information" and "discipline" those who breach the administration's views, prompting outrage from some scientists. Cox has obtained funding to develop the new AI reviewer from the American Chemistry Council (ACC), the nation's largest chemical industry advocacy group, which counts oil and chemical giants such as Exxon and DuPont as members.
Google Gemini is coming for your private apps. Here's how to stop it
Google recently informed some users that Gemini AI will have access to numerous new apps starting July 7th, 2025. These include messaging apps and messengers such as WhatsApp, and it applies regardless of whether you actually use Gemini as an app assistant or not. In an email shared by Android Authority, Google states that they've "made it easier for Gemini to interact with your [Android] device" and that Gemini will "help you use" various apps "whether your Gemini Apps Activity is on or off." If you don't want this, you'll have to disable the feature in the Apps settings page, but Google hasn't yet provided an explanation of how this will work. Due to the vague wording in the email, the associated data privacy concerns, and the very sudden introduction of this change, many users are understandably concerned, especially since it doesn't seem to make any difference whether Gemini is activated or not.
Consistent Zero-shot 3D Texture Synthesis Using Geometry-aware Diffusion and Temporal Video Models
Kang, Donggoo, Kim, Jangyeong, Jeong, Dasol, Choi, Junyoung, Wi, Jeonga, Lee, Hyunmin, Gwon, Joonho, Paik, Joonki
Current texture synthesis methods, which generate textures from fixed viewpoints, suffer from inconsistencies due to the lack of global context and geometric understanding. Meanwhile, recent advancements in video generation models have demonstrated remarkable success in achieving temporally consistent videos. In this paper, we introduce Video-T ex, a novel framework for seamless texture synthesis that leverages video generation models to address both spatial and temporal inconsistencies in 3D textures. Our approach incorporates geometry-aware conditions, enabling precise utilization of 3D mesh structures. Additionally, we propose a structure-wise UV diffusion strategy, which enhances the generation of occluded areas by preserving semantic information, resulting in smoother and more coherent textures. VideoT ex not only achieves smoother transitions across UV boundaries but also ensures high-quality, temporally stable textures across video frames. Extensive experiments demonstrate that VideoT ex outperforms existing methods in texture fidelity, seam blending, and stability, paving the way for dynamic real-time applications that demand both visual quality and temporal coherence.
Poster: Enhancing GNN Robustness for Network Intrusion Detection via Agent-based Analysis
Zhan, Zhonghao, Zhou, Huichi, Haddadi, Hamed
--Graph Neural Networks (GNNs) show great promise for Network Intrusion Detection Systems (NIDS), particularly in IoT environments, but suffer performance degradation due to distribution drift and lack robustness against realistic adversarial attacks. Current robustness evaluations often rely on unrealistic synthetic perturbations and lack demonstrations on systematic analysis of different kinds of adversarial attack, which encompass both black-box and white-box scenarios. This work proposes a novel approach to enhance GNN robustness and generalization by employing Large Language Models (LLMs) in an agentic pipeline as simulated cybersecurity expert agents. These agents scrutinize graph structures derived from network flow data, identifying and potentially mitigating suspicious or adversarially perturbed elements before GNN processing. Our experiments, using a framework designed for realistic evaluation and testing with a variety of adversarial attacks including a dataset collected from physical testbed experiments, demonstrate that integrating LLM analysis can significantly improve the resilience of GNN-based NIDS against challenges, showcasing the potential of LLM agent as a complementary layer in intrusion detection architectures.
Optimising Language Models for Downstream Tasks: A Post-Training Perspective
Language models (LMs) have demonstrated remarkable capabilities in NLP, yet adapting them efficiently and robustly to specific tasks remains challenging. As their scale and complexity grow, fine-tuning LMs on labelled data often underutilizes available unlabelled data, leads to overfitting on small task-specific sets, and imposes significant computational costs. These limitations hamper their application to the open-ended landscape of real-world language tasks. This thesis proposes a series of methods to better adapt LMs to downstream applications. First, we explore strategies for extracting task-relevant knowledge from unlabelled data, introducing a novel continued pre-training technique that outperforms state-of-the-art semi-supervised approaches. Next, we present a parameter-efficient fine-tuning method that substantially reduces memory and compute costs while maintaining competitive performance. We also introduce improved supervised fine-tuning methods that enable LMs to better follow instructions, especially when labelled data is scarce, enhancing their performance across a range of NLP tasks, including open-ended generation. Finally, we develop new evaluation methods and benchmarks, such as multi-hop spatial reasoning tasks, to assess LM capabilities and adaptation more comprehensively. Through extensive empirical studies across diverse NLP tasks, our results demonstrate that these approaches substantially improve LM robustness, efficiency, and generalization, making them more adaptable to a broad range of applications. These advances mark a significant step towards more robust and efficient LMs, bringing us closer to the goal of artificial general intelligence.
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents
Men, Tianyi, Jin, Zhuoran, Cao, Pengfei, Chen, Yubo, Liu, Kang, Zhao, Jun
As Multimodal Large Language Models (MLLMs) advance, multimodal agents show promise in real-world tasks like web navigation and embodied intelligence. However, due to limitations in a lack of external feedback, these agents struggle with self-correction and generalization. A promising approach is to use reward models as external feedback, but there is no clear on how to select reward models for agents. Thus, there is an urgent need to build a reward bench targeted at agents. To address these challenges, we propose Agent-RewardBench, a benchmark designed to evaluate reward modeling ability in MLLMs. The benchmark is characterized by three key features: (1) Multiple dimensions and real-world agent scenarios evaluation. It covers perception, planning, and safety with 7 scenarios; (2) Step-level reward evaluation. It allows for the assessment of agent capabilities at the individual steps of a task, providing a more granular view of performance during the planning process; and (3) Appropriately difficulty and high-quality. We carefully sample from 10 diverse models, difficulty control to maintain task challenges, and manual verification to ensure the integrity of the data. Experiments demonstrate that even state-of-the-art multimodal models show limited performance, highlighting the need for specialized training in agent reward modeling. Code is available at github.