Retail
DRBench: A Realistic Benchmark for Enterprise Deep Research
Abaskohi, Amirhossein, Chen, Tianyi, Muñoz-Mármol, Miguel, Fox, Curtis, Ramesh, Amrutha Varshini, Marcotte, Étienne, Lù, Xing Han, Chapados, Nicolas, Gella, Spandana, Pal, Christopher, Drouin, Alexandre, Laradji, Issam H.
We introduce DRBench, a benchmark for evaluating AI agents on complex, open-ended deep research tasks in enterprise settings. Unlike prior benchmarks that focus on simple questions or web-only queries, DRBench evaluates agents on multi-step queries (for example, ``What changes should we make to our product roadmap to ensure compliance with this standard?") that require identifying supporting facts from both the public web and private company knowledge base. Each task is grounded in realistic user personas and enterprise context, spanning a heterogeneous search space that includes productivity software, cloud file systems, emails, chat conversations, and the open web. Tasks are generated through a carefully designed synthesis pipeline with human-in-the-loop verification, and agents are evaluated on their ability to recall relevant insights, maintain factual accuracy, and produce coherent, well-structured reports. We release 15 deep research tasks across 10 domains, such as Sales, Cybersecurity, and Compliance. We demonstrate the effectiveness of DRBench by evaluating diverse DR agents across open- and closed-source models (such as GPT, Llama, and Qwen) and DR strategies, highlighting their strengths, weaknesses, and the critical path for advancing enterprise deep research. Code is available at https://github.com/ServiceNow/drbench.
Full Moon October 2025: When To See The 'Harvest Supermoon' Rise
The harvest moon is the ninth of 12 full moons in 2025. A solar year is 365.24 days, while a lunar year is around 354.37 days, so sometimes there are 13 full moons in one calendar (solar) year -- as in 2023 and next in 2028. Of the 12 full moons in 2025, three will be "supermoons" -- of which the harvest moon is the first -- with two "blood moon" total lunar eclipses (the first happened on March 13-14 and the second on Sept. 7-8). The next full moon will be the beaver moon, the year's biggest "supermoon" (and the biggest since 2019), on Wednesday, Nov. 5, 2025.
Silent Leaks: Implicit Knowledge Extraction Attack on RAG Systems through Benign Queries
Wang, Yuhao, Qu, Wenjie, Zhai, Shengfang, Jiang, Yanze, Liu, Zichen, Liu, Yue, Dong, Yinpeng, Zhang, Jiaheng
Retrieval-Augmented Generation (RAG) systems enhance large language models (LLMs) by incorporating external knowledge bases, but this may expose them to extraction attacks, leading to potential copyright and privacy risks. However, existing extraction methods typically rely on malicious inputs such as prompt injection or jailbreaking, making them easily detectable via input- or output-level detection. In this paper, we introduce Implicit Knowledge Extraction Attack (IKEA), which conducts Knowledge Extraction on RAG systems through benign queries. Specifically, IKEA first leverages anchor concepts-keywords related to internal knowledge-to generate queries with a natural appearance, and then designs two mechanisms that lead anchor concepts to thoroughly "explore" the RAG's knowledge: (1) Experience Reflection Sampling, which samples anchor concepts based on past query-response histories, ensuring their relevance to the topic; (2) Trust Region Directed Mutation, which iteratively mutates anchor concepts under similarity constraints to further exploit the embedding space. Extensive experiments demonstrate IKEA's effectiveness under various defenses, surpassing baselines by over 80% in extraction efficiency and 90% in attack success rate. Moreover, the substitute RAG system built from IKEA's extractions shows comparable performance to the original RAG and outperforms those based on baselines across multiple evaluation tasks, underscoring the stealthy copyright infringement risk in RAG systems.
AI Playing Business Games: Benchmarking Large Language Models on Managerial Decision-Making in Dynamic Simulations
The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions. One of the most recent trends in AI benchmarking is performance of Large Language Models (LLMs) over longer time horizons. While LLMs excel at tasks involving natural language and pattern recognition, their capabilities in multi-step, strategic business decision-making remain largely unexplored. Few studies demonstrated how results can be different from benchmarks in short-term tasks, as Vending-Bench revealed. Meanwhile, there is a shortage of alternative benchmarks for long-term coherence. This research analyses a novel benchmark using a business game for the decision making in business. The research contributes to the recent literature on AI by proposing a reproducible, open-access management simulator to the research community for LLM benchmarking. This novel framework is used for evaluating the performance of five leading LLMs available in free online interface: Gemini, ChatGPT, Meta AI, Mistral AI, and Grok. LLM makes decisions for a simulated retail company. A dynamic, month-by-month management simulation provides transparently in spreadsheet model as experimental environment. In each of twelve months, the LLMs are provided with a structured prompt containing a full business report from the previous period and are tasked with making key strategic decisions: pricing, order size, marketing budget, hiring, dismissal, loans, training expense, R&D expense, sales forecast, income forecast The methodology is designed to compare the LLMs on quantitative metrics: profit, revenue, and market share, and other KPIs. LLM decisions are analyzed in their strategic coherence, adaptability to market changes, and the rationale provided for their decisions. This approach allows to move beyond simple performance metrics for assessment of the long-term decision-making.
Breville early Amazon Prime Day deals: Treat yourself to a high-end espresso machine or smart oven
Invest in a new espresso machine, smart oven, or other kitchen appliances from Breville and you won't want to eat out anymore. We may earn revenue from the products available on this page and participate in affiliate programs. The oven in my stove has been broken for a few weeks and it has barely affected my life because I use a Breville smart oven/air fryer almost exclusively when I cook at home . Right now, Amazon has a ton of Breville's high-end kitchen appliances on sale during its early Prime Big Deal Day sale. The official shopping holiday starts on October 7th, but these devices are on sale now, so you can order a new espresso machine and spend your Prime Day sipping while you scroll.
40 Best Early Amazon Prime Day Deals on WIRED-Tested Gear (2025)
Amazon Prime Day is back on October 7, but we've already found good deals on WIRED-approved gear. All products featured on WIRED are independently selected by our editors. However, we may receive compensation from retailers and/or from purchases of products through these links. It's that time of year again, and Prime Day deals are back. The Amazon Prime Big Deal Days event--also known as Amazon Prime Day 2--is officially arriving on October 7 and 8, but early deals have already started.
11 Best White Noise Machines (2025): Lectrofan, Snooz, Hatch, and More
The Best White-Noise Machines for a Blissful Night's Sleep Help the whole family catch more Z's with soothing background noise from our favorite sound machines. All products featured on WIRED are independently selected by our editors. However, we may receive compensation from retailers and/or from purchases of products through these links. The Best White noise machine isn't a complex device, even as companies constantly add more bells and whistles. Nowadays, they come in all shapes and sizes, outfitted with the capacity to play other noise frequencies and nature sounds while at home or in a more portable, on-the-go form. They're not just for kids or babies anymore--if you're like us, trying to drown out your internal monologue so that you can finally drift off, this is the article for you. But if you're building up your arsenal of sleep gadgets, with a white noise machine among them, we've tried out everything from the best sleep trackers, best sunrise alarm clocks, the best mattresses, and the best extreme alarm clocks .
Everything Amazon Announced Today at Its Fall Hardware Event (2025)
Amazon's next-gen Alexa+ chatbot is now available in four new Echo devices and a bevy of Ring cameras. The company also debuted three new Kindle Scribe tablets, one with a color screen. All products featured on WIRED are independently selected by our editors. However, we may receive compensation from retailers and/or from purchases of products through these links. It got a large language model power-up earlier this year in the form of Alexa+ (a paid upgrade for non-Amazon Prime subscribers), and now, Amazon has fresh hardware to take advantage of the assistant's new capabilities.
We Asked Audio Pros to Blind Test Headphones. The Results Were Surprising
We Asked Audio Pros to Blind Test Headphones. When it comes to choosing a good pair of headphones, what happens when you take features, design, and brand awareness out of the equation, and leave it all up to the sound? All products featured on WIRED are independently selected by our editors. However, we may receive compensation from retailers and/or from purchases of products through these links. What makes a really great pair of headphones? The basic answer used to be sound quality, but modern headphones offer so much more than just audio chops.
Grocery to General Merchandise: A Cross-Pollination Recommender using LLMs and Real-Time Cart Context
Kekuda, Akshay, Dandu, Murali Mohana Krishna, Lahiri, Rimita, Cai, Shiqin, Subramaniam, Sinduja, Korpeoglu, Evren, Achan, Kannan
Modern e-commerce platforms strive to enhance customer experience by providing timely and contextually relevant recommendations. However, recommending general merchandise to customers focused on grocery shopping -- such as pairing milk with a milk frother -- remains a critical yet under-explored challenge. This paper introduces a cross-pollination (XP) framework, a novel approach that bridges grocery and general merchandise cross-category recommendations by leveraging multi-source product associations and real-time cart context. Our solution employs a two-stage framework: (1) A candidate generation mechanism that uses co-purchase market basket analysis and LLM-based approach to identify novel item-item associations; and (2) a transformer-based ranker that leverages the real-time sequential cart context and optimizes for engagement signals such as add-to-carts. Offline analysis and online A/B tests show an increase of 36\% add-to-cart rate with LLM-based retrieval on the item page, and 15\% lift in add-to-cart using cart context-based ranker on the cart page. Our work contributes practical techniques for cross-category recommendations and broader insights for e-commerce systems.