Large Language Model
ScholarBench: A Bilingual Benchmark for Abstraction, Comprehension, and Reasoning Evaluation in Academic Contexts
Noh, Dongwon, Koh, Donghyeok, Yuk, Junghun, Kim, Gyuwan, Lee, Jaeyong, Lim, Kyungtae, Park, Cheoneum
Prior benchmarks for evaluating the domain-specific knowledge of large language models (LLMs) lack the scalability to handle complex academic tasks. To address this, we introduce \texttt{ScholarBench}, a benchmark centered on deep expert knowledge and complex academic problem-solving, which evaluates the academic reasoning ability of LLMs and is constructed through a three-step process. \texttt{ScholarBench} targets more specialized and logically complex contexts derived from academic literature, encompassing five distinct problem types. Unlike prior benchmarks, \texttt{ScholarBench} evaluates the abstraction, comprehension, and reasoning capabilities of LLMs across eight distinct research domains. To ensure high-quality evaluation data, we define category-specific example attributes and design questions that are aligned with the characteristic research methodologies and discourse structures of each domain. Additionally, this benchmark operates as an English-Korean bilingual dataset, facilitating simultaneous evaluation for linguistic capabilities of LLMs in both languages. The benchmark comprises 5,031 examples in Korean and 5,309 in English, with even state-of-the-art models like o3-mini achieving an average evaluation score of only 0.543, demonstrating the challenging nature of this benchmark.
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
Meterez, Alexandru, Morwani, Depen, Wu, Jingfeng, Oncescu, Costin-Andrei, Pehlevan, Cengiz, Kakade, Sham
Increasing the batch size during training -- a ''batch ramp'' -- is a promising strategy to accelerate large language model pretraining. While for SGD, doubling the batch size can be equivalent to halving the learning rate, the optimal strategy for adaptive optimizers like Adam is less clear. As a result, any batch-ramp scheduling, if used at all, is typically tuned heuristically. This work develops a principled framework for batch-size scheduling and introduces Seesaw: whenever a standard scheduler would halve the learning rate, Seesaw instead multiplies it by $1/\sqrt{2}$ and doubles the batch size, preserving loss dynamics while reducing serial steps. Theoretically, we provide, to our knowledge, the first finite-sample proof of equivalence between learning-rate decay and batch-size ramp-up for SGD on noisy linear regression, and we extend this equivalence to normalized SGD, a tractable proxy for Adam, under a variance-dominated regime observed in practice. Empirically, on 150M/300M/600M-parameter models trained at Chinchilla scale using a constant (critical) batch size, Seesaw matches cosine decay at equal FLOPs while reducing wall-clock time by $\approx 36\%$, approaching the theoretical limit implied by our analysis.
LLM Prompt Duel Optimizer: Efficient Label-Free Prompt Optimization
Wu, Yuanchen, Verma, Saurabh, Lee, Justin, Xiong, Fangzhou, Zhang, Poppy, Awadelkarim, Amel, Chen, Xu, Yuan, Yubai, Hill, Shawndra
Large language models (LLMs) are highly sensitive to their input prompts, making prompt design a central challenge. While automatic prompt optimization (APO) reduces manual engineering, most approaches assume access to ground-truth references such as labeled validation data. In practice, however, collecting high-quality labels is costly and slow. We propose the Prompt Duel Optimizer (PDO), a sample-efficient framework for label-free prompt optimization. PDO formulates the problem as a dueling-bandit setting, where supervision signal comes from pairwise preference feedback provided by an LLM judge. The framework combines Double Thompson Sampling (D-TS), which prioritizes informative prompt comparisons, with Top-Performer Guided Mutation, which expands the candidate pool by mutating high-performing prompts. PDO naturally operates in label-free settings and can also incorporate partial labels to mitigate judge noise. Experiments on BIG-bench Hard (BBH) and MS MARCO show that PDO consistently outperforms baseline methods. Ablation studies further demonstrate the effectiveness of both D-TS and prompt mutation.
Chatbots Are Becoming More Sexually Explicit in a Bid to Attract Usership and Paying Customers
The eighteen plus symbol (18+) appears on a smartphone screen, and the OpenAI logo displays as the background on a laptop screen in this photo illustration in Athens, Greece, on October 16, 2025. The eighteen plus symbol (18+) appears on a smartphone screen, and the OpenAI logo displays as the background on a laptop screen in this photo illustration in Athens, Greece, on October 16, 2025. In August, OpenAI CEO Sam Altman said on a podcast that he was "proud" that his company had not gotten "distracted" by putting features like a "sexbot avatar" into ChatGPT. But on Tuesday, he announced that adult users will be able to access explicit interactive experiences, marking a major shift in the company's practices. "In December, as we roll out age-gating more fully and as part of our'treat adult users like adults' principle, we will allow even more, like erotica for verified adults," Altman said in a post on X.
Barrister found to have used AI to prepare for hearing after citing 'fictitious' cases
The judge said: 'I am bound to observe that one of the cases cited has recently been wrongly deployed by ChatGPT in support of similar arguments.' The judge said: 'I am bound to observe that one of the cases cited has recently been wrongly deployed by ChatGPT in support of similar arguments.' Barrister found to have used AI to prepare for hearing after citing'fictitious' cases Judge rules Chowdhury Rahman used ChatGPT-like software and then tried to hide it, wasting immigration tribunal's time Thu 16 Oct 2025 09.47 EDTFirst published on Thu 16 Oct 2025 09.33 EDT An immigration barrister was found by a judge to be using AI to do his work for a tribunal hearing after citing cases that were "entirely fictitious" or "wholly irrelevant". Chowdhury Rahman was discovered using ChatGPT-like software to prepare his legal research, a tribunal heard. Rahman was found not only to have used AI to prepare his work, but "failed thereafter to undertake any proper checks on the accuracy".
Microsoft supercharges Copilot with Google integration, smarter vision
When you purchase through links in our articles, we may earn a small commission. Microsoft's Copilot AI technologies will be able to see more and connect to a greater range of files. Copilot Vision's eyesight is improving, as the integrated Windows AI technology will soon be able to see entire documents, plus link to apps like Google Drive via a new connectors function. Separately, Microsoft is adding Copilot to the Windows 11 taskbar and making "Hey Copilot" a wake word for the Windows AI app. It's part of the company's effort to expand its presence across your PC.
Meet Copilot Actions, Windows 11's most revolutionary AI feature yet
When you purchase through links in our articles, we may earn a small commission. Microsoft wants to redefine the Windows AI PC. Copilot Actions is the first step. Microsoft's Copilot Actions is what happens when Microsoft begins rethinking the future of Windows and how AI is integrated into the operating system. Imagine agentic AI being turned loose inside your PC and performing tasks without your supervision.
Japan's government asks OpenAI to seek permission amid Sora 2 copyright concerns
In a time of both misinformation and too much information, quality journalism is more crucial than ever. By subscribing, you can help us get the story right. With your current subscription plan you can comment on stories. However, before writing your first comment, please create a display name in the Profile section of your subscriber account page. Your subscription plan doesn't allow commenting.
Deliberate Lab: A Platform for Real-Time Human-AI Social Experiments
Qian, Crystal, Tsai, Vivian, Behr, Michael, Hussein, Nada, Laugier, Léo, Thain, Nithum, Dixon, Lucas
Social and behavioral scientists increasingly aim to study how humans interact, collaborate, and make decisions alongside artificial intelligence. However, the experimental infrastructure for such work remains underdeveloped: (1) few platforms support real-time, multi-party studies at scale; (2) most deployments require bespoke engineering, limiting replicability and accessibility, and (3) existing tools do not treat AI agents as first-class participants. We present Deliberate Lab, an open-source platform for large-scale, real-time behavioral experiments that supports both human participants and large language model (LLM)-based agents. We report on a 12-month public deployment of the platform (N=88 experimenters, N=9195 experiment participants), analyzing usage patterns and workflows. Case studies and usage scenarios are aggregated from platform users, complemented by in-depth interviews with select experimenters. By lowering technical barriers and standardizing support for hybrid human-AI experimentation, Deliberate Lab expands the methodological repertoire for studying collective decision-making and human-centered AI.
Developing and Validating the Arabic Version of the Attitudes Toward Large Language Models Scale
Barajeeh, Basad, Yankouskaya, Ala, AlShakhsi, Sameha, Ho, Chun Sing Maxwell, Xu, Guandong, Ali, Raian
As the use of large language models (LLMs) becomes increasingly global, understanding public attitudes toward these systems requires tools that are adapted to local contexts and languages. In the Arab world, LLM adoption has grown rapidly with both globally dominant platforms and regional ones like Fanar and Jais offering Arabic-specific solutions. This highlights the need for culturally and linguistically relevant scales to accurately measure attitudes toward LLMs in the region. Tools assessing attitudes toward artificial intelligence (AI) can provide a base for measuring attitudes specific to LLMs. The 5-item Attitudes Toward Artificial Intelligence (ATAI) scale, which measures two dimensions, the AI Fear and the AI Acceptance, has been recently adopted and adapted to develop new instruments in English using a sample from the UK: the Attitudes Toward General LLMs (AT-GLLM) and Attitudes Toward Primary LLM (AT-PLLM) scales. In this paper, we translate the two scales, AT-GLLM and AT-PLLM, and validate them using a sample of 249 Arabic-speaking adults. The results show that the scale, translated into Arabic, is a reliable and valid tool that can be used for the Arab population and language. Psychometric analyses confirmed a two-factor structure, strong measurement invariance across genders, and good internal reliability. The scales also demonstrated strong convergent and discriminant validity. Our scales will support research in a non-Western context, a much-needed effort to help draw a global picture of LLM perceptions, and will also facilitate localized research and policy-making in the Arab region.