Goto

Collaborating Authors

 Large Language Model


OpenAI Goes MAGA

The Atlantic - Technology

Things were not looking great for OpenAI at the end of last year. The company had been struggling with major delays on its long-awaited GPT-5 and hemorrhaging key talent--notably, Chief Scientist Ilya Sutskever, Chief Technology Officer Mira Murati, and Alec Radford, the researcher who'd set the company on the path of developing GPTs in the first place. Several people who left either joined OpenAI competitors or launched new ones. The start-up's relationship with Microsoft, its biggest backer and a crucial provider of the computing infrastructure needed to train and deploy its AI models, was being investigated by the Federal Trade Commission. And then there was Elon Musk.


OpenAI's Operator Lets ChatGPT Use the Web for You

WIRED

OpenAI is letting some users try a new ChatGPT feature that uses its artificial intelligence to operate a web browser to book trips, buy groceries, hunt for bargains, and do many other online chores. The new tool, called Operator, is an AI agent: It relies on an AI model trained on both text and images to interpret commands and figure out how to use a web browser to execute them. OpenAI claims it has the potential to automate many day-to-day tasks and workday errands. OpenAI's Operator follows rival releases by both Google and Anthropic, which have demonstrated ones capable of using the web. AI agents are widely seen as the next evolutionary stage for AI following chatbots, and many companies have hopped on the hype train by touting them.


OpenAI launches Operator--an agent that can use a computer for you

MIT Technology Review

OpenAI claims that Operator outperforms similar rival tools, including Anthropic's Computer Use (a version of Claude 3.5 Sonnet that can carry out simple tasks on a computer) and Google DeepMind's Mariner (a web-browsing agent built on top of Gemini 2.0). The fact that three of the world's top AI firms have converged on the same vision of what agent-based models could be makes one thing clear. The battle for AI supremacy has a new frontier--and it's our computer screens. "Moving from generating text and images to doing things is the right direction," says Ali Farhadi, CEO of the Allen Institute for AI (AI2). "It unlocks business, solves new problems."


ChatGPT down as thousands report issues worldwide

BBC News

Thousands of people have reported they are unable to use the world's best-known artificial intelligence (AI) tool, ChatGPT Downdetector, which tracks website outages, showed more than 10,000 people reported the AI chatbot, which is made by OpenAI, was not working in the UK on Thursday. Users attempting to use it were met with a message reading "the web server reported a bad gateway error". OpenAI has not yet publicly commented on the outage. The BBC has contacted the firm for comment. However, many people have taken to social media to highlight the problem and the disruption it is causing them.


ChatGPT is down across the US as users report 'bad gateway' error when using the AI tool

Daily Mail - Science & tech

ChatGPT is down across the US as users report seeing a'bad gateway' error message when using the AI tool. Downdetector, a site that monitors online outages, shows issues hit the OpenAI-owned platform around 7am ET. The error message indicates that one server received an invalid response from another, creating a communication breakdown. Many users have shared their frustrations on X, saying they'feel lost without it.' 'Seriously, how am I supposed to brainstorm, write, and research without my AI assistant,' an X user posted.


Review for NeurIPS paper: Modeling Task Effects on Meaning Representation in the Brain via Zero-Shot MEG Prediction

Neural Information Processing Systems

Summary and Contributions: This paper presents a re-analysis of the MEG experiment of Sudre et al (2012), where participants were tasked with responding to a question about the meaning of an object concept word (e.g. In the original Sudre et al analysis, the focus was on testing the predictive power of different perceptual and semantic feature models of the concept word for the MEG data. In the current study, the focus is on the role of the task question that precedes the concept word, and in particular whether and how the semantics of the task question modulates the subsequent processing and neural activity time-locked to the stimulus word. This is an interesting neurocognitive question, as it sheds light on how lexical-semantic representation and access can be modulated by the preceding context, and how the timing of processing of the target concept word that is independent of the task demands relates to the timing of the processing that involves integrating that conceptual knowledge with the task requirements in order to respond on the task. To analyze the data, the authors construct vector-based semantic models of both the concept words and the task questions, using human responses from separate questions and concepts where the participants rated the truth of the task questions for the concepts.


Review for NeurIPS paper: Modeling Task Effects on Meaning Representation in the Brain via Zero-Shot MEG Prediction

Neural Information Processing Systems

Understanding how the tasks that we perform while perceiving a stimulus modulate brain activity is of wide interest to neuroscience. Reviewers found the experimental setup interesting, with the clear hypotheses about how tasks can impact neural activity. The small effect sizes observed were identified as a key limitation in drawing conclusions from this experiment. Reviewers found this worrisome particularly when coupled with marginal accuracies and few subjects. Reviewers identified that BERT results were not very conclusive, and significantly more could be done.


Revealed: Microsoft deepened ties with Israeli military to provide tech support during Gaza war

The Guardian

The Israeli military's reliance on Microsoft's cloud technology and artificial intelligence systems surged during the most intensive phase of its bombardment of Gaza, leaked documents reveal. The files offer an inside view of how Microsoft deepened its relationship with Israel's defence establishment after 7 October 2023, supplying the military with greater computing and storage services and striking at least 10m in deals to provide thousands of hours of technical support. Microsoft's deep ties with Israel's military are revealed in an investigation by the Guardian with the Israeli-Palestinian publication 972 Magazine and a Hebrew-language outlet, Local Call. It is based in part on documents obtained by Drop Site News, which has published its own story. The investigation, which also draws on interviews with sources from across Israel's defence and intelligence establishment, sheds new light on how the Israel Defense Forces (IDF) turned to major US tech companies to meet the technological demands of war. After launching its offensive in Gaza in October 2023, the IDF faced a sudden rush in demand for storage and computing power, leading it to swiftly expand its computing infrastructure and embrace what one commander described as "the wonderful world of cloud providers".


Reasoning Language Models: A Blueprint

arXiv.org Artificial Intelligence

Reasoning language models (RLMs), also known as Large Reasoning Models (LRMs), such as OpenAI's o1 and o3, DeepSeek-V3, and Alibaba's QwQ, have redefined AI's problem-solving capabilities by extending LLMs with advanced reasoning mechanisms. Yet, their high costs, proprietary nature, and complex architectures - uniquely combining Reinforcement Learning (RL), search heuristics, and LLMs - present accessibility and scalability challenges. To address these, we propose a comprehensive blueprint that organizes RLM components into a modular framework, based on a survey and analysis of all RLM works. This blueprint incorporates diverse reasoning structures (chains, trees, graphs, and nested forms), reasoning strategies (e.g., Monte Carlo Tree Search, Beam Search), RL concepts (policy, value models and others), supervision schemes (Outcome-Based and Process-Based Supervision), and other related concepts (e.g., Test-Time Compute, Retrieval-Augmented Generation, agent tools). We also provide detailed mathematical formulations and algorithmic specifications to simplify RLM implementation. By showing how schemes like LLaMA-Berry, QwQ, Journey Learning, and Graph of Thoughts fit as special cases, we demonstrate the blueprint's versatility and unifying potential. To illustrate its utility, we introduce x1, a modular implementation for rapid RLM prototyping and experimentation. Using x1 and a literature review, we provide key insights, such as multi-phase training for policy and value models, and the importance of familiar training distributions. Finally, we discuss scalable RLM cloud deployments and we outline how RLMs can integrate with a broader LLM ecosystem. Our work demystifies RLM construction, democratizes advanced reasoning capabilities, and fosters innovation, aiming to mitigate the gap between "rich AI" and "poor AI" by lowering barriers to RLM design and experimentation.


Mining Social Determinants of Health for Heart Failure Patient 30-Day Readmission via Large Language Model

arXiv.org Artificial Intelligence

Heart Failure (HF) affects millions of Americans and leads to high readmission rates, posing significant healthcare challenges. While Social Determinants of Health (SDOH) such as socioeconomic status and housing stability play critical roles in health outcomes, they are often underrepresented in structured EHRs and hidden in unstructured clinical notes. This study leverages advanced large language models (LLMs) to extract SDOHs from clinical text and uses logistic regression to analyze their association with HF readmissions.