Large Language Model
OpenAI has released a new ChatGPT bot that you can talk to
The voice mode is powered by OpenAI's new GPT-4o model, which combines voice, text, and vision capabilities. To gather feedback, the company is initially launching the chatbot to a "small group of users" paying for ChatGPT Plus, but it says it will make the bot available to all ChatGPT Plus subscribers this fall. OpenAI says it will notify customers who are part of the first rollout wave in the ChatGPT app and provide instructions on how to use the new model. The new voice feature, which was announced in May, is being launched a month later than originally planned because the company said it needed more time to improve safety features, such as the model's ability to detect and refuse unwanted content. The company also said it was preparing its infrastructure to offer real-time responses to millions of users.
UK regulator looks at Google's partnership with Anthropic
The Competition and Markets Authority has begun a preliminary investigation into a partnership between Google and the AI startup Anthropic, marking the latest in a string of investigations into deals between big tech companies and smallerAI ones. Google invested 2bn (about 1.56bn) into Anthropic in 2023, shortly after signing a cloud computing agreement with the startup, which develops the Claude LLM and chatbot. The CMA is now considering whether the partnership has "resulted in the creation of a relevant merger situation" which would allow the agency to begin a formal investigation. It is inviting comments over the next two weeks. The move comes amid broader concerns about competition in the generative AI sector. A deal between Amazon and Anthropic is also being investigated by the CMA as a potential merger after Amazon took a 4bn stake in the company and signed a deal to become one of the startup's cloud computing providers.
Perplexity will put ads in its AI search engine and share revenue with publishers
When people type a question into Perplexity, the two-year-old search engine scours the internet and uses information from multiple sources, including online publishers, to synthesize an answer using AI. Soon, Perplexity will start sharing revenue with some publishers as part of an advertising platform it plans to launch around the end of September, the company announced on Tuesday. The initiative, known as the Perplexity Publishers' Program, comes less than two months after the San Francisco-based startup backed by investors like Jeff Bezos and NVIDIA, and valued at 3 billion, came under fire from Forbes, Wired, and Condé Nast for allegedly scraping content without permission and ignoring robots.txt, Perplexity's initial partners include TIME, Fortune, The Texas Tribune, Der Spiegel and Automattic, the company behind Wordpress.com. It's not clear exactly how much revenue Perplexity will share with publishers.
TechScape: Will OpenAI's 5bn gamble on chatbots pay off? Only if you use them
What if you build it and they don't come? The Guardian's journalism is independent. We will earn a commission if you buy something through an affiliate link. It's fair to say the shine is coming off the AI boom. Soaring valuations are starting to look unstable next to the sky-high spending required to sustain them.
How machines that can solve complex math problems might usher in more powerful AI
But the news item that really stood out to me was one that didn't get as much attention as it should have. It has the potential to usher in more powerful AI and scientific discovery than previously possible. Last Thursday, Google DeepMind announced it had built AI systems that can solve complex math problems. The systems--called AlphaProof and AlphaGeometry 2--worked together to successfully solve four out of six problems from this year's International Mathematical Olympiad, a prestigious competition for high school students. Their performance was the equivalent of winning a silver medal.
Diffusion Augmented Agents: A Framework for Efficient Exploration and Transfer Learning
Di Palo, Norman, Hasenclever, Leonard, Humplik, Jan, Byravan, Arunkumar
We introduce Diffusion Augmented Agents (DAAG), a novel framework that leverages large language models, vision language models, and diffusion models to improve sample efficiency and transfer learning in reinforcement learning for embodied agents. DAAG hindsight relabels the agent's past experience by using diffusion models to transform videos in a temporally and geometrically consistent way to align with target instructions with a technique we call Hindsight Experience Augmentation. The framework reduces the amount of rewardlabeled data needed to 1) finetune a vision language model that acts as a reward detector, and 2) train RL agents on new tasks. We demonstrate the sample efficiency gains of DAAG in simulated robotics environments involving manipulation and navigation. Our results show that DAAG improves learning of reward detectors, transferring past experience, and acquiring new tasks - key abilities for developing efficient lifelong learning agents. The most recent notable breakthroughs in AI have come from the combination of large models trained on enormous datasets (Firoozi et al., 2023; Brown et al., 2020; Hoffmann et al., 2022; Reed et al., 2022; Gemini-Team, 2023). However, despite efforts to scale up data collection (Collaboration, 2023; Reed et al., 2022; Bousmalis et al., 2023), data in embodied AI settings is still prohibitively scarce because such agents need to interact with physical environments where sensors and actuators present major bottlenecks (Cabi et al., 2020; Lee et al., 2022). This data scarcity issue is especially pronounced in reinforcement learning scenarios, where rewards are often sparse or completely absent in realistic settings (Ecoffet et al., 2021).
Cocobo: Exploring Large Language Models as the Engine for End-User Robot Programming
Ge, Yate, Dai, Yi, Shan, Run, Li, Kechun, Hu, Yuanda, Sun, Xiaohua
End-user development allows everyday users to tailor service robots or applications to their needs. One user-friendly approach is natural language programming. However, it encounters challenges such as an expansive user expression space and limited support for debugging and editing, which restrict its application in end-user programming. The emergence of large language models (LLMs) offers promising avenues for the translation and interpretation between human language instructions and the code executed by robots, but their application in end-user programming systems requires further study. We introduce Cocobo, a natural language programming system with interactive diagrams powered by LLMs. Cocobo employs LLMs to understand users' authoring intentions, generate and explain robot programs, and facilitate the conversion between executable code and flowchart representations. Our user study shows that Cocobo has a low learning curve, enabling even users with zero coding experience to customize robot programs successfully.
The Responsible Development of Automated Student Feedback with Generative AI
Lindsay, Euan D, Zhang, Mike, Johri, Aditya, Bjerva, Johannes
Abstract--Contribution: This paper identifies four critical ethical considerations for implementing generative AI tools to provide automated feedback to students. Background: Providing rich feedback to students is essential for supporting student learning. Recent advances in generative AI, particularly with large language models (LLMs), provide the opportunity to deliver repeatable, scalable and instant automatically generated feedback to students, making abundant a previously scarce and expensive learning resource. A visualisation of Bloom's revised taxonomy, modified from [6]. Intended Outcomes: The goal of this work is to enable the use of AI systems to automate mundane assessment and feedback tasks, without introducing a "tyranny of the majority", where HE release of powerful language technology tools based on generative language modelling (e.g., ChatGPT, GPT-are going to use AI tools in their working lives, we should 4(o), Claude, Gemini, Llama; [1]-[3]), marked a significant aim to train them in their use. For example, While assessment is a clear space of development for days after the release of ChatGPT, students, educators, and this type of educational technology, we argue that the real the public alike discovered the potential of the application potential of generative language modelling can be found in for assisting with a range of teaching and learning tasks, but student feedback. E. D. Lindsay is with the UNESCO Centre for Problem Based Learning M. Zhang is with the Department of Computer Science, Aalborg University, A.C. Meyers Vænge 15, 2450 København SV, Denmark. A. Johri is the Director of the Technocritical Research on AI, Learning J. Bjerva is with the Department of Computer Science, Aalborg University, Manuscript revised on July 31, 2024. Hence, this current state has common patterns of student answers and standardize responses effectively locked some engineering courses into a focus, to them, rather than having to make bespoke responses to where a particular set of questions are iterated over.
Adapting Safe-for-Work Classifier for Malaysian Language Text: Enhancing Alignment in LLM-Ops Framework
Razak, Aisyah, Nazhan, Ariff, Adha, Kamarul, Adzlan, Wan Adzhar Faiq, Ahmad, Mas Aisyah, Azman, Ammar
As large language models (LLMs) become increasingly integrated into operational workflows (LLM-Ops), there is a pressing need for effective guardrails to ensure safe and aligned interactions, including the ability to detect potentially unsafe or inappropriate content across languages. However, existing safe-for-work classifiers are primarily focused on English text. To address this gap for the Malaysian language, we present a novel safe-for-work text classifier tailored specifically for Malaysian language content. By curating and annotating a first-of-its-kind dataset of Malaysian text spanning multiple content categories, we trained a classification model capable of identifying potentially unsafe material using state-of-the-art natural language processing techniques. This work represents an important step in enabling safer interactions and content filtering to mitigate potential risks and ensure responsible deployment of LLMs.
Segment Anything for Videos: A Systematic Survey
Zhang, Chunhui, Cui, Yawen, Lin, Weilin, Huang, Guanjie, Rong, Yan, Liu, Li, Shan, Shiguang
The recent wave of foundation models has witnessed tremendous success in computer vision (CV) and beyond, with the segment anything model (SAM) having sparked a passion for exploring task-agnostic visual foundation models. Empowered by its remarkable zero-shot generalization, SAM is currently challenging numerous traditional paradigms in CV, delivering extraordinary performance not only in various image segmentation and multi-modal segmentation (\eg, text-to-mask) tasks, but also in the video domain. Additionally, the latest released SAM 2 is once again sparking research enthusiasm in the realm of promptable visual segmentation for both images and videos. However, existing surveys mainly focus on SAM in various image processing tasks, a comprehensive and in-depth review in the video domain is notably absent. To address this gap, this work conducts a systematic review on SAM for videos in the era of foundation models. As the first to review the progress of SAM for videos, this work focuses on its applications to various tasks by discussing its recent advances, and innovation opportunities of developing foundation models on broad applications. We begin with a brief introduction to the background of SAM and video-related research domains. Subsequently, we present a systematic taxonomy that categorizes existing methods into three key areas: video understanding, video generation, and video editing, analyzing and summarizing their advantages and limitations. Furthermore, comparative results of SAM-based and current state-of-the-art methods on representative benchmarks, as well as insightful analysis are offered. Finally, we discuss the challenges faced by current research and envision several future research directions in the field of SAM for video and beyond.