Large Language Model
Join me at EmTech Digital this week!
Between the world leaders gathering in Seoul for the second AI Safety Summit this week and Google and OpenAI's launches of their supercharged new models, Astra and GPT-4o, the timing could not be better. AI feels hotter than ever. This year's EmTech will be all about how we can harness the power of generative AI while mitigating its risks,and how the technology will affect the workforce, competitiveness, and democracy. We will also get a sneak peek into the AI labs of Google, OpenAI, Adobe, AWS, and others. This year's top speakers include Nick Clegg, the president of global affairs at Meta, who will talk about what the platform intends to do to curb misinformation.
Scarlett Johansson 'Angered' By ChatGPT Voice That Sounded 'Eerily' Like Her
Scarlett Johansson said Monday that she was "shocked, angered and in disbelief" when she heard that OpenAI used a voice "eerily similar" to hers for its new ChatGPT 4.0 chatbot, even after she had declined to provide her voice. Earlier on Monday, OpenAI announced on X that it would pause the AI voice, known as "Sky," while it addresses "questions about how we chose the voices in ChatGPT." The company said in a blog post that the "Sky" voice was "not an imitation" of Johansson's voice, but that it was recorded by a different professional actor, whose identity the company would not reveal to protect her privacy. But Johansson said in a statement to NPR on Monday that OpenAI's Chief Executive Officer Sam Altman had asked her in September to voice the ChatGPT 4.0 system because he thought her "voice would be comforting to people." She declined, but nine months later, her friends, family and the public noticed how the "Sky" voice resembled hers.
EchoPT: A Pretrained Transformer Architecture that Predicts 2D In-Air Sonar Images for Mobile Robotics
Steckel, Jan, Jansen, Wouter, Huebel, Nico
The predictive brain hypothesis suggests that perception can be interpreted as the process of minimizing the error between predicted perception tokens generated by an internal world model and actual sensory input tokens. When implementing working examples of this hypothesis in the context of in-air sonar, significant difficulties arise due to the sparse nature of the reflection model that governs ultrasonic sensing. Despite these challenges, creating consistent world models using sonar data is crucial for implementing predictive processing of ultrasound data in robotics. In an effort to enable robust robot behavior using ultrasound as the sole exteroceptive sensor modality, this paper introduces EchoPT, a pretrained transformer architecture designed to predict 2D sonar images from previous sensory data and robot ego-motion information. We detail the transformer architecture that drives EchoPT and compare the performance of our model to several state-of-the-art techniques. In addition to presenting and evaluating our EchoPT model, we demonstrate the effectiveness of this predictive perception approach in two robotic tasks.
Towards Retrieval-Augmented Architectures for Image Captioning
Sarto, Sara, Cornia, Marcella, Baraldi, Lorenzo, Nicolosi, Alessandro, Cucchiara, Rita
The objective of image captioning models is to bridge the gap between the visual and linguistic modalities by generating natural language descriptions that accurately reflect the content of input images. In recent years, researchers have leveraged deep learning-based models and made advances in the extraction of visual features and the design of multimodal connections to tackle this task. This work presents a novel approach towards developing image captioning models that utilize an external kNN memory to improve the generation process. Specifically, we propose two model variants that incorporate a knowledge retriever component that is based on visual similarities, a differentiable encoder to represent input images, and a kNN-augmented language model to predict tokens based on contextual cues and text retrieved from the external memory. We experimentally validate our approach on COCO and nocaps datasets and demonstrate that incorporating an explicit external memory can significantly enhance the quality of captions, especially with a larger retrieval corpus. This work provides valuable insights into retrieval-augmented captioning models and opens up new avenues for improving image captioning at a larger scale.
Towards Responsible Development of Generative AI for Education: An Evaluation-Driven Approach
Jurenka, Irina, Kunesch, Markus, McKee, Kevin R., Gillick, Daniel, Zhu, Shaojian, Wiltberger, Sara, Phal, Shubham Milind, Hermann, Katherine, Kasenberg, Daniel, Bhoopchand, Avishkar, Anand, Ankit, Pรฎslar, Miruna, Chan, Stephanie, Wang, Lisa, She, Jennifer, Mahmoudieh, Parsa, Rysbek, Aliya, Ko, Wei-Jen, Huber, Andrea, Wiltshire, Brett, Elidan, Gal, Rabin, Roni, Rubinovitz, Jasmin, Pitaru, Amit, McAllister, Mac, Wilkowski, Julia, Choi, David, Engelberg, Roee, Hackmon, Lidan, Levin, Adva, Griffin, Rachel, Sears, Michael, Bar, Filip, Mesar, Mia, Jabbour, Mana, Chaudhry, Arslan, Cohan, James, Thiagarajan, Sridhar, Levine, Nir, Brown, Ben, Gorur, Dilan, Grant, Svetlana, Hashimoshoni, Rachel, Weidinger, Laura, Hu, Jieru, Chen, Dawn, Dolecki, Kuba, Akbulut, Canfer, Bileschi, Maxwell, Culp, Laura, Dong, Wen-Xin, Marchal, Nahema, Van Deman, Kelsie, Misra, Hema Bajaj, Duah, Michael, Ambar, Moran, Caciularu, Avi, Lefdal, Sandra, Summerfield, Chris, An, James, Kamienny, Pierre-Alexandre, Mohdi, Abhinit, Strinopoulous, Theofilos, Hale, Annie, Anderson, Wayne, Cobo, Luis C., Efron, Niv, Ananda, Muktha, Mohamed, Shakir, Heymans, Maureen, Ghahramani, Zoubin, Matias, Yossi, Gomes, Ben, Ibrahim, Lila
A major challenge facing the world is the provision of equitable and universal access to quality education. Recent advances in generative AI (gen AI) have created excitement about the potential of new technologies to offer a personal tutor for every learner and a teaching assistant for every teacher. The full extent of this dream, however, has not yet materialised. We argue that this is primarily due to the difficulties with verbalising pedagogical intuitions into gen AI prompts and the lack of good evaluation practices, reinforced by the challenges in defining excellent pedagogy. Here we present our work collaborating with learners and educators to translate high level principles from learning science into a pragmatic set of seven diverse educational benchmarks, spanning quantitative, qualitative, automatic and human evaluations; and to develop a new set of fine-tuning datasets to improve the pedagogical capabilities of Gemini, introducing LearnLM-Tutor. Our evaluations show that LearnLM-Tutor is consistently preferred over a prompt tuned Gemini by educators and learners on a number of pedagogical dimensions. We hope that this work can serve as a first step towards developing a comprehensive educational evaluation framework, and that this can enable rapid progress within the AI and EdTech communities towards maximising the positive impact of gen AI in education.
Securing the Future of GenAI: Policy and Technology
Christodorescu, Mihai, Craven, Ryan, Feizi, Soheil, Gong, Neil, Hoffmann, Mia, Jha, Somesh, Jiang, Zhengyuan, Kamarposhti, Mehrdad Saberi, Mitchell, John, Newman, Jessica, Probasco, Emelia, Qi, Yanjun, Shams, Khawaja, Turek, Matthew
The rise of Generative AI (GenAI) brings about transformative potential across sectors, but its dual-use nature also amplifies risks. Governments globally are grappling with the challenge of regulating GenAI, balancing innovation against safety. China, the United States (US), and the European Union (EU) are at the forefront with initiatives like the Management of Algorithmic Recommendations, the Executive Order, and the AI Act, respectively. However, the rapid evolution of GenAI capabilities often outpaces the development of comprehensive safety measures, creating a gap between regulatory needs and technical advancements. A workshop co-organized by Google, University of Wisconsin, Madison (UW-Madison), and Stanford University aimed to bridge this gap between GenAI policy and technology. The diverse stakeholders of the GenAI space -- from the public and governments to academia and industry -- make any safety measures under consideration more complex, as both technical feasibility and regulatory guidance must be realized. This paper summarizes the discussions during the workshop which addressed questions, such as: How regulation can be designed without hindering technological progress? How technology can evolve to meet regulatory standards? The interplay between legislation and technology is a very vast topic, and we don't claim that this paper is a comprehensive treatment on this topic. This paper is meant to capture findings based on the workshop, and hopefully, can guide discussion on this topic.
AI in Manufacturing: Market Analysis and Opportunities
In this paper, we explore the transformative impact of Artificial Intelligence (AI) in the manufacturing sector, highlighting its potential to revolutionize industry practices and enhance operational efficiency. We delve into various applications of AI in manufacturing, with a particular emphasis on human-machine interfaces (HMI) and AI-powered milling machines, showcasing how these technologies contribute to more intuitive operations and precision in production processes. Through rigorous market analysis, the paper presents insightful data on AI adoption rates among German manufacturers, comparing these figures with global trends and exploring the specific uses of AI in production, maintenance, customer service, and more. In addition, the paper examines the emerging field of Generative AI and the potential applications of large language models in manufacturing processes. The findings indicate a significant increase in AI adoption from 6% in 2020 to 13.3% in 2023 among German companies, with a projection of substantial economic impact by 2030. The study also addresses the challenges faced by companies, such as data quality and integration hurdles, providing a balanced view of the opportunities and obstacles in AI implementation.
EyeFound: A Multimodal Generalist Foundation Model for Ophthalmic Imaging
Shi, Danli, Zhang, Weiyi, Chen, Xiaolan, Liu, Yexin, Yang, Jiancheng, Huang, Siyu, Tham, Yih Chung, Zheng, Yingfeng, He, Mingguang
Artificial intelligence (AI) is vital in ophthalmology, tackling tasks like diagnosis, classification, and visual question answering (VQA). However, existing AI models in this domain often require extensive annotation and are task-specific, limiting their clinical utility. While recent developments have brought about foundation models for ophthalmology, they are limited by the need to train separate weights for each imaging modality, preventing a comprehensive representation of multi-modal features. This highlights the need for versatile foundation models capable of handling various tasks and modalities in ophthalmology. To address this gap, we present EyeFound, a multimodal foundation model for ophthalmic images. Unlike existing models, EyeFound learns generalizable representations from unlabeled multimodal retinal images, enabling efficient model adaptation across multiple applications. Trained on 2.78 million images from 227 hospitals across 11 ophthalmic modalities, EyeFound facilitates generalist representations and diverse multimodal downstream tasks, even for detecting challenging rare diseases. It outperforms previous work RETFound in diagnosing eye diseases, predicting systemic disease incidents, and zero-shot multimodal VQA. EyeFound provides a generalizable solution to improve model performance and lessen the annotation burden on experts, facilitating widespread clinical AI applications for retinal imaging.
More Distinctively Black and Feminine Faces Lead to Increased Stereotyping in Vision-Language Models
Lee, Messi H. J., Montgomery, Jacob M., Lai, Calvin K.
Vision Language Models (VLMs), exemplified by GPT-4V, adeptly integrate text and vision modalities. This integration enhances Large Language Models' ability to mimic human perception, allowing them to process image inputs. Despite VLMs' advanced capabilities, however, there is a concern that VLMs inherit biases of both modalities in ways that make biases more pervasive and difficult to mitigate. Our study explores how VLMs perpetuate homogeneity bias and trait associations with regards to race and gender. When prompted to write stories based on images of human faces, GPT-4V describes subordinate racial and gender groups with greater homogeneity than dominant groups and relies on distinct, yet generally positive, stereotypes. Importantly, VLM stereotyping is driven by visual cues rather than group membership alone such that faces that are rated as more prototypically Black and feminine are subject to greater stereotyping. These findings suggest that VLMs may associate subtle visual cues related to racial and gender groups with stereotypes in ways that could be challenging to mitigate. We explore the underlying reasons behind this behavior and discuss its implications and emphasize the importance of addressing these biases as VLMs come to mirror human perception.
Bring Your Own KG: Self-Supervised Program Synthesis for Zero-Shot KGQA
Agarwal, Dhruv, Das, Rajarshi, Khosla, Sopan, Gangadharaiah, Rashmi
We present BYOKG, a universal question-answering (QA) system that can operate on any knowledge graph (KG), requires no human-annotated training data, and can be ready to use within a day -- attributes that are out-of-scope for current KGQA systems. BYOKG draws inspiration from the remarkable ability of humans to comprehend information present in an unseen KG through exploration -- starting at random nodes, inspecting the labels of adjacent nodes and edges, and combining them with their prior world knowledge. In BYOKG, exploration leverages an LLM-backed symbolic agent that generates a diverse set of query-program exemplars, which are then used to ground a retrieval-augmented reasoning procedure to predict programs for arbitrary questions. BYOKG is effective over both small- and large-scale graphs, showing dramatic gains in QA accuracy over a zero-shot baseline of 27.89 and 58.02 F1 on GrailQA and MetaQA, respectively. On GrailQA, we further show that our unsupervised BYOKG outperforms a supervised in-context learning method, demonstrating the effectiveness of exploration. Lastly, we find that performance of BYOKG reliably improves with continued exploration as well as improvements in the base LLM, notably outperforming a state-of-the-art fine-tuned model by 7.08 F1 on a sub-sampled zero-shot split of GrailQA.