Generative AI
"Is ChatGPT a Better Explainer than My Professor?": Evaluating the Explanation Capabilities of LLMs in Conversation Compared to a Human Baseline
Li, Grace, Alshomary, Milad, Muresan, Smaranda
Explanations form the foundation of knowledge sharing and build upon communication principles, social dynamics, and learning theories. We focus specifically on conversational approaches for explanations because the context is highly adaptive and interactive. Our research leverages previous work on explanatory acts, a framework for understanding the different strategies that explainers and explainees employ in a conversation to both explain, understand, and engage with the other party. We use the 5-Levels dataset was constructed from the WIRED YouTube series by Wachsmuth et al., and later annotated by Booshehri et al. with explanatory acts. These annotations provide a framework for understanding how explainers and explainees structure their response when crafting a response. With the rise of generative AI in the past year, we hope to better understand the capabilities of Large Language Models (LLMs) and how they can augment expert explainer's capabilities in conversational settings. To achieve this goal, the 5-Levels dataset (We use Booshehri et al.'s 2023 annotated dataset with explanatory acts.) allows us to audit the ability of LLMs in engaging in explanation dialogues. To evaluate the effectiveness of LLMs in generating explainer responses, we compared 3 different strategies, we asked human annotators to evaluate 3 different strategies: human explainer response, GPT4 standard response, GPT4 response with Explanation Moves.
Generative artificial intelligence in ophthalmology: multimodal retinal images for the diagnosis of Alzheimer's disease with convolutional neural networks
Slootweg, I. R., Thach, M., Curro-Tafili, K. R., Verbraak, F. D., Bouwman, F. H., Pijnenburg, Y. A. L., Boer, J. F., de Kwisthout, J. H. P., Bagheriye, L., Gonzรกlez, P. J.
Background/Aim. This study aims to predict Amyloid Positron Emission Tomography (AmyloidPET) status with multimodal retinal imaging and convolutional neural networks (CNNs) and to improve the performance through pretraining with synthetic data. Methods. Fundus autofluorescence, optical coherence tomography (OCT), and OCT angiography images from 328 eyes of 59 AmyloidPET positive subjects and 108 AmyloidPET negative subjects were used for classification. Denoising Diffusion Probabilistic Models (DDPMs) were trained to generate synthetic images and unimodal CNNs were pretrained on synthetic data and finetuned on real data or trained solely on real data. Multimodal classifiers were developed to combine predictions of the four unimodal CNNs with patient metadata. Class activation maps of the unimodal classifiers provided insight into the network's attention to inputs. Results. DDPMs generated diverse, realistic images without memorization. Pretraining unimodal CNNs with synthetic data improved AUPR at most from 0.350 to 0.579. Integration of metadata in multimodal CNNs improved AUPR from 0.486 to 0.634, which was the best overall best classifier. Class activation maps highlighted relevant retinal regions which correlated with AD. Conclusion. Our method for generating and leveraging synthetic data has the potential to improve AmyloidPET prediction from multimodal retinal imaging. A DDPM can generate realistic and unique multimodal synthetic retinal images. Our best performing unimodal and multimodal classifiers were not pretrained on synthetic data, however pretraining with synthetic data slightly improved classification performance for two out of the four modalities.
BASS: Batched Attention-optimized Speculative Sampling
Qian, Haifeng, Gonugondla, Sujan Kumar, Ha, Sungsoo, Shang, Mingyue, Gouda, Sanjay Krishna, Nallapati, Ramesh, Sengupta, Sudipta, Ma, Xiaofei, Deoras, Anoop
Speculative decoding has emerged as a powerful method to improve latency and throughput in hosting large language models. However, most existing implementations focus on generating a single sequence. Real-world generative AI applications often require multiple responses and how to perform speculative decoding in a batched setting while preserving its latency benefits poses non-trivial challenges. This paper describes a system of batched speculative decoding that sets a new state of the art in multi-sequence generation latency and that demonstrates superior GPU utilization as well as quality of generations within a time budget. For example, for a 7.8B-size model on a single A100 GPU and with a batch size of 8, each sequence is generated at an average speed of 5.8ms per token, the overall throughput being 1.1K tokens per second. These results represent state-of-the-art latency and a 2.15X speed-up over optimized regular decoding. Within a time budget that regular decoding does not finish, our system is able to generate sequences with HumanEval Pass@First of 43% and Pass@All of 61%, far exceeding what's feasible with single-sequence speculative decoding. Our peak GPU utilization during decoding reaches as high as 15.8%, more than 3X the highest of that of regular decoding and around 10X of single-sequence speculative decoding.
OpenAI has delayed its seductive ChatGPT voice assistants
If you've been dreaming about spending your summer whispering sweet nothings into the digital ears of one of the seductive ChatGPT voice assistants that OpenAI showed off last month, you'll have to dream a little longer. On Tuesday, the company announced that its "advanced Voice Mode" feature needs more time in the oven "to reach our bar to launch." The feature will be available to a small group of users to gather feedback, and then launch to all paying ChatGPT customers in the fall. "We're improving the model's ability to detect and refuse certain content," OpenAI posted on X. "We're also working on improving the user experience and preparing our infrastructure to scale to millions while maintaining real-time responses." We're sharing an update on the advanced Voice Mode we demoed during our Spring Update, which we remain very excited about: We had planned to start rolling this out in alpha to a small group of ChatGPT Plus users in late June, but need one more month to reach our bar to launch.โฆ Voices have been a part of ChatGPT since 2023.
Toys 'R' Us uses OpenAI's Sora to make a brand film about its origin story and it's horrifying
The rise of artificial intelligence in our media and entertainment industries has raised a lot of concerns about programs like Open Al's text-to-video maker Sora replacing the artistic endeavors and aspirations of humans. If those AI made movies are anything like a new brand film about the Toys'R' Us toy store chain's origin story, the only thing we'll have to fear is watching them. Toys'R' Us's current owner WHP Global worked with the Emmy nominated creative agency Native Foreign to create a short brand film called The Origin of Toys'R' Us using OpenAI's text-to-video creator Sora. The film premiered at the 2024 Cannes Lions International Festival of Creativity and can currently be viewed on the toy retailer's website. The Origin of Toys'R' Us is only a little over a minute long but it's a mix of confusing and eerie.
ChatGPT for macOS no longer requires a subscription
The macOS ChatGPT desktop app is now available to everyone. That is, provided you're running an Apple Silicon Mac (sorry, Intel users) and your computer is on macOS Sonoma or higher. OpenAI rolled out the app gradually, starting with Plus subscribers last month. ChatGPT now has an official macOS client before it has a Windows one. Of course, Windows 11 has the OpenAI-powered Microsoft CoPilot baked into its OS, which likely explains the omission.
OpenAI will block people in China from using its services
OpenAI plans to block people from using ChatGPT in China, a country where its services aren't officially available, but where users and developers access it via the company's API anyway. Securities Times, a Chinese state-owned newspaper reported on Tuesday that OpenAI had started sending emails to users in China outlining its plans to block access starting July 9, according to Reuters. "We are taking additional taps to block API traffic from regions where we do not support access to OpenAI's services," an OpenAI spokesperson told the publication. The move could impact several Chinese startups which have built applications using OpenAI's large language models. Although OpenAI's services are available in more than 160 countries, China isn't one of them.
Claude 3.5 suggests AI's looming ubiquity could be a good thing
The frontier of AI just got pushed a little further forward. On Friday, Anthropic, the AI lab set up by a team of disgruntled OpenAI staffers, released the latest version of its Claude LLM. The company said Thursday that the new model โ the technology that underpins its popular chatbot Claude โ is twice as fast as its most powerful previous version. Anthropic said in its evaluations, the model outperforms leading competitors like OpenAI on several key intelligence capabilities, such as coding and text-based reasoning. Anthropic only released the previous version of Claude, 3.0, in March.
Generative AI Systems: A Systems-based Perspective on Generative AI
Large Language Models (LLMs) have revolutionized AI systems by enabling communication with machines using natural language. Recent developments in Generative AI (GenAI) like Vision-Language Models (GPT-4V) and Gemini have shown great promise in using LLMs as multimodal systems. This new research line results in building Generative AI systems, GenAISys for short, that are capable of multimodal processing and content creation, as well as decision-making. GenAISys use natural language as a communication means and modality encoders as I/O interfaces for processing various data sources. They are also equipped with databases and external specialized tools, communicating with the system through a module for information retrieval and storage. This paper aims to explore and state new research directions in Generative AI Systems, including how to design GenAISys (compositionality, reliability, verifiability), build and train them, and what can be learned from the system-based perspective. Cross-disciplinary approaches are needed to answer open questions about the inner workings of GenAI systems.
Transforming Software Development: Evaluating the Efficiency and Challenges of GitHub Copilot in Real-World Projects
Pandey, Ruchika, Singh, Prabhat, Wei, Raymond, Shankar, Shaila
Generative AI technologies promise to transform the product development lifecycle. This study evaluates the efficiency gains, areas for improvement, and emerging challenges of using GitHub Copilot, an AI-powered coding assistant. We identified 15 software development tasks and assessed Copilot's benefits through real-world projects on large proprietary code bases. Our findings indicate significant reductions in developer toil, with up to 50% time saved in code documentation and autocompletion, and 30-40% in repetitive coding tasks, unit test generation, debugging, and pair programming. However, Copilot struggles with complex tasks, large functions, multiple files, and proprietary contexts, particularly with C/C++ code. We project a 33-36% time reduction for coding-related tasks in a cloud-first software development lifecycle. This study aims to quantify productivity improvements, identify underperforming scenarios, examine practical benefits and challenges, investigate performance variations across programming languages, and discuss emerging issues related to code quality, security, and developer experience.