Generative AI
What to expect from Microsoft Build 2024: The Surface event, Windows 11 and AI
If you can't tell by now, just about every tech company is eager to pray at the altar of AI, for better or worse. Google's recent I/O developer conference was dominated by AI features, like its seemingly life-like Project Astra assistant. Just before that, OpenAI debuted GPT 4o, a free and conversational AI model that's disturbingly flirty. Next up is Microsoft Build 2024, the company's developer conference that's kicking off next week in Seattle. Normally, Build is a fairly straightforward celebration of Microsoft's devotion to productivity, with a dash of on-stage coding to excite the developer crowd.
Unlocking the trillion-dollar potential of generative AI
In this session, experts from Amazon Web Services (AWS) and QuantumBlack, AI by McKinsey, discuss the drivers fueling the massive potential impact of generative AI. Plus, they look at key industries set to capture the largest share of this value and practical strategies for effectively upskilling their workforces to take advantage of these productivity gains. Learn how to seamlessly integrate generative AI into your organization's workflows while fostering a skilled and adaptable workforce. Register now to learn how to unlock the trillion-dollar potential of generative AI.
What's up with ChatGPT's new sexy persona? Arwa Mahdawi
"Any sufficiently advanced technology is indistinguishable from magic," Arthur C Clarke famously said. And this could certainly be said of the impressive OpenAI update to ChatGPT, called GPT-4o, which was released on Monday. With the slight caveat that it felt a lot like the magician was a horny 12-year-old boy who had just watched the Spike Jonze movie Her. If you aren't up to speed on GPT-4o (the o stands for "omni") it's basically an all-singing, all-dancing, all-seeing version of the original chatbot. It can give you advice, it can rate your jokes, it can describe your surroundings, it can banter with you.
How Far Are We From AGI
Feng, Tao, Jin, Chuanyang, Liu, Jingyu, Zhu, Kunlun, Tu, Haoqin, Cheng, Zirui, Lin, Guanyu, You, Jiaxuan
The evolution of artificial intelligence (AI) has profoundly impacted human society, driving significant advancements in multiple sectors. Yet, the escalating demands on AI have highlighted the limitations of AI's current offerings, catalyzing a movement towards Artificial General Intelligence (AGI). AGI, distinguished by its ability to execute diverse real-world tasks with efficiency and effectiveness comparable to human intelligence, reflects a paramount milestone in AI evolution. While existing works have summarized specific recent advancements of AI, they lack a comprehensive discussion of AGI's definitions, goals, and developmental trajectories. Different from existing survey papers, this paper delves into the pivotal questions of our proximity to AGI and the strategies necessary for its realization through extensive surveys, discussions, and original perspectives. We start by articulating the requisite capability frameworks for AGI, integrating the internal, interface, and system dimensions. As the realization of AGI requires more advanced capabilities and adherence to stringent constraints, we further discuss necessary AGI alignment technologies to harmonize these factors. Notably, we emphasize the importance of approaching AGI responsibly by first defining the key levels of AGI progression, followed by the evaluation framework that situates the status-quo, and finally giving our roadmap of how to reach the pinnacle of AGI. Moreover, to give tangible insights into the ubiquitous impact of the integration of AI, we outline existing challenges and potential pathways toward AGI in multiple domains. In sum, serving as a pioneering exploration into the current state and future trajectory of AGI, this paper aims to foster a collective comprehension and catalyze broader public discussions among researchers and practitioners on AGI.
GPT Store Mining and Analysis
Su, Dongxun, Zhao, Yanjie, Hou, Xinyi, Wang, Shenao, Wang, Haoyu
As a pivotal extension of the renowned ChatGPT, the GPT The development of Large Language Models (LLMs) has been Store serves as a dynamic marketplace for various Generative a transformative force in human life, reshaping interactions, Pre-trained Transformer (GPT) models, shaping the frontier enhancing communication, and influencing decision-making of conversational AI. This paper presents an in-depth measurement processes. A notable manifestation of this impact is ChatGPT, study of the GPT Store, with a focus on the categorization which, since its inception, has garnered widespread popularity, of GPTs by topic, factors influencing GPT popularity, evidenced by its millions of active users and its profound and the potential security risks. Our investigation starts with integration into various sectors such as education, business, assessing the categorization of GPTs in the GPT Store, analyzing and entertainment [17]. This surge in popularity not only how they are organized by topics, and evaluating the highlights the effectiveness of ChatGPT in understanding effectiveness of the classification system. We then examine and generating human-like text but also underscores the the factors that affect the popularity of specific GPTs, looking growing public interest in AI-driven solutions.
IGOT: Information Gain Optimized Tokenizer on Domain Adaptive Pretraining
Feng, Dawei, Zhang, Yihai, Xu, Zhixuan
Pretrained Large Language Models (LLM) such as ChatGPT, Claude, etc. have demonstrated strong capabilities in various fields of natural language generation. However, there are still many problems when using LLM in specialized domain-specific fields. When using generative AI to process downstream tasks, a common approach is to add new knowledge (e.g., private domain knowledge, cutting-edge information) to a pretrained model through continued training or fine-tuning. However, whether there is a universal paradigm for domain adaptation training is still an open question. In this article, we proposed Information Gain Optimized Tokenizer (IGOT), which analyzes the special token set of downstream tasks, constructs a new subset using heuristic function $\phi$ with the special token and its information gain, to build new domain-specific tokenizer, and continues pretraining on the downstream task data. We explored the many positive effects of this method's customized tokenizer on domain-adaptive pretraining and verified this method can perform better than the ordinary method of just collecting data and fine-tuning. Based on our experiment, the continued pretraining process of IGOT with LLaMA-7B achieved 11.9\% token saving, 12.2\% training time saving, and 5.8\% maximum GPU VRAM usage saving, combined with the T5 model, we can even reach a 31.5\% of training time saving, making porting general generative AI to specific domains more effective than before. In domain-specific tasks, supervised $IGOT_\tau$ shows great performance on reducing both the convergence radius and convergence point during keep pretraining.
Dynamic In-context Learning with Conversational Models for Data Extraction and Materials Property Prediction
The advent of natural language processing and large language models (LLMs) has revolutionized the extraction of data from unstructured scholarly papers. However, ensuring data trustworthiness remains a significant challenge. In this paper, we introduce PropertyExtractor, an open-source tool that leverages advanced conversational LLMs like Google Gemini-Pro and OpenAI GPT-4, blends zero-shot with few-shot in-context learning, and employs engineered prompts for the dynamic refinement of structured information hierarchies, enabling autonomous, efficient, scalable, and accurate identification, extraction, and verification of material property data. Our tests on material data demonstrate precision and recall exceeding 93% with an error rate of approximately 10%, highlighting the effectiveness and versatility of the toolkit. We apply PropertyExtractor to generate a database of 2D material thicknesses, a critical parameter for device integration. The rapid evolution of the field has outpaced both experimental measurements and computational methods, creating a significant data gap. Our work addresses this gap and showcases the potential of PropertyExtractor as a reliable and efficient tool for the autonomous generation of diverse material property databases, advancing the field.
The AI Collaborator: Bridging Human-AI Interaction in Educational and Professional Settings
Samadi, Mohammad Amin, JaQuay, Spencer, Gu, Jing, Nixon, Nia
In the rapidly evolving landscape of artificial intelligence, significant advancements are being made, impacting a broad spectrum of fields ranging from Education [Becker et al.(2018)] to road transit [Banks and Stanton(2019)]. Looking ahead, these advancements are poised to significantly influence the dynamics of team environments. While research on teams only a few years ago highlighted the potential usefulness of AI integration in both research and practical settings, it also acknowledged the limitations of AI technologies in fully mimicking and comprehending the complex aspects of human-team interactions at the time [Seeber et al.(2020)]. However, with recent developments in generative AI and Large Language Models i.e., (OpenAI's GPT-4 [OpenAI(2023)], Google's Bard [Manyika and Hsiao(2023)] and Gemini [Team et al.(2023)]), we are approaching a level where AI-human teams can collaborate more effectively e.g., [Lakhnati et al.(2023)]. This progression prompts a critical question: How can we harness the evolving capabilities of AI to effectively enhance and integrate it into human-AI team dynamics, particularly in settings where traditional automation tools face limitations?
Rethinking Multi-User Semantic Communications with Deep Generative Models
Grassucci, Eleonora, Choi, Jinho, Park, Jihong, Gramaccioni, Riccardo F., Cicchetti, Giordano, Comminiello, Danilo
In recent years, novel communication strategies have emerged to face the challenges that the increased number of connected devices and the higher quality of transmitted information are posing. Among them, semantic communication obtained promising results especially when combined with state-of-the-art deep generative models, such as large language or diffusion models, able to regenerate content from extremely compressed semantic information. However, most of these approaches focus on single-user scenarios processing the received content at the receiver on top of conventional communication systems. In this paper, we propose to go beyond these methods by developing a novel generative semantic communication framework tailored for multi-user scenarios. This system assigns the channel to users knowing that the lost information can be filled in with a diffusion model at the receivers. Under this innovative perspective, OFDMA systems should not aim to transmit the largest part of information, but solely the bits necessary to the generative model to semantically regenerate the missing ones. The thorough experimental evaluation shows the capabilities of the novel diffusion model and the effectiveness of the proposed framework, leading towards a GenAI-based next generation of communications.
Prepare to Get Manipulated by Emotionally Expressive Chatbots
It's nothing new for computers to mimic human social etiquette, emotion, or humor. We just aren't used to them doing it very well. OpenAI's presentation of an all-new version of ChatGPT on Monday suggests that's about to change. It's built around an updated AI model called GPT-4o, which OpenAI says is better able to make sense of visual and auditory input, describing it as "multimodal." You can point your phone at something, like a broken coffee cup or differential equation, and ask ChatGPT to suggest what to do.