Generative AI
Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey
Chen, Liang, Wang, Zekun, Ren, Shuhuai, Li, Lei, Zhao, Haozhe, Li, Yunshui, Cai, Zefan, Guo, Hongcheng, Zhang, Lei, Xiong, Yizhe, Zhang, Yichi, Wu, Ruoyu, Dong, Qingxiu, Zhang, Ge, Yang, Jian, Meng, Lingwei, Hu, Shujie, Chen, Yulong, Lin, Junyang, Bai, Shuai, Vlachos, Andreas, Tan, Xu, Zhang, Minjia, Xiao, Wen, Yee, Aaron, Liu, Tianyu, Chang, Baobao
Building on the foundations of language modeling in natural language processing, Next Token Prediction (NTP) has evolved into a versatile training objective for machine learning tasks across various modalities, achieving considerable success. As Large Language Models (LLMs) have advanced to unify understanding and generation tasks within the textual modality, recent research has shown that tasks from different modalities can also be effectively encapsulated within the NTP framework, transforming the multimodal information into tokens and predict the next one given the context. This survey introduces a comprehensive taxonomy that unifies both understanding and generation within multimodal learning through the lens of NTP. The proposed taxonomy covers five key aspects: Multimodal tokenization, MMNTP model architectures, unified task representation, datasets \& evaluation, and open challenges. This new taxonomy aims to aid researchers in their exploration of multimodal intelligence. An associated GitHub repository collecting the latest papers and repos is available at https://github.com/LMM101/Awesome-Multimodal-Next-Token-Prediction
The Synergy of Automated Pipelines with Prompt Engineering and Generative AI in Web Crawling
Web crawling is a critical technique for extracting online data, yet it poses challenges due to webpage diversity and anti-scraping mechanisms. This study investigates the integration of generative AI tools Claude AI (Sonnet 3.5) and ChatGPT4.0 with prompt engineering to automate web scraping. Using two prompts, PROMPT I (general inference, tested on Yahoo News) and PROMPT II (element-specific, tested on Coupons.com), we evaluate the code quality and performance of AI-generated scripts. Claude AI consistently outperformed ChatGPT-4.0 in script quality and adaptability, as confirmed by predefined evaluation metrics, including functionality, readability, modularity, and robustness. Performance data were collected through manual testing and structured scoring by three evaluators. Visualizations further illustrate Claude AI's superiority. Anti-scraping solutions, including undetected_chromedriver, Selenium, and fake_useragent, were incorporated to enhance performance. This paper demonstrates how generative AI combined with prompt engineering can simplify and improve web scraping workflows.
Low-Overhead Channel Estimation via 3D Extrapolation for TDD mmWave Massive MIMO Systems Under High-Mobility Scenarios
Zhou, Binggui, Yang, Xi, Ma, Shaodan, Gao, Feifei, Yang, Guanghua
In time division duplexing (TDD) millimeter wave (mmWave) massive multiple-input multiple-output (MIMO) systems, downlink channel state information (CSI) can be obtained from uplink channel estimation thanks to channel reciprocity. However, under high-mobility scenarios, frequent uplink channel estimation is needed due to channel aging. Additionally, large amounts of antennas and subcarriers result in high-dimensional CSI matrices, aggravating pilot training overhead. To address this, we propose a three-domain (3D) channel extrapolation framework across spatial, frequency, and temporal domains. First, considering the effectiveness of traditional knowledge-driven channel estimation methods and the marginal effects of pilots in the spatial and frequency domains, a knowledge-and-data driven spatial-frequency channel extrapolation network (KDD-SFCEN) is proposed for uplink channel estimation via joint spatial-frequency channel extrapolation to reduce spatial-frequency domain pilot overhead. Then, leveraging channel reciprocity and temporal dependencies, we propose a temporal uplink-downlink channel extrapolation network (TUDCEN) powered by generative artificial intelligence for slot-level channel extrapolation, aiming to reduce the tremendous temporal domain pilot overhead caused by high mobility. Numerical results demonstrate the superiority of the proposed framework in significantly reducing the pilot training overhead by 16 times and improving the system's spectral efficiency under high-mobility scenarios compared with state-of-the-art channel estimation/extrapolation methods.
FairDiffusion: Enhancing Equity in Latent Diffusion Models via Fair Bayesian Perturbation
Luo, Yan, Khan, Muhammad Osama, Wen, Congcong, Afzal, Muhammad Muneeb, Wuermeling, Titus Fidelis, Shi, Min, Tian, Yu, Fang, Yi, Wang, Mengyu
Recent progress in generative AI, especially diffusion models, has demonstrated significant utility in text-to-image synthesis. Particularly in healthcare, these models offer immense potential in generating synthetic datasets and training medical students. However, despite these strong performances, it remains uncertain if the image generation quality is consistent across different demographic subgroups. To address this critical concern, we present the first comprehensive study on the fairness of medical text-to-image diffusion models. Our extensive evaluations of the popular Stable Diffusion model reveal significant disparities across gender, race, and ethnicity. To mitigate these biases, we introduce FairDiffusion, an equity-aware latent diffusion model that enhances fairness in both image generation quality as well as the semantic correlation of clinical features. In addition, we also design and curate FairGenMed, the first dataset for studying the fairness of medical generative models. Complementing this effort, we further evaluate FairDiffusion on two widely-used external medical datasets: HAM10000 (dermatoscopic images) and CheXpert (chest X-rays) to demonstrate FairDiffusion's effectiveness in addressing fairness concerns across diverse medical imaging modalities. Together, FairDiffusion and FairGenMed significantly advance research in fair generative learning, promoting equitable benefits of generative AI in healthcare.
OpenAI lays out plan to shift to for-profit corporate structure
OpenAI has laid out a plan to revamp its corporate structure next year, saying it would create a public benefit corporation to manage its growing business and ease the restrictions imposed by its current non-profit parent. Rumors have swirled that OpenAI was in the process of shifting to a largely for-profit company, but this is the first time it has detailed the proposal publicly. Under the proposed structure, the public benefit corporation, which is a for-profit corporate entity, will run and control OpenAI's operations and business, while the non-profit will hire a leadership team and staff for charitable initiatives in sectors such as healthcare, education and science. This new structure will give the for-profit arm of OpenAI much more control. In a blogpost, the company said it is "a stronger non-profit supported by the for-profit's success".
OpenAI's for-profit plan includes a public benefit corporation
Following months of speculation, OpenAI has finally shared how it plans to become a for-profit company. In a blog post penned by its board of directors, OpenAI said Thursday it plans to transform its for-profit arm into a Public Benefit Corporation sometime in 2025. PBCs or B Corps are for-profit organizations that attempt to balance the interests of their stakeholders while making a positive impact on society. "As we enter 2025, we will have to become more than a lab and a startup -- we have to become an enduring company," OpenAI said, adding that many of its competitors are registered as PBCs, including Anthropic and even Elon Musk's own xAI. "[The move] would enable us to raise the necessary capital with conventional terms like others in this space."
OpenAI whistleblower's mother wants suicide death investigation reopened
If you or someone you know is having thoughts of suicide, please contact the Suicide & Crisis Lifeline at 988 or 1-800-273-TALK (8255). Balaji's death on November 26 was ruled a suicide, and Fox News Digital previously reported that the San Francisco Police Department found no evidence of foul play. But the 26-year-old's mother is urging police to reopen their investigation, saying it "doesn't look like a normal situation." Bereaved mother Poornima Ramarao told Business Insider that a private autopsy commissioned by Balaji's family and completed in early December produced concerning results. Now, they are working with an attorney to urge the department to conduct a "proper investigation."
Estimation of System Parameters Including Repeated Cross-Sectional Data through Emulator-Informed Deep Generative Model
Cho, Hyunwoo, Cho, Sung Woong, Jo, Hyeontae, Hwang, Hyung Ju
Differential equations (DEs) are crucial for modeling the evolution of natural or engineered systems. Traditionally, the parameters in DEs are adjusted to fit data from system observations. However, in fields such as politics, economics, and biology, available data are often independently collected at distinct time points from different subjects (i.e., repeated cross-sectional (RCS) data). Conventional optimization techniques struggle to accurately estimate DE parameters when RCS data exhibit various heterogeneities, leading to a significant loss of information. To address this issue, we propose a new estimation method called the emulator-informed deep-generative model (EIDGM), designed to handle RCS data. Specifically, EIDGM integrates a physics-informed neural network-based emulator that immediately generates DE solutions and a Wasserstein generative adversarial network-based parameter generator that can effectively mimic the RCS data. We evaluated EIDGM on exponential growth, logistic population models, and the Lorenz system, demonstrating its superior ability to accurately capture parameter distributions. Additionally, we applied EIDGM to an experimental dataset of Amyloid beta 40 and beta 42, successfully capturing diverse parameter distribution shapes. This shows that EIDGM can be applied to model a wide range of systems and extended to uncover the operating principles of systems based on limited data.
Global Search of Optimal Spacecraft Trajectories using Amortization and Deep Generative Models
Beeson, Ryne, Li, Anjian, Sinha, Amlan
The preliminary spacecraft trajectory design phase can be posed as a parameterized global search problem for optimal spacecraft trajectories. At each stage of the preliminary design, the mission objectives, requirements, and constraints may change, resulting in variations of the global search problem parameters. Parameters may also change to represent increased modeling fidelity. The aim at any stage of the preliminary design is to solve for a large set of high quality spacecraft trajectories with diverse, or similarly qualitatively different, features. High quality is naturally defined by the value of a solution's objective value relative to the best known. Examples of qualitatively different features may include trajectories that have a different number of revolutions around a central body, a different number or sequence of gravity assist flybys, solutions that avoid radiation belts or other hazards, or solutions that depart the original or target orbital planes. The benefit of having different qualitative solutions is that it allows mission designers to trade different priorities in their design and reflects the fact that not all relevant objectives and constraints can be incorporated into the optimal spacecraft trajectory problem so early or readily in the design phase (i.e., without prior knowledge of what is relevant and when designing at a quick cadence). In the simplest of cases, a mission designer's past experience may be sufficient to guide them in finding a high quality set of solutions.
ErgoChat: a Visual Query System for the Ergonomic Risk Assessment of Construction Workers
Fan, Chao, Mei, Qipei, Wang, Xiaonan, Li, Xinming
In the construction sector, workers often endure prolonged periods of high-intensity physical work and prolonged use of tools, resulting in injuries and illnesses primarily linked to postural ergonomic risks, a longstanding predominant health concern. To mitigate these risks, researchers have applied various technological methods to identify the ergonomic risks that construction workers face. However, traditional ergonomic risk assessment (ERA) techniques do not offer interactive feedback. The rapidly developing vision-language models (VLMs), capable of generating textual descriptions or answering questions about ergonomic risks based on image inputs, have not yet received widespread attention. This research introduces an interactive visual query system tailored to assess the postural ergonomic risks of construction workers. The system's capabilities include visual question answering (VQA), which responds to visual queries regarding workers' exposure to postural ergonomic risks, and image captioning (IC), which generates textual descriptions of these risks from images. Additionally, this study proposes a dataset designed for training and testing such methodologies. Systematic testing indicates that the VQA functionality delivers an accuracy of 96.5%. Moreover, evaluations using nine metrics for IC and assessments from human experts indicate that the proposed approach surpasses the performance of a method using the same architecture trained solely on generic datasets. This study sets a new direction for future developments in interactive ERA using generative artificial intelligence (AI) technologies. Keywords: Generative Artificial Intelligence; Vision-Language Model; Large language model; Ergonomic Risk Assessment; Construction Safety 1 Introduction Prompt and effective identification and mitigation of workplace hazards are essential for maintaining safety, health, and productivity within the work environment. In the construction industry, workers are often subject to conditions that require awkward body postures, repetitive motions, and intense physical effort, which can detrimentally impact their health [1]. Such conditions in construction tasks usually lead to the emergence of work-related musculoskeletal disorders (WMSDs). Statistics from the United States Bureau of Labor Statistics show that the construction industry's injuries and illnesses caused by WMSDs ranked fifth among all industries. Moreover, in the same year, WMSDs represented 30% of all occupational injuries and illnesses [1]. According to the Association of Workers' Compensation Boards of Canada, the manufacturing and construction sectors reported the second and third-highest rates of losttime injury claims in 2021, representing 13.6% and 10.4% of claims, respectively [2]. European Agency for Safety and Health at Work indicated that the construction and manufacturing sectors reported the highest sick leave rates due to WMSDs [3].