Goto

Collaborating Authors

 Generative AI


How Nvidia's CEO got me excited about our 'agentic AI' future

PCWorld

"The era we're in now is called the era of reasoning AI which is going to be the foundation layer of the next era of AI, or agentic AI." This was a statement from Nvidia's CEO Jensen Huang at HP's Amplify Conference in Nashville, last week. And in reflecting on it, it struck me as poignant for the following reason: Up until now I've viewed AI on a single timeline starting at the point where I first saw simple AI tools pop up, to now, where they are a lot more complex and capable than they used to be. In doing so, I've failed to see the true potential of AI -- the fact that we're about to enter a whole new age of AI that will make even those smart generative AI tools we have today pale in comparison to what we'll soon have on our PCs and other connected devices. It's not that I haven't been impressed with the advancements I've seen in AI tools so far.


Test-Time Reasoning Through Visual Human Preferences with VLMs and Soft Rewards

arXiv.org Artificial Intelligence

Can Visual Language Models (VLMs) effectively capture huma n visual preferences? This work addresses this question by training VLMs to think about preferences at test time, employing reinforcement learnin g methods inspired by DeepSeek R1 and OpenAI O1. Using datasets such as ImageRewar d and Human Preference Score v2 (HPSv2), our models achieve accurac ies of 64.9% on the ImageReward test set (trained on ImageReward official sp lit) and 65.4% on HPSv2 (trained on approximately 25% of its data). These resu lts match traditional encoder-based models while providing transparent r easoning and enhanced generalization. This approach allows to use not only rich VL M world knowledge, but also its potential to think, yielding interpretable out comes that help decision-making processes. By demonstrating that human visual prefe rences reasonable by current VLMs, we introduce efficient soft-reward strateg ies for image ranking, outperforming simplistic selection or scoring methods. Th is reasoning capability enables VLMs to rank arbitrary images--regardless of aspect ratio or complexity--thereby potentially amplifying the effectiveness of v isual Preference Optimization. By reducing the need for extensive markup while im proving reward generalization and explainability, our findings can be a str ong mile-stone that will enhance text-to-vision models even further.


Open Deep Search: Democratizing Search with Open-source Reasoning Agents

arXiv.org Artificial Intelligence

We introduce Open Deep Search (ODS) to close the increasing gap between the proprietary search AI solutions, such as Perplexity's Sonar Reasoning Pro and OpenAI's GPT-4o Search Preview, and their open-source counterparts. The main innovation introduced in ODS is to augment the reasoning capabilities of the latest open-source LLMs with reasoning agents that can judiciously use web search tools to answer queries. Concretely, ODS consists of two components that work with a base LLM chosen by the user: Open Search Tool and Open Reasoning Agent. Open Reasoning Agent interprets the given task and completes it by orchestrating a sequence of actions that includes calling tools, one of which is the Open Search Tool. Open Search Tool is a novel web search tool that outperforms proprietary counterparts. Together with powerful open-source reasoning LLMs, such as DeepSeek-R1, ODS nearly matches and sometimes surpasses the existing state-of-the-art baselines on two benchmarks: SimpleQA and FRAMES. For example, on the FRAMES evaluation benchmark, ODS improves the best existing baseline of the recently released GPT-4o Search Preview by 9.7% in accuracy. ODS is a general framework for seamlessly augmenting any LLMs -- for example, DeepSeek-R1 that achieves 82.4% on SimpleQA and 30.1% on FRAMES -- with search and reasoning capabilities to achieve state-of-the-art performance: 88.3% on SimpleQA and 75.3% on FRAMES.


Guarding against artificial intelligence--hallucinated citations: the case for full-text reference deposit

arXiv.org Artificial Intelligence

The tendency of generative artificial intelligence (AI) sys tems to "hallucinate" false information is well-known; AI-generated cit ations to nonexistent sources have made their way into the reference list s of peer-reviewed publications. Here, I propose a solution to this pr oblem, taking inspiration from the T ransparency and Openness Promotion ( TOP) data sharing guidelines, the clash of generative AI with the Amer ican judiciary, and the precedent set by submissions of prior art to the Unite d States Patent and T rademark Office. Journals should require authors to sub mit the full text of each cited source along with their manuscripts, ther eby preventing authors from citing any material whose full text they cannot produce. This solution requires limited additional work on the part of aut hors or editors while effectively immunizing journals against hallucinat ed references. Within the same month, commenters on Pub-Peer raised concerns regarding the article's reference list.


Membership Inference Attacks on Large-Scale Models: A Survey

arXiv.org Artificial Intelligence

The adoption of the Large Language Model (LLM) has accelerated dramatically since the ChatGPT from OpenAI went online in November 2022. Recent advances in Large Multimodal Models (LMMs), which process diverse data types and enable interaction through various channels, have expanded beyond the text-to-text limitations of early LLMs, attracting significant and concurrent attention from both researchers and industry. While LLMs and LMMs are starting to spread widely, concerns about their privacy risks are increasing as well. Membership Inference Attacks (MIAs), techniques used to determine whether a particular data point was part of a model's training set, serve as a key metric for assessing the privacy vulnerabilities of machine learning models. Hu et al. show that various machine learning algorithms are vulnerable to MIA. Despite extensive studies on MIAs in traditional models, there remains a lack of systematic surveys addressing their effectiveness and implications in modern large-scale models like LLMs and LMMs. In this paper, we systematically reviewed recent studies of MIA against LLMs and LMMs. We analyzed and categorized each attack based on their methodology and scenario and discussed the limitations in existing research. Additionally, we examine privacy concerns associated with the fine-tuning process. Finally, we provided some suggestions for future research in this direction.


PALATE: Peculiar Application of the Law of Total Expectation to Enhance the Evaluation of Deep Generative Models

arXiv.org Artificial Intelligence

Deep generative models (DGMs) have caused a paradigm shift in the field of machine learning, yielding noteworthy advancements in domains such as image synthesis, natural language processing, and other related areas. However, a comprehensive evaluation of these models that accounts for the trichotomy between fidelity, diversity, and novelty in generated samples remains a formidable challenge. A recently introduced solution that has emerged as a promising approach in this regard is the Feature Likelihood Divergence (FLD), a method that offers a theoretically motivated practical tool, yet also exhibits some computational challenges. In this paper, we propose PALATE, a novel enhancement to the evaluation of DGMs that addresses limitations of existing metrics. Our approach is based on a peculiar application of the law of total expectation to random variables representing accessible real data. When combined with the MMD baseline metric and DINOv2 feature extractor, PALATE offers a holistic evaluation framework that matches or surpasses state-of-the-art solutions while providing superior computational efficiency and scalability to large-scale datasets. Through a series of experiments, we demonstrate the effectiveness of the PALATE enhancement, contributing a computationally efficient, holistic evaluation approach that advances the field of DGMs assessment, especially in detecting sample memorization and evaluating generalization capabilities.


Generative AI in Knowledge Work: Design Implications for Data Navigation and Decision-Making

arXiv.org Artificial Intelligence

Our study of 20 knowledge workers revealed a common challenge: the difficulty of synthesizing unstructured information scattered across multiple platforms to make informed decisions. Drawing on their vision of an ideal knowledge synthesis tool, we developed Yodeai, an AI-enabled system, to explore both the opportunities and limitations of AI in knowledge work. Through a user study with 16 product managers, we identified three key requirements for Generative AI in knowledge work: adaptable user control, transparent collaboration mechanisms, and the ability to integrate background knowledge with external information. However, we also found significant limitations, including overreliance on AI, user isolation, and contextual factors outside the AI's reach. As AI tools become increasingly prevalent in professional settings, we propose design principles that emphasize adaptability to diverse workflows, accountability in personal and collaborative contexts, and context-aware interoperability to guide the development of human-centered AI systems for product managers and knowledge workers.


OpenAI's Sora Is Plagued by Sexist, Racist, and Ableist Biases

WIRED

Despite recent leaps forward in image quality, the biases found in videos generated by AI tools, like OpenAI's Sora, are as conspicuous as ever. A WIRED investigation, which included a review of hundreds of AI-generated videos, has found that Sora's model perpetuates sexist, racist, and ableist stereotypes in its results. In Sora's world, everyone is good-looking. Pilots, CEOs, and college professors are men, while flight attendants, receptionists, and childcare workers are women. Disabled people are wheelchair users, interracial relationships are tricky to generate, and fat people don't run.


HH4AI: A methodological Framework for AI Human Rights impact assessment under the EUAI ACT

arXiv.org Artificial Intelligence

This paper introduces the HH4AI Methodology, a structured approach to assessing the impact of AI systems on human rights, focusing on compliance with the EU AI Act and addressing technical, ethical, and regulatory challenges. The paper highlights AIs transformative nature, driven by autonomy, data, and goal-oriented design, and how the EU AI Act promotes transparency, accountability, and safety. A key challenge is defining and assessing "high-risk" AI systems across industries, complicated by the lack of universally accepted standards and AIs rapid evolution. To address these challenges, the paper explores the relevance of ISO/IEC and IEEE standards, focusing on risk management, data quality, bias mitigation, and governance. It proposes a Fundamental Rights Impact Assessment (FRIA) methodology, a gate-based framework designed to isolate and assess risks through phases including an AI system overview, a human rights checklist, an impact assessment, and a final output phase. A filtering mechanism tailors the assessment to the system's characteristics, targeting areas like accountability, AI literacy, data governance, and transparency. The paper illustrates the FRIA methodology through a fictional case study of an automated healthcare triage service. The structured approach enables systematic filtering, comprehensive risk assessment, and mitigation planning, effectively prioritizing critical risks and providing clear remediation strategies. This promotes better alignment with human rights principles and enhances regulatory compliance.


A Novel Hat-Shaped Device-Cloud Collaborative Inference Framework for Large Language Models

arXiv.org Artificial Intelligence

Abstract--Recent advancements in large language models (LLMs) have catalyzed a substantial surge in demand for LLM services. While traditional cloud-based LLM services satisfy high-accuracy requirements, they fall short in meeting critical demands for low delay and enhanced privacy . T o address these limitations, we propose HA T, a novel device-cloud collaborative inference framework that leverages the complementary strengths of U-shaped inference and speculative decoding. HA T partitions the LLM into three submodels, and the input and output submodels, stacked with a lightweight adapter network, are deployed as a small language model (SLM) on each end device. Meanwhile, the middle submodel, encompassing the majority of the LLM's decoder layers, is hosted in the cloud to perform speculative decoding with on-device SLMs. During inference, HA T exchanges hidden states (rather than raw tokens) of input or draft tokens between devices and the cloud, thereby incurring substantial communication delays. Besides, processing hidden states of long prompts will exacerbate computation delays in the cloud, further compromising inference efficiency . T o improve efficiency, we introduce a prompt chunking mechanism that segments long prompts into shorter chunks, enabling parallel transmission and processing. Furthermore, HA T is implemented to dynamically determine optimal chunk sizes for devices handling long prompts, thereby improving overall inference speed. Extensive experiments are conducted on a physical testbed comprising 30 NVIDIA Jetson devices and a server with 8 NVIDIA A6000 GPUs. Experimental results demonstrate that HA T achieves promising performance improvements, reducing TTFT by 41% to 54% and TBT by 41% to 77% compared to the baselines. Recent advancements in large language models (LLMs) have revolutionized the field of natural language processing, demonstrating unprecedented capabilities across various tasks and triggering exponential growth of LLM services [1], [2]. For instance, OpenAI's ChatGPT provides various services, e.g., chat-based interaction, and automated writing, to approximately 180 million users, and processes over 1.6 billion requests monthly [3]. The underlying architecture of LLM services mainly operates through an autore-gressive process, which involves a prefill phase followed by a decode phase. In prefill phase, the LLM processes all input prompt tokens simultaneously, leveraging parallel computation to generate the initial output token.