Goto

Collaborating Authors

 Industry


LAFA: Agentic LLM-Driven Federated Analytics over Decentralized Data Sources

arXiv.org Artificial Intelligence

Abstract--Large Language Models (LLMs) have shown great promise in automating data analytics tasks by interpreting natural language queries and generating multi-operation execution plans. However, existing LLM-agent-based analytics frameworks operate under the assumption of centralized data access, offering little to no privacy protection. In contrast, federated analytics (F A) enables privacy-preserving computation across distributed data sources, but lacks support for natural language input and requires structured, machine-readable queries. In this work, we present LAF A, the first system that integrates LLM-agent-based data analytics with F A. LAF A introduces a hierarchical multi-agent architecture that accepts natural language queries and transforms them into optimized, executable F A workflows. T o improve execution efficiency, an optimizer agent rewrites and merges multiple DAGs, eliminating redundant operations and minimizing computational and communicational overhead. Our experiments demonstrate that LAF A consistently outperforms baseline prompting strategies by achieving higher execution plan success rates and reducing resource-intensive F A operations by a substantial margin. This work establishes a practical foundation for privacy-preserving, LLM-driven analytics that supports natural language input in the F A setting. The rapid development of Large Language Models (LLMs) has offered unprecedented capabilities in natural language understanding, reasoning, and planning [1], significantly transforming the landscape of data analytics. LLMs can interpret complex analytical intents, generate structured code, and orchestrate multi-step tasks by interacting with external environments such as databases and computation sandboxes. These capabilities have led to the emergence of LLM-based agents that decompose high-level queries, plan analytical workflows, and execute or verify results through tool interactions.


CARE: Contrastive Alignment for ADL Recognition from Event-Triggered Sensor Streams

arXiv.org Artificial Intelligence

Abstract--The recognition of Activities of Daily Living (ADLs) from event-triggered ambient sensors is an essential task in Ambient Assisted Living, yet existing methods remain constrained by representation-level limitations. Sequence-based approaches preserve temporal order of sensor activations but are sensitive to noise and lack spatial awareness, while image-based approaches capture global patterns and implicit spatial correlations but compress fine-grained temporal dynamics and distort sensor layouts. Na ฤฑve fusion (e.g., feature concatenation) fail to enforce alignment between sequence-and image-based representation views, under-utilizing their complementary strengths. We propose C ontrastive A lignment for ADL R ecognition from E vent-Triggered Sensor Streams (CARE), an end-to-end framework that jointly optimizes representation learning via Sequence-Image Contrastive Alignment (SICA) and classification via cross-entropy, ensuring both cross-representation alignment and task-specific discriminability. CARE integrates (i) time-aware, noise-resilient sequence encoding with (ii) spatially-informed and frequency-sensitive image representations, and employs (iii) a joint contrastive-classification objective for end-to-end learning of aligned and discriminative embeddings. Evaluated on three CASAS datasets, CARE achieves state-of-the-art performance (89.8% on Milan, 88.9% on Cairo, and 73.3% on Kyoto7) and demonstrates robustness to sensor malfunctions and layout variability, highlighting its potential for reliable ADL recognition in smart homes. Global increases in life expectancy are leading to aging societies, with a rising number of older adults who require continuous support from healthcare providers and their family members [30]. However, given the critical shortage of healthcare personnel, it is essential to support older adults in maintaining independence for as long as possible. These functional abilities often decline with aging, and can be further deteriorated by aging-related chronic conditions [32]. Ambient Assisted Living (AAL) technologies have emerged to support ADL performance, encompassing systems for activity recognition, anomaly detection, and personalized prompting.


Prompt-MII: Meta-Learning Instruction Induction for LLMs

arXiv.org Artificial Intelligence

A popular method to adapt large language models (LLMs) to new tasks is in-context learning (ICL), which is effective but incurs high inference costs as context length grows. In this paper we propose a method to perform instruction induction, where we take training examples and reduce them to a compact but descriptive prompt that can achieve performance comparable to ICL over the full training set. Specifically, we propose PROMPT-MII, a reinforcement learning (RL) based framework to meta-learn an instruction induction model that can generate compact instructions on the fly for an arbitrary new dataset. We train on over 3,000 diverse classification datasets from the HuggingFace hub, and evaluate on 90 unseen tasks. PROMPT-MII improves downstream model quality by 4-9 F1 points (10-20% relative), matching ICL performance while requiring 3-13x fewer tokens.


A tutorial on discovering and quantifying the effect of latent causal sources of multimodal EHR data

arXiv.org Artificial Intelligence

We provide an accessible description of a peer-reviewed generalizable causal machine learning pipeline to (i) discover latent causal sources of large-scale electronic health records observations, and (ii) quantify the source causal effects on clinical outcomes. We illustrate how imperfect multimodal clinical data can be processed, decomposed into probabilistic independent latent sources, and used to train taskspecific causal models from which individual causal effects can be estimated. We summarize the findings of the two real-world applications of the approach to date as a demonstration of its versatility and utility for medical discovery at scale.


DARTS-GT: Differentiable Architecture Search for Graph Transformers with Quantifiable Instance-Specific Interpretability Analysis

arXiv.org Artificial Intelligence

Abstract--Graph Transformers (GTs) have emerged as powerful architectures for graph-structured data, yet remain constrained by rigid designs and lack quantifiable interpretability methods. Current state-of-the-art GTs commit to fixed GNN types across all layers, missing potential benefits of depth-specific component selection, while their increasingly complex architectures become opaque black boxes where performance gains cannot be distinguished between meaningful structural patterns and spurious correlations. We redesign the GT attention mechanism through asymmetry, decoupling structural encoding from feature representation. Queries derive directly from node features, while keys and values come from graph neural network (GNN) transformations, separating how the model learns features from how it encodes graph structure. Within this asymmetric framework, we use Differentiable ARchiT ecture Search (DARTS) to select optimal GNN operators at each layer, enabling depth-wise heterogeneity inside the transformer attention itself, hence the name DARTS-GT . T o understand these discovered architectures, we develop the first quantitative interpretability framework for GTs through causal ablation that identifies which heads and nodes actually drive predictions. Our metrics: Head-deviation, Specialization, and Focus, reveal the specific components responsible for each prediction while enabling broader model comparison. Experiments across eight benchmarks demonstrate that DARTS-GT achieves state-of-the-art performance on four datasets while remaining competitive on others, with discovered architectures revealing dataset-specific patterns ranging from highly specialized to balanced GNN distributions. Our inter-pretability analysis reveals that visual attention salience and causal importance do not necessarily correlate, indicating that widely used visualization approaches may miss the components that actually matter for predictions. Crucially, the heterogeneous architectures found by DARTS-GT consistently produced more interpretable models than baseline GTs, establishing that Graph Transformers do not need to choose between performance and interpretability. For graph-structured data, Graph Transformers (GTs) have become a dominant architectural choice, combining attention mechanisms with graph-awareness [1], [2]. Their success spans protein structure-to-function prediction [3], drug discovery [4], and materials design [5], where understanding complex structural patterns is crucial. Current state-of-the-art GTs incorporate graph structure through GNN-Transformer combinations [2], [6], specialized positional encodings [7], and attention augmentation with structural biases [8].


When In Doubt, Abstain: The Impact of Abstention on Strategic Classification

arXiv.org Artificial Intelligence

Algorithmic decision making is increasingly prevalent, but often vulnerable to strategic manipulation by agents seeking a favorable outcome. Prior research has shown that classifier abstention (allowing a classifier to decline making a decision due to insufficient confidence) can significantly increase classifier accuracy. This paper studies abstention within a strategic classification context, exploring how its introduction impacts strategic agents' responses and how principals should optimally leverage it. We model this interaction as a Stackelberg game where a principal, acting as the classifier, first announces its decision policy, and then strategic agents, acting as followers, manipulate their features to receive a desired outcome. Here, we focus on binary classifiers where agents manipulate observable features rather than their true features, and show that optimal abstention ensures that the principal's utility (or loss) is no worse than in a non-abstention setting, even in the presence of strategic agents. We also show that beyond improving accuracy, abstention can also serve as a deterrent to manipulation, making it costlier for agents, especially those less qualified, to manipulate to achieve a positive outcome when manipulation costs are significant enough to affect agent behavior. These results highlight abstention as a valuable tool for reducing the negative effects of strategic behavior in algorithmic decision making systems.


Generative AI and Firm Productivity: Field Experiments in Online Retail

arXiv.org Artificial Intelligence

We quantify the impact of Generative Artificial Intelligence (GenAI) on firm productivity through a series of large-scale randomized field experiments involving millions of users and products at a leading cross-border online retail platform. Over six months in 2023-2024, GenAI-based enhancements were integrated into seven consumer-facing business workflows. We find that GenAI adoption significantly increases sales, with treatment effects ranging from $0\%$ to $16.3\%$, depending on GenAI's marginal contribution relative to existing firm practices. Because inputs and prices were held constant across experimental arms, these gains map directly into total factor productivity improvements. Across the four GenAI applications with positive effects, the implied annual incremental value is approximately $\$ 5$ per consumer-an economically meaningful impact given the retailer's scale and the early stage of GenAI adoption. The primary mechanism operates through higher conversion rates, consistent with GenAI reducing frictions in the marketplace and improving consumer experience. We also document substantial heterogeneity: smaller and newer sellers, as well as less experienced consumers, exhibit disproportionately larger gains. Our findings provide novel, large-scale causal evidence on the productivity effects of GenAI in online retail, highlighting both its immediate value and broader potential.


SVTime: Small Time Series Forecasting Models Informed by "Physics" of Large Vision Model Forecasters

arXiv.org Artificial Intelligence

Time series AI is crucial for analyzing dynamic web content, driving a surge of pre-trained large models known for their strong knowledge encoding and transfer capabilities across diverse tasks. However, given their energy-intensive training, inference, and hardware demands, using large models as a one-fits-all solution raises serious concerns about carbon footprint and sustainability. For a specific task, a compact yet specialized, high-performing model may be more practical and affordable, especially for resource-constrained users such as small businesses. This motivates the question: Can we build cost-effective lightweight models with large-model-like performance on core tasks such as forecasting? This paper addresses this question by introducing SVTime, a novel Small model inspired by large Vision model (LVM) forecasters for long-term Time series forecasting (LTSF). Recently, LVMs have been shown as powerful tools for LTSF. We identify a set of key inductive biases of LVM forecasters -- analogous to the "physics" governing their behaviors in LTSF -- and design small models that encode these biases through meticulously crafted linear layers and constraint functions. Across 21 baselines spanning lightweight, complex, and pre-trained large models on 8 benchmark datasets, SVTime outperforms state-of-the-art (SOTA) lightweight models and rivals large models with 10^3 fewer parameters than LVMs, while enabling efficient training and inference in low-resource settings.


Artificially intelligent agents in the social and behavioral sciences: A history and outlook

arXiv.org Artificial Intelligence

We review the historical development and current trends of artificially intelligent agents (agentic AI) in the social and behavioral sciences: from the first programmable computers, and social simulations soon thereafter, to today's experiments with large language models. This overview emphasizes the role of AI in the scientific process and the changes brought about, both through technological advancements and the broader evolution of science from around 1950 to the present. Some of the specific points we cover include: the challenges of presenting the first social simulation studies to a world unaware of computers, the rise of social systems science, intelligent game theoretic agents, the age of big data and the epistemic upheaval in its wake, and the current enthusiasm around applications of generative AI, and many other topics. A pervasive theme is how deeply entwined we are with the technologies we use to understand ourselves.


Deciphering Invariant Feature Decoupling in Source-free Time Series Forecasting with Proxy Denoising

arXiv.org Artificial Intelligence

The proliferation of mobile devices generates a massive volume of time series across various domains, where effective time series forecasting enables a variety of real-world applications. This study focuses on a new problem of source-free domain adaptation for time series forecasting. It aims to adapt a pretrained model from sufficient source time series to the sparse target time series domain without access to the source data, embracing data protection regulations. To achieve this, we propose TimePD, the first source-free time series forecasting framework with proxy denoising, where large language models (LLMs) are employed to benefit from their generalization capabilities. Specifically, TimePD consists of three key components: (1) dual-branch invariant disentangled feature learning that enforces representation-and gradient-wise invariance by means of season-trend decomposition; (2) lightweight, parameter-free proxy denoising that dynamically calibrates systematic biases of LLMs; and (3) knowledge distillation that bidirectionally aligns the denoised prediction and the original target prediction. Extensive experiments on real-world datasets offer insight into the effectiveness of the proposed TimePD, outperforming SOT A baselines by 9.3% on average. The widespread deployment of Internet-of-Things (IoT) sensors has produced massive time series data across domains (Sun et al., 2025; Wang et al., 2024a), including traffic (Kieu et al., 2024; Cirstea et al., 2022), weather (Hettige et al., 2024), and energy (Wu et al., 2020). Accurate time series forecasting is crucial, enabling effective decision-making across diverse domains (Liu et al., 2025a; 2024a; Campos et al., 2023). We are seeing impressive advances in machine learning, especially in deep learning, that are successful in effective feature extraction and value creation (Hettige et al., 2024; Liu et al., 2025b).