Goto

Collaborating Authors

 Large Language Model


Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs

arXiv.org Artificial Intelligence

Large Language Models (LLMs) are increasingly used as chatbots, yet their ability to personalize responses to user preferences remains limited. We introduce PrefEval, a benchmark for evaluating LLMs' ability to infer, memorize and adhere to user preferences in a long-context conversational setting. PrefEval comprises 3,000 manually curated user preference and query pairs spanning 20 topics. PrefEval contains user personalization or preference information in both explicit and implicit forms, and evaluates LLM performance using a generation and a classification task. With PrefEval, we evaluated the aforementioned preference following capabilities of 10 open-source and proprietary LLMs in multi-session conversations with varying context lengths up to 100k tokens. We benchmark with various prompting, iterative feedback, and retrieval-augmented generation methods. Our benchmarking effort reveals that state-of-the-art LLMs face significant challenges in proactively following users' preferences during conversations. In particular, in zero-shot settings, preference following accuracy falls below 10% at merely 10 turns (~3k tokens) across most evaluated models. Even with advanced prompting and retrieval methods, preference following still deteriorates in long-context conversations. Furthermore, we show that fine-tuning on PrefEval significantly improves performance. We believe PrefEval serves as a valuable resource for measuring, understanding, and enhancing LLMs' preference following abilities, paving the way for personalized conversational agents. Our code and dataset are available at https://prefeval.github.io/.


Fine-Tuning Foundation Models with Federated Learning for Privacy Preserving Medical Time Series Forecasting

arXiv.org Artificial Intelligence

Federated Learning (FL) provides a decentralized machine learning approach, where multiple devices or servers collaboratively train a model without sharing their raw data, thus enabling data privacy. This approach has gained significant interest in academia and industry due to its privacy-preserving properties, which are particularly valuable in the medical domain where data availability is often protected under strict regulations. A relatively unexplored area is the use of FL to fine-tune Foundation Models (FMs) for time series forecasting, potentially enhancing model efficacy by overcoming data limitation while maintaining privacy. In this paper, we fine-tuned time series FMs with Electrocardiogram (ECG) and Impedance Cardiography (ICG) data using different FL techniques. We then examined various scenarios and discussed the challenges FL faces under different data heterogeneity configurations. Our empirical results demonstrated that while FL can be effective for fine-tuning FMs on time series forecasting tasks, its benefits depend on the data distribution across clients. We highlighted the trade-offs in applying FL to FM fine-tuning.


Task Generalization With AutoRegressive Compositional Structure: Can Learning From $\d$ Tasks Generalize to $\d^{T}$ Tasks?

arXiv.org Machine Learning

Large language models (LLMs) exhibit remarkable task generalization, solving tasks they were never explicitly trained on with only a few demonstrations. This raises a fundamental question: When can learning from a small set of tasks generalize to a large task family? In this paper, we investigate task generalization through the lens of AutoRegressive Compositional (ARC) structure, where each task is a composition of $T$ operations, and each operation is among a finite family of $\d$ subtasks. This yields a total class of size~\( \d^\TT \). We first show that generalization to all \( \d^\TT \) tasks is theoretically achievable by training on only \( \tilde{O}(\d) \) tasks. Empirically, we demonstrate that Transformers achieve such exponential task generalization on sparse parity functions via in-context learning (ICL) and Chain-of-Thought (CoT) reasoning. We further demonstrate this generalization in arithmetic and language translation, extending beyond parity functions.


Theoretical Benefit and Limitation of Diffusion Language Model

arXiv.org Machine Learning

Diffusion language models have emerged as a promising approach for text generation. One would naturally expect this method to be an efficient replacement for autoregressive models since multiple tokens can be sampled in parallel during each diffusion step. However, its efficiency-accuracy trade-off is not yet well understood. In this paper, we present a rigorous theoretical analysis of a widely used type of diffusion language model, the Masked Diffusion Model (MDM), and find that its effectiveness heavily depends on the target evaluation metric. Under mild conditions, we prove that when using perplexity as the metric, MDMs can achieve near-optimal perplexity in sampling steps regardless of sequence length, demonstrating that efficiency can be achieved without sacrificing performance. However, when using the sequence error rate--which is important for understanding the "correctness" of a sequence, such as a reasoning chain--we show that the required sampling steps must scale linearly with sequence length to obtain "correct" sequences, thereby eliminating MDM's efficiency advantage over autoregressive models. Our analysis establishes the first theoretical foundation for understanding the benefits and limitations of MDMs. All theoretical findings are supported by empirical studies.


In-Context Learning of Linear Dynamical Systems with Transformers: Error Bounds and Depth-Separation

arXiv.org Machine Learning

This paper investigates approximation-theoretic aspects of the in-context learning capability of the transformers in representing a family of noisy linear dynamical systems. Our first theoretical result establishes an upper bound on the approximation error of multi-layer transformers with respect to an $L^2$-testing loss uniformly defined across tasks. This result demonstrates that transformers with logarithmic depth can achieve error bounds comparable with those of the least-squares estimator. In contrast, our second result establishes a non-diminishing lower bound on the approximation error for a class of single-layer linear transformers, which suggests a depth-separation phenomenon for transformers in the in-context learning of dynamical systems. Moreover, this second result uncovers a critical distinction in the approximation power of single-layer linear transformers when learning from IID versus non-IID data.


Can AI solve your romantic issues? Study finds couples therapy can be done by ChatGPT

Daily Mail - Science & tech

Romantic issues might be dealt with using ChatGPT in the future. People can rarely tell the difference between couples therapy from a professional or advice supplied by AI, a study has found. In fact, responses to relationship problems written by ChatGPT were generally rated more highly for including key psychotherapy principles. Researchers asked 830 people, almost a fifth of whom had previously had couples therapy, to look at responses to relationship issues provided by either a therapist or AI. They had to identify whether they thought the answer came from the human expert or the AI tool in each case.


'DeepSeek moved me to tears': How young Chinese find therapy in AI

BBC News

Before she goes to bed each night, Holly Wang logs on to DeepSeek for "therapy sessions". Ever since January, when the breakout Chinese AI app launched, the 28-year-old has brought her dilemmas and sorrows, including the recent death of her grandmother, to the chatbot. Its responses have resonated so deeply they have at times brought her to tears. "DeepSeek has been such an amazing counsellor. It has helped me look at things from different perspectives and does a better job than the paid counselling services I have tried," says Holly, who asked for her real name to be withheld to protect her privacy. From writing reports and Excel formulas to planning trips, workouts and learning new skills, AI apps have found their way into many people's lives across the world.


OpenAI will offer free ChatGPT users unlimited access to GPT-5

Engadget

OpenAI's upcoming GPT-5 release will integrate its o3 reasoning model and be available to free users, CEO Sam Altman revealed in a roadmap he shared on X. He said the company is also working to simplify how users interact with ChatGPT. "We want AI to'just work' for you; we realize how complicated our model and product offerings have gotten," Altman wrote. "We hate the model picker as much as you do and want to return to magic unified intelligence." In its current iteration, forcing ChatGPT to use a specific model, such as o3-mini, involves either tapping the "Reason" button in the prompt bar or one of the options present in the model picker, which appears after the chatbot answers a question.


The Dirty Truth Behind Musk and Altman's Mud Fight

Slate

Sign up for the Slatest to get the most insightful analysis, criticism, and advice out there, delivered to your inbox daily. On Monday, Elon Musk and a group of investors made an unsolicited offer to buy ChatGPT parent company OpenAI for 97.4 billion. This ticked off OpenAI CEO Sam Altman not only because the company isn't for sale, but because the offer is insultingly low. OpenAI is reportedly in talks to raise new money in a funding round led by SoftBank at a 300 billion valuation, which would make it the most valuable privately held company in the world. Altman's return fire was equally petty, with him refusing via tweet before offering to buy Musk's social media company X for a decimal-sliding 9.74 billion.


Watch out, Nvidia. OpenAI's proprietary AI chip is coming along

PCWorld

According to a new report from Reuters (spotted by Thurrott), OpenAI could finalize the design of its first 3nm AI chip in the coming months, with the goal of starting mass production at TSMC in 2026. The chip is being developed by a team of 40 OpenAI employees in collaboration with Broadcom. The project is being led by Richard Ho, OpenAI's new head of hardware, who previously worked on solutions for Google's infrastructure and cloud services. According to Reuters, OpenAI's chip will be able to both train and run AI models, but initially it'll be used mainly for inference (running AI models) and to a limited extent within the company's infrastructure. Demand for Nvidia's AI chips remains extremely high right now, with companies like OpenAI, Microsoft, Meta, and Google investing billions in AI data centers.