Goto

Collaborating Authors

 Deep Learning


Improving Patient Subtyping on Longitudinal Data using Representations from Mamba-based Architecture

arXiv.org Machine Learning

Effective sub-typing (also known as grouping or clustering) of patients using their electronic health record (EHR) data can greatly inform precision medicine efforts. However, subtyping temporal EHR datasets is known to be challenging due to inherent EHR issues, including complexity and irregularity. In this study, we propose a self-supervised Mamba-based model that learns effective EHR representations and enables enhanced patient subtyping. We evaluate the proposed model on public and private real-world EHR datasets to classify the data based on the available labels and subtype patients based on the representations learned from the model. Through an extensive set of experiments, we demonstrate that our model's design choices lead to better performance compared to competitive baseline models for prediction. Moreover, we evaluate several clustering techniques to demonstrate that our findings offer valuable insights into subtyping patients based on temporal records from EHR models\footnote{Our implementations are available at https://github.com/healthylaife/triplet_mamba.


Perspectives on Latent Factor Indeterminacy and its Implications for Data Representation

arXiv.org Machine Learning

The common factor analytic model is related to Helmholtz and Boltzmann machines, can be conceived as a linear autoencoder, or can be thought of as a single-hidden-layer generative neural network. We thus consider it a basal generative representation learner that can be used as a minimal model for studying the foundational characteristics of (deep) generative model architectures. We focus on the fundamental problem of indeterminacy in latent factor projections. This indeterminacy implies that, even when the intrinsic dimension of the latent vector is known, regularity conditions are met, and rotational indeterminacy is resolved, an inherent indefiniteness in the retrieval of causative latent sources remains: they will be uncertain, distributionally deviant, and non-unique. This can have major implications for data representation but remains an elusive issue, even to practitioners and theorists well-versed in the factor model. Moreover, this classic psychometric problem is intricately related to the modern issue of latent variable collapse in the variational autoencoder framework for deep generative modeling. Here, we assess this indeterminacy from various perspectives and show how these are mathematically and conceptually related and we discuss subsequent implications for the Psychometrics, Statistics, and Artificial Intelligence communities. We show that one has latent factor determinacy across all its facets when the feature-dimension grows to infinity. This feeds into an essentially distribution-free estimation approach in the sample case when the number of features grows very large. We conclude, as these are emergent properties at scale, that the factor model is suited for representation learning of very-high-dimensional data.


What LLMs explain is not what they believe: Evaluating explanation sufficiency under models' own input beliefs

arXiv.org Machine Learning

Large language models (LLMs) are increasingly deployed in high-stakes domains, where free-text explanations such as chain-of-thought and post-hoc rationales are used to justify model outputs. Yet it remains unclear whether these explanations are sufficient, i.e., if they contain enough information to explain the model's output-generating process. We generalize classical sufficiency from feature attributions to arbitrary explanations and prove that explanation sufficiency can change depending on the input distribution, which must be explicitly defined for LLM explanations. We propose using the LLM itself to generate alternative inputs conditioned on an explanation, capturing its beliefs about possible inputs. We formalize self-consistent sufficiency as a goal for free-text explanations and introduce an information-theoretic metric, SCSuff, that enables evaluation of free-text explanations without relying on predefined biases or shortcuts. Our experiments show that SCSuff agrees with targeted perturbation tests where applicable and demonstrate that explanation sufficiency can vary with the input distribution. We find LLM explanations are generally insufficient and weakly correlated with model size, accuracy, or output entropy. Analysis of final-token hidden states shows that top and bottom SCSuff scores can be predicted from internal representations, suggesting that SCSuff can guide detection and improvement of sufficient LLM explanations. The code for this paper is available at https://github.com/rajesh-lab/self-consistent-sufficiency .


Weighted universal approximation of differentiable maps on infinite-dimensional manifolds

arXiv.org Machine Learning

We generalize the universal approximation theorem for functional input neural networks (FNN) to differentiable maps by including the approximation of the derivatives. A FNN maps the input from a possibly infinite-dimensional weighted manifold to the real-valued hidden layer, on which a non-linear scalar activation function is applied, and then returns the output into a Banach space via some linear readouts. By proving a weighted Nachbin theorem, we establish a universal approximation theorem for differentiable maps, which goes beyond the usual formulation on compact sets and also includes the approximation of the derivatives. This leads us to approximation results for non-anticipative functionals including the horizontal and vertical derivatives. As a further application, we show that linear functions of the signature are able to approximate path space functionals including their directional derivatives.


Sample Complexity of Scientific Discovery: PAC Learnability of Compositional Function Trees

arXiv.org Machine Learning

Scientific discovery via symbolic regression is often viewed as statistically and computationally intractable because the hypothesis space of expressions grows combinatorially with depth. This paper revisits the statistical side through the lens of PAC learning, focusing on compositional function trees built from a finite vocabulary of smooth operators (e.g., $\{+,\times,\sin,\exp\}$ and affine maps). We prove that the relevant generalization quantity, Rademacher complexity, hence the excess risk, does not necessarily blow up exponentially with the number of distinct symbolic structures, but is controlled by (i) the depth $d$ and (ii) the Lipschitz constants of the base operators along the composed computation graph. Concretely, under mild Lipschitz conditions on operators and bounded affine leaves, a finite-union bound over a vocabulary of size $K=|\mathcal{H}_{\mathrm{base}}|$ together with Maurer-type vector contraction yields $\mathfrak{R}_n(\mathcal{H}_{\mathrm{comp}}^{d}) \leq (Kb\sqrt{2}L)^{d-1}\mathfrak{R}_n(\mathcal{H}_{\mathrm{comp}}^{1})$ with arity bound $b$; corresponding high-probability risk bounds scale as $\mathcal{O}(L^{d}/\sqrt{n})$ when $K,b=O(1)$ and $\mathfrak{R}_n(\mathcal{H}_{\mathrm{comp}}^{1})=O(n^{-1/2})$. We complement the theory with a modular codebase that trains differentiable operator trees (not MLPs) on synthetic "physics-like" targets of controlled depth and shows that the empirical generalization gap correlates positively with the predicted complexity term $(\widehat{L}^{d})/\sqrt{n}$.


Convergence of Continual Learning in Homogeneous Deep Networks

arXiv.org Machine Learning

We characterize weakly regularized continual classification in homogeneous models as sequential projections onto task margin sets. This result generalizes prior analyses restricted to either stationary (single-task) deep models or continual linear models. We show that global convergence generally fails, even for simple models linear in data but nonlinear in parameters. Nevertheless, by leveraging results from nonconvex projection theory, we identify regularity properties of homogeneous deep networks that guarantee local linear convergence under random and cyclic task sequences. Finally, we extend our analysis to continual regression, unifying the framework for homogeneous models.


Are Humanoid Robots Ready to Be Deployed?

The New Yorker

Are Humanoid Robots Ready to Be Deployed? Neo and a dozen other robots with human forms are scheduled to hit the market. "The same robot that can land a backflip might not be able to walk up a flight of stairs," a researcher said. On a recent sunny day in Silicon Valley, I visited the industrial headquarters of 1X Technologies. Security was tight, so I had to put a sticker over my cellphone's camera and talk my way out of signing an N.D.A. before I was brought into an enormous space to meet Neo, the company's home robot. Neo stands five feet six and has no facial features except for two black cameras in place of eyes. The robot is a humanoid--its design is inspired by the human form--and its proportions are a blend of those of the median American male and those of the median American female. But Neo has no skin. Instead, it wears a beige nylon turtleneck bodysuit, gloves, and padded shoes over a see-through carapace. Under that is a skeleton made up of more than a hundred whizzing motors and cordlike artificial tendons that control Neo's limbs. Neo's cozy, minimalist aesthetic allows it to blend into the background. If it served me an espresso at a café, I'm not certain I would look up from my phone. The robot weighs just sixty-six pounds, and I was able to pick it up in a bridal carry. It communicates through a speaker in its chest, using several different voices; the default one is in a calm but authoritative masculine register, an A.I.-modulated mixture of several voice actors. Neo can talk, listen, and respond to commands.


My Boyfriend Said His Family Friend Is Like a "Sister" to Him. Uh, That's Not What His Computer History Says.

Slate

How to Do It My Boyfriend Said His Family Friend Is Like a "Sister" to Him. Uh, That's Not What His Computer History Says. Sign up for the Slatest to get the most insightful analysis, criticism, and advice out there, delivered to your inbox daily. I've been with my boyfriend for a year and things have been smooth sailing so far, with very little contention between us two. But about a week ago, he left his laptop open on the bed and I decided to take a look at his ChatGPT history.


Don't pay 20/month for ChatGPT--1 year of ChatOn gives you GPT, Gemini & Claude for just 30

PCWorld

When you purchase through links in our articles, we may earn a small commission. Don't pay $20/month for ChatGPT--1 year of ChatOn gives you GPT, Gemini & Claude for just $30 Try the biggest AI models in one place for a year. Get ChatOn Premium for $29.99 (MSRP $39.99) through June 28 and access GPT, Claude, Gemini, Sonar, and more from a single app. Most people don't need another AI subscription. They need one app that does the job .


Apple executive in charge of Vision Pro is reportedly leaving for OpenAI

Engadget

Paul Meade will start OpenAI's hardware division, 'Bloomberg' says. Paul Meade, an Apple VP who heads the Vision Products Group, is reportedly leaving the company next week for OpenAI. According to Bloomberg, the top executive in charge of the Vision Pro headset and Apple's smart glasses projects will be starting up the AI company's hardware unit. OpenAI has been developing AI-powered devices with Jony Ive's startup since 2025. While Ive's io merged with OpenAI in a $6.5 billion deal, it remains independent.