Goto

Collaborating Authors

 Statistical Learning


Music Recommendation on Spotify using Deep Learning

arXiv.org Artificial Intelligence

Hosting about 50 million songs and 4 billion playlists, there is an enormous amount of data generated at Spotify every single day - upwards of 600 gigabytes of data (harvard.edu). Since the algorithms that Spotify uses in recommendation systems is proprietary and confidential, code for big data analytics and recommendation can only be speculated. However, it is widely theorized that Spotify uses two main strategies to target users' playlists and personalized mixes that are infamous for their retention - exploration and exploitation (kaggle.com). This paper aims to appropriate the filtering using the approach of deep learning for maximum user likeability. The architecture achieves 98.57% and 80% training and validation accuracy respectively.


Classification with Partially Private Features

arXiv.org Artificial Intelligence

Privacy of data has become increasingly important in large-scale machine learning applications, where data points correspond to individuals that seek privacy. Classifiers are often trained over data of individuals with sensitive attributes about individuals: income, education, marital status, and so on for instance. A wellaccepted way to incorporate privacy into machine learning is the framework of differential privacy [9, 7, 10]. The key idea is to add noise either to individual data items or to the output of the classifier so that the distribution of classifiers produced is mathematically close when an arbitrary individual is added or removed from the dataset. This provides a quantifiable way in which sensitive data about any individual is information theoretically secure during the classification process. All this comes at a cost: Adding noise leads to a loss in accuracy of the classifier, and as we elaborate below, a large body of work has studied the privacy-accuracy trade-off both theoretically and empirically.


Deep Quantum Error Correction

arXiv.org Artificial Intelligence

Quantum error correction codes (QECC) are a key component for realizing the potential of quantum computing. QECC, as its classical counterpart (ECC), enables the reduction of error rates, by distributing quantum logical information across redundant physical qubits, such that errors can be detected and corrected. In this work, we efficiently train novel {\emph{end-to-end}} deep quantum error decoders. We resolve the quantum measurement collapse by augmenting syndrome decoding to predict an initial estimate of the system noise, which is then refined iteratively through a deep neural network. The logical error rates calculated over finite fields are directly optimized via a differentiable objective, enabling efficient decoding under the constraints imposed by the code. Finally, our architecture is extended to support faulty syndrome measurement, by efficient decoding of repeated syndrome sampling. The proposed method demonstrates the power of neural decoders for QECC by achieving state-of-the-art accuracy, outperforming {for small distance topological codes,} the existing {end-to-end }neural and classical decoders, which are often computationally prohibitive.


Almost Equivariance via Lie Algebra Convolutions

arXiv.org Machine Learning

Recently, the equivariance of models with respect to a group action has become an important topic of research in machine learning. Analysis of the built-in equivariance of existing neural network architectures, as well as the study of building models that explicitly "bake in" equivariance, have become significant research areas in their own right. However, imbuing an architecture with a specific group equivariance imposes a strong prior on the types of data transformations that the model expects to see. While strictly-equivariant models enforce symmetries, real-world data does not always conform to such strict equivariances. In such cases, the prior of strict equivariance can actually prove too strong and cause models to underperform. Therefore, in this work we study a closely related topic, that of almost equivariance. We provide a definition of almost equivariance and give a practical method for encoding almost equivariance in models by appealing to the Lie algebra of a Lie group. Specifically, we define Lie algebra convolutions and demonstrate that they offer several benefits over Lie group convolutions, including being well-defined for non-compact Lie groups having non-surjective exponential map. From there, we demonstrate connections between the notions of equivariance and isometry and those of almost equivariance and almost isometry. We prove two existence theorems, one showing the existence of almost isometries within bounded distance of isometries of a manifold, and another showing the converse for Hilbert spaces. We extend these theorems to prove the existence of almost equivariant manifold embeddings within bounded distance of fully equivariant embedding functions, subject to certain constraints on the group action and the function class. Finally, we demonstrate the validity of our approach by benchmarking against datasets in fully equivariant and almost equivariant settings.


Can Learning Be Explained By Local Optimality In Low-rank Matrix Recovery?

arXiv.org Artificial Intelligence

We explore the local landscape of low-rank matrix recovery, aiming to reconstruct a $d_1\times d_2$ matrix with rank $r$ from $m$ linear measurements, some potentially noisy. When the true rank is unknown, overestimation is common, yielding an over-parameterized model with rank $k\geq r$. Recent findings suggest that first-order methods with the robust $\ell_1$-loss can recover the true low-rank solution even when the rank is overestimated and measurements are noisy, implying that true solutions might emerge as local or global minima. Our paper challenges this notion, demonstrating that, under mild conditions, true solutions manifest as \textit{strict saddle points}. We study two categories of low-rank matrix recovery, matrix completion and matrix sensing, both with the robust $\ell_1$-loss. For matrix sensing, we uncover two critical transitions. With $m$ in the range of $\max\{d_1,d_2\}r\lesssim m\lesssim \max\{d_1,d_2\}k$, none of the true solutions are local or global minima, but some become strict saddle points. As $m$ surpasses $\max\{d_1,d_2\}k$, all true solutions become unequivocal global minima. In matrix completion, even with slight rank overestimation and mild noise, true solutions either emerge as non-critical or strict saddle points.


VAE-IF: Deep feature extraction with averaging for unsupervised artifact detection in routine acquired ICU time-series

arXiv.org Artificial Intelligence

Artifacts are a common problem in physiological time-series data collected from intensive care units (ICU) and other settings. They affect the quality and reliability of clinical research and patient care. Manual annotation of artifacts is costly and time-consuming, rendering it impractical. Automated methods are desired. Here, we propose a novel unsupervised approach to detect artifacts in clinical-standard minute-by-minute resolution ICU data without any prior labeling or signal-specific knowledge. Our approach combines a variational autoencoder (VAE) and an isolation forest (iForest) model to learn features and identify anomalies in different types of vital signs, such as blood pressure, heart rate, and intracranial pressure. We evaluate our approach on a real-world ICU dataset and compare it with supervised models based on long short-term memory (LSTM) and XGBoost. We show that our approach achieves comparable sensitivity and generalizes well to an external dataset. We also visualize the latent space learned by the VAE and demonstrate its ability to disentangle clean and noisy samples. Our approach offers a promising solution for cleaning ICU data in clinical research and practice without the need for any labels whatsoever.


Adaptive Parameter Selection for Kernel Ridge Regression

arXiv.org Artificial Intelligence

This paper focuses on parameter selection issues of kernel ridge regression (KRR). Due to special spectral properties of KRR, we find that delicate subdivision of the parameter interval shrinks the difference between two successive KRR estimates. Based on this observation, we develop an early-stopping type parameter selection strategy for KRR according to the so-called Lepskii-type principle. Theoretical verifications are presented in the framework of learning theory to show that KRR equipped with the proposed parameter selection strategy succeeds in achieving optimal learning rates and adapts to different norms, providing a new record of parameter selection for kernel methods.


TaBIIC: Taxonomy Building through Iterative and Interactive Clustering

arXiv.org Artificial Intelligence

Building taxonomies is often a significant part of building an ontology, and many attempts have been made to automate the creation of such taxonomies from relevant data. The idea in such approaches is either that relevant definitions of the intension of concepts can be extracted as patterns in the data (e.g. in formal concept analysis) or that their extension can be built from grouping data objects based on similarity (clustering). In both cases, the process leads to an automatically constructed structure, which can either be too coarse and lacking in definition, or too fined-grained and detailed, therefore requiring to be refined into the desired taxonomy. In this paper, we explore a method that takes inspiration from both approaches in an iterative and interactive process, so that refinement and definition of the concepts in the taxonomy occur at the time of identifying those concepts in the data. We show that this method is applicable on a variety of data sources and leads to taxonomies that can be more directly integrated into ontologies.


Detecting Toxic Flow

arXiv.org Artificial Intelligence

In foreign exchange (FX), as in other asset classes, broker-client relationships are ubiquitous. The broker streams bid and ask quotes to her clients and the clients decide when to trade on these quotes, so the broker bears the risk of adverse selection when trading with better informed clients. These risks are borne by both liquidity providers who stream quotes to individual parties and by market participants who provide liquidity in the books of electronic exchanges. However, in contrast to electronic order books in which trading is anonymous for all participants (e.g., in Nasdaq, LSE, Euronext), in broker-client relationships the broker knows which client executed the order. This privileged information can be used by the broker to classify flow, i.e., toxic or benign, and to devise strategies that mitigate adverse selection costs. In the literature, models generally classify traders as informed or uninformed; see e.g., Bagehot (1971), Copeland and Galai (1983), Grossman and Stiglitz (1980), Amihud and Mendelson (1980), Kyle (1989), Kyle (1985), and Glosten and Milgrom (1985). In equity markets, many studies focus on informed flow (i.e., asymmetry of information) across various traded stocks, see e.g., Easley et al. (1996) who study the probability of informed trading at the stock level, while our study focuses on We thank Andrew Stewart, Alistair Sturgiss, Fayçal Drissi, Patrick Chang, Álvaro Arroyo, Sergio Calvo Ordoñez, and participants at the Oxford Victoria Seminar for comments. ChatGPT suggested the name PULSE for our algorithm.


ICTSurF: Implicit Continuous-Time Survival Functions with Neural Networks

arXiv.org Artificial Intelligence

Survival analysis, also known as time-to-event analysis, aims at estimating the survival distributions of a specific event and time-of-interests. Typically, the estimation of survival probability involves modeling a relationship between covariates and time-to-event that is typically partially observed; e.g., it may not be possible to observe the event status of the same sample. This presents one of the key challenges in the field of survival analysis. The conventional approaches commonly employed in survival analysis include the Cox Proportional Hazards (CPH) model, as proposed by Cox [6]. Although the CPH model is widely used, it is burdened by a substantial assumption of a consistent proportional hazard throughout the entire lifespan and a predetermined relationship between covariates. Other conventional methods, such as Weibull or Log-Normal distribution, also model a relationship between time and covariates based on a strong parametric assumption. Recently, due to the success of DNN-based models, the majority of research in survival analysis has shifted towards models built on DNNs, demonstrating superior performance compared to traditional approaches. Recent studies have shown that the majority of the survival models are an extension of the conventional CPH model [28].