Goto

Collaborating Authors

 Statistical Learning


LIGHTCODE: Light Analytical and Neural Codes for Channels with Feedback

arXiv.org Artificial Intelligence

The design of reliable and efficient codes for channels with feedback remains a longstanding challenge in communication theory. While significant improvements have been achieved by leveraging deep learning techniques, neural codes often suffer from high computational costs, a lack of interpretability, and limited practicality in resource-constrained settings. We focus on designing low-complexity coding schemes that are interpretable and more suitable for communication systems. We advance both analytical and neural codes. First, we demonstrate that POWERBLAST, an analytical coding scheme inspired by Schalkwijk-Kailath (SK) and Gallager-Nakiboglu (GN) schemes, achieves notable reliability improvements over both SK and GN schemes, outperforming neural codes in high signal-to-noise ratio (SNR) regions. Next, to enhance reliability in low-SNR regions, we propose LIGHTCODE, a lightweight neural code that achieves state-of-the-art reliability while using a fraction of memory and compute compared to existing deep-learning-based codes. Finally, we systematically analyze the learned codes, establishing connections between LIGHTCODE and POWERBLAST, identifying components crucial for performance, and providing interpretation aided by linear regression analysis.


Quantum Machine Learning with HQC Architectures using non-Classically Simulable Feature Maps

arXiv.org Artificial Intelligence

Hybrid Quantum-Classical (HQC) Architectures are used in near-term NISQ Quantum Computers for solving Quantum Machine Learning problems. The quantum advantage comes into picture due to the exponential speedup offered over classical computing. One of the major challenges in implementing such algorithms is the choice of quantum embeddings and the use of a functionally correct quantum variational circuit. In this paper, we present an application of QSVM (Quantum Support Vector Machines) to predict if a person will require mental health treatment in the tech world in the future using the dataset from OSMI Mental Health Tech Surveys. We achieve this with non-classically simulable feature maps and prove that NISQ HQC Architectures for Quantum Machine Learning can be used alternatively to create good performance models in near-term real-world applications.


ALICE: Combining Feature Selection and Inter-Rater Agreeability for Machine Learning Insights

arXiv.org Machine Learning

The use of Machine Learning models for decision-making has become the new norm not only in tech but any business field imaginable, covering any possible task at hand be it search engine recommendations, customer churn prediction, credit risk scoring, energy load forecasting, or the deployment of personalized AI assistants. This comes at a time when developing ML models has become increasingly easier with the rise of open-source, free and user-friendly Python libraries such as Keras, scikit-learn, PyTorch and as generative AI-based conversational chatbots such as ChatGPT, Gemini and Claude that can provide coding assistance -- if not ready-made code for modeling -- are evolving rapidly. Such developments yet again beg the question of interpretability in machine learning, which has been formulated in various ways in literature and been offered multiple proposed solutions such as exploring causality (see Section 2.1), explainability (see Section 2.2) or abandoning black box ML models altogether. But to make a philosophical argument, it is hard to see the benefits of highly model or domain-specific, post-hoc, or complex solutions to obtain insights into the inner-doings of machine learning models when the modeling task itself is growing ever more accessible to laypeople. Common thought on categorizing ML models in this regard would argue that parametric models descending from the fields of statistics and econometrics such as Linear or Logistic Regression are by nature more interpretable than their data-driven and non-parametric counterparts such as tree-based models or neural networks.


Concentration properties of fractional posterior in 1-bit matrix completion

arXiv.org Machine Learning

The problem of estimating a matrix based on a set of its observed entries is commonly referred to as the matrix completion problem. In this work, we specifically address the scenario of binary observations, often termed as 1-bit matrix completion. While numerous studies have explored Bayesian and frequentist methods for real-value matrix completion, there has been a lack of theoretical exploration regarding Bayesian approaches in 1-bit matrix completion. We tackle this gap by considering a general, non-uniform sampling scheme and providing theoretical assurances on the efficacy of the fractional posterior. Our contributions include obtaining concentration results for the fractional posterior and demonstrating its effectiveness in recovering the underlying parameter matrix. We accomplish this using two distinct types of prior distributions: low-rank factorization priors and a spectral scaled Student prior, with the latter requiring fewer assumptions. Importantly, our results exhibit an adaptive nature by not mandating prior knowledge of the rank of the parameter matrix. Our findings are comparable to those found in the frequentist literature, yet demand fewer restrictive assumptions.


A Large Scale Survey of Motivation in Software Development and Analysis of its Validity

arXiv.org Artificial Intelligence

Context: Motivation is known to improve performance. In software development in particular, there has been considerable interest in the motivation of contributors to open source. Objective: We identify 11 motivators from the literature (enjoying programming, ownership of code, learning, self use, etc.), and evaluate their relative effect on motivation. Since motivation is an internal subjective feeling, we also analyze the validity of the answers. Method: We conducted a survey with 66 questions on motivation which was completed by 521 developers. Most of the questions used an 11 point scale. We evaluated the validity of the answers validity by comparing related questions, comparing to actual behavior on GitHub, and comparison with the same developer in a follow up survey. Results: Validity problems include moderate correlations between answers to related questions, as well as self promotion and mistakes in the answers. Despite these problems, predictive analysis, investigating how diverse motivators influence the probability of high motivation, provided valuable insights. The correlations between the different motivators are low, implying their independence. High values in all 11 motivators predict increased probability of high motivation. In addition, improvement analysis shows that an increase in most motivators predicts an increase in general motivation.


Federated Optimization with Doubly Regularized Drift Correction

arXiv.org Artificial Intelligence

Federated learning is a distributed optimization paradigm that allows training machine learning models across decentralized devices while keeping the data localized. The standard method, FedAvg, suffers from client drift which can hamper performance and increase communication costs over centralized methods. Previous works proposed various strategies to mitigate drift, yet none have shown uniformly improved communication-computation trade-offs over vanilla gradient descent. In this work, we revisit DANE, an established method in distributed optimization. We show that (i) DANE can achieve the desired communication reduction under Hessian similarity constraints. Furthermore, (ii) we present an extension, DANE+, which supports arbitrary inexact local solvers and has more freedom to choose how to aggregate the local updates. We propose (iii) a novel method, FedRed, which has improved local computational complexity and retains the same communication complexity compared to DANE/DANE+. This is achieved by using doubly regularized drift correction.


Under pressure: learning-based analog gauge reading in the wild

arXiv.org Artificial Intelligence

We propose an interpretable framework for reading analog gauges that is deployable on real world robotic systems. Our framework splits the reading task into distinct steps, such that we can detect potential failures at each step. Our system needs no prior knowledge of the type of gauge or the range of the scale and is able to extract the units used. We show that our gauge reading algorithm is able to extract readings with a relative reading error of less than 2%.


Constrained C-Test Generation via Mixed-Integer Programming

arXiv.org Artificial Intelligence

This work proposes a novel method to generate C-Tests; a deviated form of cloze tests (a gap filling exercise) where only the last part of a word is turned into a gap. In contrast to previous works that only consider varying the gap size or gap placement to achieve locally optimal solutions, we propose a mixed-integer programming (MIP) approach. This allows us to consider gap size and placement simultaneously, achieving globally optimal solutions, and to directly integrate state-of-the-art models for gap difficulty prediction into the optimization problem. A user study with 40 participants across four C-Test generation strategies (including GPT-4) shows that our approach (MIP) significantly outperforms two of the baseline strategies (based on gap placement and GPT-4); and performs on-par with the third (based on gap size). Our analysis shows that GPT-4 still struggles to fulfill explicit constraints during generation and that MIP produces C-Tests that correlate best with the perceived difficulty. We publish our code, model, and collected data consisting of 32 English C-Tests with 20 gaps each (totaling 3,200 individual gap responses) under an open source license.


Relational Prompt-based Pre-trained Language Models for Social Event Detection

arXiv.org Artificial Intelligence

Social Event Detection (SED) aims to identify significant events from social streams, and has a wide application ranging from public opinion analysis to risk management. In recent years, Graph Neural Network (GNN) based solutions have achieved state-of-the-art performance. However, GNN-based methods often struggle with noisy and missing edges between messages, affecting the quality of learned message embedding. Moreover, these methods statically initialize node embedding before training, which, in turn, limits the ability to learn from message texts and relations simultaneously. In this paper, we approach social event detection from a new perspective based on Pre-trained Language Models (PLMs), and present RPLM_SED (Relational prompt-based Pre-trained Language Models for Social Event Detection). We first propose a new pairwise message modeling strategy to construct social messages into message pairs with multi-relational sequences. Secondly, a new multi-relational prompt-based pairwise message learning mechanism is proposed to learn more comprehensive message representation from message pairs with multi-relational prompts using PLMs. Thirdly, we design a new clustering constraint to optimize the encoding process by enhancing intra-cluster compactness and inter-cluster dispersion, making the message representation more distinguishable. We evaluate the RPLM_SED on three real-world datasets, demonstrating that the RPLM_SED model achieves state-of-the-art performance in offline, online, low-resource, and long-tail distribution scenarios for social event detection tasks.


Enhancing Fairness and Performance in Machine Learning Models: A Multi-Task Learning Approach with Monte-Carlo Dropout and Pareto Optimality

arXiv.org Artificial Intelligence

The term bias was first introduced in the machine learning domain by Tom Mitchell in his 1980 paper titled "The need for biases in learning generalizations" Mitchell [1980]. The concept of bias refers to giving importance to particular features to improve generalization. This general idea of bias in machine learning is positive and necessary for models to perform, eliminating the risk of hyper-focusing on specific samples over others. On the contrary, bias can also be negative in machine learning. Negative bias can be defined as an inaccurate assumption made by a machine learning algorithm that is systematically or historically prejudiced against certain groups of people Zanna et al. [2022]. Decisions made by these biased algorithms could cause adverse effects on particular social groups, for example, those defined by sex, race, age, marital status, handicaps, etc., when used to make autonomous decisions in life-changing cases such as health, hiring, education, criminal sentencing, etc. Negative bias can be introduced into the machine pipeline in two main ways, through the data or the algorithm itself Blanzeisky and Cunningham [2021]. Bias due to data, also known as a negative legacy Cunningham and Delany [2021], Kamishima et al. [2012], can be caused by an imbalance in the representation of different population categories