Overview
Privacy in the Age of AI: A Taxonomy of Data Risks
Billiris, Grace, Gill, Asif, Bandara, Madhushi
Artificial Intelligence (AI) systems introduce unprecedented privacy challenges as they process increasingly sensitive data. Traditional privacy frameworks prove inadequate for AI technologies due to unique characteristics such as autonomous learning and black-box decision-making. This paper presents a taxonomy classifying AI privacy risks, synthesised from 45 studies identified through systematic review. We identify 19 key risks grouped under four categories: Dataset-Level, Model-Level, Infrastructure-Level, and Insider Threat Risks. Findings reveal a balanced distribution across these dimensions, with human error (9.45%) emerging as the most significant factor. This taxonomy challenges conventional security approaches that typically prioritise technical controls over human factors, highlighting gaps in holistic understanding. By bridging technical and behavioural dimensions of AI privacy, this paper contributes to advancing trustworthy AI development and provides a foundation for future research.
An Investigation into the Performance of Non-Contrastive Self-Supervised Learning Methods for Network Intrusion Detection
Fard, Hamed, Schalau, Tobias, Wunder, Gerhard
Network intrusion detection, a well-explored cybersecurity field, has predominantly relied on supervised learning algorithms in the past two decades. However, their limitations in detecting only known anomalies prompt the exploration of alternative approaches. Motivated by the success of self-supervised learning in computer vision, there is a rising interest in adapting this paradigm for network intrusion detection. While prior research mainly delved into contrastive self-supervised methods, the efficacy of non-contrastive methods, in conjunction with encoder architectures serving as the representation learning backbone and augmentation strategies that determine what is learned, remains unclear for effective attack detection. This paper compares the performance of five non-contrastive self-supervised learning methods using three encoder architectures and six augmentation strategies. Ninety experiments are systematically conducted on two network intrusion detection datasets, UNSW-NB15 and 5G-NIDD. For each self-supervised model, the combination of encoder architecture and augmentation method yielding the highest average precision, recall, F1-score, and AUCROC is reported.
e96ed478dab8595a7dbda4cbcbee168f-Reviews.html
First provide a summary of the paper, and then address the following criteria: Quality, clarity, originality and significance. This paper proposes a simple latent factor model for one-shot learning with continuous outputs where very few observations are available. Specifically, it derives risk approximations in an asymptotic regime where the number of training examples is fixed and the number of features in the X space diverges. Based on principal component regression (PCR) estimator, two estimators including the bias-corrected estimator and the so-called oracle estimator are proposed and the bounds for the risks of these estimators are derived. These bounds provide insights into the significance of various parameters relevant to one-shot learning.
e5f6ad6ce374177eef023bf5d0c018b6-Reviews.html
First provide a summary of the paper, and then address the following criteria: Quality, clarity, originality and significance. This paper develops a model for multifurcating trees with edge lengths and observed data at the tree leaves; the model is based on the beta coalescent from the probability literature. The authors develop an MCMC inference scheme for their model, in which they draw on existing work that uses belief propagation to perform inference for the Kingman coalescent (an edge case of the beta coalescent in which all trees are binary). The particular challenge for inference here is that there are many more possible parent-child node relationships when parents can have multiple children (not just two). The authors seem to use a Dirichlet Process mixture model (DPMM) at each node to narrow down the space of possible children subsets to consider. As the authors note, even inference with the Kingman coalescent is a hard problem. In experiments, they compare to the Kingman coalescent and hierarchical agglomerative clustering. The Kingman coalescent is a popular modeling tool, so it is great to see a practical extension of the Kingman coalescent to the multifurcating case being explored for inference.
428fca9bc1921c25c5121f9da7815cde-Reviews.html
First provide a summary of the paper, and then address the following criteria: Quality, clarity, originality and significance. The set in question in this paper is the set of function with bounded partial derivatives up to order K. The authors' technique mimics the work of Thaler et al (ICALP 2012) only the authors decompose the queries not into regular polynomials (Chebyshev polynomials in the case of Thaler et al), but rather to trigonometric polynomial in this case. The bulk of the work is indeed to show that the abovementioned set of queries can be well-approximated by trigonometric polynomials. Having established that, adding Laplace noise to each monomial suffices to guarantee differential privacy.
2dffbc474aa176b6dc957938c15d0c8b-Reviews.html
First provide a summary of the paper, and then address the following criteria: Quality, clarity, originality and significance. This paper presents a Bayesian approach to state and parameter estimation in nonlinear state-space models, while also learning the transition dynamics through the use of a Gaussian process (GP) prior. The inference mechanism is based on particle Markov chain Monte Carlo (PMCMC) with the recently-introduced idea of ancestor sampling. The paper also discusses computational efficiencies to be had with respect to sparsity and low-rank Cholesky updates. This is a technically sound and strong paper with clear and accessible presentation.
291597a100aadd814d197af4f4bab3a7-Reviews.html
First provide a summary of the paper, and then address the following criteria: Quality, clarity, originality and significance. Online Learning with Costly Features and Labels Summary: The paper discusses a version of sequential prediction where there is a cost associated to obtaining features and labels. First, the case where labels are given but features are bought. Here the regret bound is of the form sqrt{2^d T}. The time dependence is as desired, and a lower bound shows that the exponential dependence in the dimension of the features space cannot be reduced in general.
28fc2782ea7ef51c1104ccf7b9bea13d-Reviews.html
First provide a summary of the paper, and then address the following criteria: Quality, clarity, originality and significance. The paper describes a method for local bandwidth selection in kernel regression models, which ensures adaptivity to local smoothness and dimension. Quality: The paper presents a useful result for adaptivity in kernel regression. The work is set out well. I think that it would be useful to have some more discussion of the bandwidth selection procedure.