Goto

Collaborating Authors

 Technology




Convolutional Monge Mapping Normalization for learning on sleep data

Neural Information Processing Systems

In many machine learning applications on signals and biomedical data, especially electroencephalogram (EEG), one major challenge is the variability of the data across subjects, sessions, and hardware devices. In this work, we propose a new method called Convolutional Monge Mapping Normalization (CMMN), which consists in filtering the signals in order to adapt their power spectrum density (PSD) to a Wasserstein barycenter estimated on training data. CMMN relies on novel closed-form solutions for optimal transport mappings and barycenters and provides individual test time adaptation to new data without needing to retrain a prediction model. Numerical experiments on sleep EEG data show that CMMN leads to significant and consistent performance gains independent from the neural network architecture when adapting between subjects, sessions, and even datasets collected with different hardware. Notably our performance gain is on par with much more numerically intensive Domain Adaptation (DA) methods and can be used in conjunction with those for even better performances.


Model Derivation We write the joint posterior as, 2|Y,Z/ (Y|X,, 2) (| 2) (2) (13) / (2) N/2exp(1 2 2 (Y Z)Tdiag(x(Z)) (Y Z)) (2) 1exp(1 2 2 T) (2) (1+

Neural Information Processing Systems

The intermediate steps can be found in [67]. Derivation of Posterior Predictive Note, this derivation takes the priors to be set as in BayesLIME or BayesSHAP, namely, with values close to zero. We apply the identity from equation 17 to derive this posterior. In these derivations, the perturbation matrices Z have elements Zij 2{ 0,1} where each Zij Bernoulli(0.5). Note, in these proofs, we take take the priors to be set as in BayesLIME and BayesSHAP, i.e., they have hyperparameter values close to 0. B.1 Proof of Theorem 3.3 Note that we use N to denote the total perturbations while S denotes the perturabtions collected so far.



Data-Efficient Instance Generation from Instance Discrimination

Neural Information Processing Systems

Generative Adversarial Networks (GANs) have significantly advanced image synthesis, however, the synthesis quality drops significantly given a limited amount of training data. To improve the data efficiency of GAN training, prior work typically employs data augmentation to mitigate the overfitting of the discriminator yet still learn the discriminator with a bi-classification (i.e., real vs.


Supplementary Material AProof of Proposition 2

Neural Information Processing Systems

Proposition 2. (From main text) The Bayes error of flow models is monotonically increasing in . That is, for 0 < 0, we have that EBayes(ห†p) EBayes(ห†p 0). B.1 Hardness of Classes In addition to measuring the difficulty of classification tasks relative to one another, it also may be of interest to evaluate the relative difficulty of individual classes within a particular task. A natural way to do this is by looking at the error of one-vs-all classification tasks. The optimal Bayes classifier in this task is CBayes(x)= 0 if logpj(x) logp j(x), 1 otherwise .


Evaluating State-of-the-Art Classification Models Against Bayes Optimality

Neural Information Processing Systems

Evaluating the inherent difficulty of a given data-driven classification problem is important for establishing absolute benchmarks and evaluating progress in the field. To this end, a natural quantity to consider is the Bayes error, which measures the optimal classification error theoretically achievable for a given data distribution. While generally an intractable quantity, we show that we can compute the exact Bayes error of generative models learned using normalizing flows. Our technique relies on a fundamental result, which states that the Bayes error is invariant under invertible transformation. Therefore, we can compute the exact Bayes error of the learned flow models by computing it for Gaussian base distributions, which can be done efficiently using Holmes-Diaconis-Ross integration. Moreover, we show that by varying the temperature of the learned flow models, we can generate synthetic datasets that closely resemble standard benchmark datasets, but with almost any desired Bayes error. We use our approach to conduct a thorough investigation of state-of-the-art classification models, and find that in some -- but not all -- cases, these models are capable of obtaining accuracy very near optimal. Finally, we use our method to evaluate the intrinsic "hardness" of standard benchmark datasets.


Optimistic Posterior Sampling for Reinforcement Learning with Few Samples and Tight Guarantees

Neural Information Processing Systems

We consider reinforcement learning in an environment modeled by an episodic, finite, stage-dependent Markov decision process of horizon H with S states, and A actions. The performance of an agent is measured by the regret after interacting with the environment for T episodes. We propose an optimistic posterior sampling algorithm for reinforcement learning (OPSRL), a simple variant of posterior sampling that only needs a number of posterior samples logarithmic in H, S, A, and T per state-action pair.