Uncertainty
Domain Adaptation and Generalization on Functional Medical Images: A Systematic Survey
Sarafraz, Gita, Behnamnia, Armin, Hosseinzadeh, Mehran, Balapour, Ali, Meghrazi, Amin, Rabiee, Hamid R.
Machine learning algorithms have revolutionized different fields, including natural language processing, computer vision, signal processing, and medical data processing. Despite the excellent capabilities of machine learning algorithms in various tasks and areas, the performance of these models mainly deteriorates when there is a shift in the test and training data distributions. This gap occurs due to the violation of the fundamental assumption that the training and test data are independent and identically distributed (i.i.d). In real-world scenarios where collecting data from all possible domains for training is costly and even impossible, the i.i.d assumption can hardly be satisfied. The problem is even more severe in the case of medical images and signals because it requires either expensive equipment or a meticulous experimentation setup to collect data, even for a single domain. Additionally, the decrease in performance may have severe consequences in the analysis of medical records. As a result of such problems, the ability to generalize and adapt under distribution shifts (domain generalization (DG) and domain adaptation (DA)) is essential for the analysis of medical data. This paper provides the first systematic review of DG and DA on functional brain signals to fill the gap of the absence of a comprehensive study in this era. We provide detailed explanations and categorizations of datasets, approaches, and architectures used in DG and DA on functional brain images. We further address the attention-worthy future tracks in this field.
Rough sets models inspired by supra-topology structures - Artificial Intelligence Review
Our aim of writing this manuscript is to found novel rough-approximation operators inspired by an abstract structure called "supra-topology". This approach is more relaxed than topological ones and extends the scope of applications because an intersection condition of topology is dispensed. Firstly, we generate eight types of supra-topologies using \(N_k\)-neighborhood systems induced from any arbitrary relation. We elucidate the relationships between them and investigate the conditions under which some of them are identical. Then, we create new rough sets models from these supra-topologies and present the main characterizations of their lower and upper approximations.
Topical Segmentation of Spoken Narratives: A Test Case on Holocaust Survivor Testimonies
Wagner, Eitan, Keydar, Renana, Pinchevski, Amit, Abend, Omri
The task of topical segmentation is well studied, but previous work has mostly addressed it in the context of structured, well-defined segments, such as segmentation into paragraphs, chapters, or segmenting text that originated from multiple sources. We tackle the task of segmenting running (spoken) narratives, which poses hitherto unaddressed challenges. As a test case, we address Holocaust survivor testimonies, given in English. Other than the importance of studying these testimonies for Holocaust research, we argue that they provide an interesting test case for topical segmentation, due to their unstructured surface level, relative abundance (tens of thousands of such testimonies were collected), and the relatively confined domain that they cover. We hypothesize that boundary points between segments correspond to low mutual information between the sentences proceeding and following the boundary. Based on this hypothesis, we explore a range of algorithmic approaches to the task, building on previous work on segmentation that uses generative Bayesian modeling and state-of-the-art neural machinery. Compared to manually annotated references, we find that the developed approaches show considerable improvements over previous work.
Multivariate Quantile Function Forecaster
Kan, Kelvin, Aubet, François-Xavier, Januschowski, Tim, Park, Youngsuk, Benidis, Konstantinos, Ruthotto, Lars, Gasthaus, Jan
We propose Multivariate Quantile Function Forecaster (MQF$^2$), a global probabilistic forecasting method constructed using a multivariate quantile function and investigate its application to multi-horizon forecasting. Prior approaches are either autoregressive, implicitly capturing the dependency structure across time but exhibiting error accumulation with increasing forecast horizons, or multi-horizon sequence-to-sequence models, which do not exhibit error accumulation, but also do typically not model the dependency structure across time steps. MQF$^2$ combines the benefits of both approaches, by directly making predictions in the form of a multivariate quantile function, defined as the gradient of a convex function which we parametrize using input-convex neural networks. By design, the quantile function is monotone with respect to the input quantile levels and hence avoids quantile crossing. We provide two options to train MQF$^2$: with energy score or with maximum likelihood. Experimental results on real-world and synthetic datasets show that our model has comparable performance with state-of-the-art methods in terms of single time step metrics while capturing the time dependency structure.
Adaptive Robust Model Predictive Control via Uncertainty Cancellation
Sinha, Rohan, Harrison, James, Richards, Spencer M., Pavone, Marco
We propose a learning-based robust predictive control algorithm that compensates for significant uncertainty in the dynamics for a class of discrete-time systems that are nominally linear with an additive nonlinear component. Such systems commonly model the nonlinear effects of an unknown environment on a nominal system. We optimize over a class of nonlinear feedback policies inspired by certainty equivalent "estimate-and-cancel" control laws pioneered in classical adaptive control to achieve significant performance improvements in the presence of uncertainties of large magnitude, a setting in which existing learning-based predictive control algorithms often struggle to guarantee safety. In contrast to previous work in robust adaptive MPC, our approach allows us to take advantage of structure (i.e., the numerical predictions) in the a priori unknown dynamics learned online through function approximation. Our approach also extends typical nonlinear adaptive control methods to systems with state and input constraints even when we cannot directly cancel the additive uncertain function from the dynamics. We apply contemporary statistical estimation techniques to certify the system's safety through persistent constraint satisfaction with high probability. Moreover, we propose using Bayesian meta-learning algorithms that learn calibrated model priors to help satisfy the assumptions of the control design in challenging settings. Finally, we show in simulation that our method can accommodate more significant unknown dynamics terms than existing methods and that the use of Bayesian meta-learning allows us to adapt to the test environments more rapidly.
Designing Ecosystems of Intelligence from First Principles
Friston, Karl J, Ramstead, Maxwell J D, Kiefer, Alex B, Tschantz, Alexander, Buckley, Christopher L, Albarracin, Mahault, Pitliya, Riddhi J, Heins, Conor, Klein, Brennan, Millidge, Beren, Sakthivadivel, Dalton A R, Smithe, Toby St Clere, Koudahl, Magnus, Tremblay, Safae Essafi, Petersen, Capm, Fung, Kaiser, Fox, Jason G, Swanson, Steven, Mapes, Dan, René, Gabriel
This white paper lays out a vision of research and development in the field of artificial intelligence for the next decade (and beyond). Its denouement is a cyber-physical ecosystem of natural and synthetic sense-making, in which humans are integral participants$\unicode{x2014}$what we call ''shared intelligence''. This vision is premised on active inference, a formulation of adaptive behavior that can be read as a physics of intelligence, and which inherits from the physics of self-organization. In this context, we understand intelligence as the capacity to accumulate evidence for a generative model of one's sensed world$\unicode{x2014}$also known as self-evidencing. Formally, this corresponds to maximizing (Bayesian) model evidence, via belief updating over several scales: i.e., inference, learning, and model selection. Operationally, this self-evidencing can be realized via (variational) message passing or belief propagation on a factor graph. Crucially, active inference foregrounds an existential imperative of intelligent systems; namely, curiosity or the resolution of uncertainty. This same imperative underwrites belief sharing in ensembles of agents, in which certain aspects (i.e., factors) of each agent's generative world model provide a common ground or frame of reference. Active inference plays a foundational role in this ecology of belief sharing$\unicode{x2014}$leading to a formal account of collective intelligence that rests on shared narratives and goals. We also consider the kinds of communication protocols that must be developed to enable such an ecosystem of intelligences and motivate the development of a shared hyper-spatial modeling language and transaction protocol, as a first$\unicode{x2014}$and key$\unicode{x2014}$step towards such an ecology.
Accelerating Inverse Learning via Intelligent Localization with Exploratory Sampling
Zhang, Jiaxin, Bi, Sirui, Fung, Victor
In the scope of "AI for Science", solving inverse problems is a longstanding challenge in materials and drug discovery, where the goal is to determine the hidden structures given a set of desirable properties. Deep generative models are recently proposed to solve inverse problems, but these currently use expensive forward operators and struggle in precisely localizing the exact solutions and fully exploring the parameter spaces without missing solutions. In this work, we propose a novel approach (called iPage) to accelerate the inverse learning process by leveraging probabilistic inference from deep invertible models and deterministic optimization via fast gradient descent. Given a target property, the learned invertible model provides a posterior over the parameter space; we identify these posterior samples as an intelligent prior initialization which enables us to narrow down the search space. We then perform gradient descent to calibrate the inverse solutions within a local region. Meanwhile, a space-filling sampling is imposed on the latent space to better explore and capture all possible solutions. We evaluate our approach on three benchmark tasks and two created datasets with real-world applications from quantum chemistry and additive manufacturing, and find our method achieves superior performance compared to several state-of-the-art baseline methods. The iPage code is available at https://github.com/jxzhangjhu/MatDesINNe.
Robustness in Fatigue Strength Estimation
Weichert, Dorina, Kister, Alexander, Houben, Sebastian, Ernis, Gunar, Wrobel, Stefan
Fatigue strength estimation is a costly manual material characterization process in which state-of-the-art approaches follow a standardized experiment and analysis procedure. In this paper, we examine a modular, Machine Learning-based approach for fatigue strength estimation that is likely to reduce the number of experiments and, thus, the overall experimental costs. Despite its high potential, deployment of a new approach in a real-life lab requires more than the theoretical definition and simulation. Therefore, we study the robustness of the approach against misspecification of the prior and discretization of the specified loads. We identify its applicability and its advantageous behavior over the state-of-the-art methods, potentially reducing the number of costly experiments.
Initial Results for Pairwise Causal Discovery Using Quantitative Information Flow
Giori, Felipe, Figueiredo, Flavio
Pairwise Causal Discovery is the task of determining causal, anticausal, confounded or independence relationships from pairs of variables. Over the last few years, this challenging task has promoted not only the discovery of novel machine learning models aimed at solving the task, but also discussions on how learning the causal direction of variables may benefit machine learning overall. In this paper, we show that Quantitative Information Flow (QIF), a measure usually employed for measuring leakages of information from a system to an attacker, shows promising results as features for the task. In particular, experiments with real-world datasets indicate that QIF is statistically tied to the state of the art. Our initial results motivate further inquiries on how QIF relates to causality and what are its limitations.
Better Peer Grading through Bayesian Inference
Zarkoob, Hedayat, d'Eon, Greg, Podina, Lena, Leyton-Brown, Kevin
Peer grading is a powerful pedagogical tool. It benefits students by giving them exposure to others' perspectives; helping them to internalize evaluation criteria by applying them critically to peer work Lu and Law (2012); and offering them feedback from equal-status learners Topping (2009). Just as importantly, it gives instructors a way to make classes more scalable by shifting (some) grading workload away from course staff; effectively, this again benefits students, by giving them more opportunities for their work to be evaluated. In order for peer grading systems to be both useful to instructors and acceptable to students, they must produce grades that are sufficiently similar to those that an instructor would have given. This is a challenging task because individual peer graders will be biased (consistently give generous or harsh grades); noisy (the same grader could grade an assignment differently on different days); and potentially strategic (some students will enter insincere peer grades unrelated to a submission's quality if they can get away with it). Addressing these interrelated challenges has been a topic of academic study in Computer Science for at least the last two decades. The first methods for aggregating peer grades--and many others introduced more recently--produce point estimates of each assignment's grade and each grader's quality (Walsh, 2014; Chakraborty et al., 2018; Prajapati et al., 2020; de Alfaro and Shavlovsky, 2014; Hamer et al., 2005). At their best, methods that produce point estimates maximize the likelihood of the data given a model, e.g., by assigning each grader a "reliability" parameter and iteratively updating these parameters to best describe the reported grades.