Bayesian Learning
Neuro-symbolic Explainable Artificial Intelligence Twin for Zero-touch IoE in Wireless Network
Munir, Md. Shirajum, Kim, Ki Tae, Adhikary, Apurba, Saad, Walid, Shetty, Sachin, Park, Seong-Bae, Hong, Choong Seon
Explainable artificial intelligence (XAI) twin systems will be a fundamental enabler of zero-touch network and service management (ZSM) for sixth-generation (6G) wireless networks. A reliable XAI twin system for ZSM requires two composites: an extreme analytical ability for discretizing the physical behavior of the Internet of Everything (IoE) and rigorous methods for characterizing the reasoning of such behavior. In this paper, a novel neuro-symbolic explainable artificial intelligence twin framework is proposed to enable trustworthy ZSM for a wireless IoE. The physical space of the XAI twin executes a neural-network-driven multivariate regression to capture the time-dependent wireless IoE environment while determining unconscious decisions of IoE service aggregation. Subsequently, the virtual space of the XAI twin constructs a directed acyclic graph (DAG)-based Bayesian network that can infer a symbolic reasoning score over unconscious decisions through a first-order probabilistic language model. Furthermore, a Bayesian multi-arm bandits-based learning problem is proposed for reducing the gap between the expected explained score and the current obtained score of the proposed neuro-symbolic XAI twin. To address the challenges of extensible, modular, and stateless management functions in ZSM, the proposed neuro-symbolic XAI twin framework consists of two learning systems: 1) an implicit learner that acts as an unconscious learner in physical space, and 2) an explicit leaner that can exploit symbolic reasoning based on implicit learner decisions and prior evidence. Experimental results show that the proposed neuro-symbolic XAI twin can achieve around 96.26% accuracy while guaranteeing from 18% to 44% more trust score in terms of reasoning and closed-loop automation.
Graph Neural Networks for Low-Energy Event Classification & Reconstruction in IceCube
Abbasi, R., Ackermann, M., Adams, J., Aggarwal, N., Aguilar, J. A., Ahlers, M., Ahrens, M., Alameddine, J. M., Alves, A. A. Jr., Amin, N. M., Andeen, K., Anderson, T., Anton, G., Argรผelles, C., Ashida, Y., Athanasiadou, S., Axani, S., Bai, X., V., A. Balagopal, Baricevic, M., Barwick, S. W., Basu, V., Bay, R., Beatty, J. J., Becker, K. -H., Tjus, J. Becker, Beise, J., Bellenghi, C., Benda, S., BenZvi, S., Berley, D., Bernardini, E., Besson, D. Z., Binder, G., Bindig, D., Blaufuss, E., Blot, S., Bontempo, F., Book, J. Y., Borowka, J., Meneguolo, C. Boscolo, Bรถser, S., Botner, O., Bรถttcher, J., Bourbeau, E., Braun, J., Brinson, B., Brostean-Kaiser, J., Burley, R. T., Busse, R. S., Campana, M. A., Carnie-Bronca, E. G., Chen, C., Chen, Z., Chirkin, D., Choi, K., Clark, B. A., Classen, L., Coleman, A., Collin, G. H., Connolly, A., Conrad, J. M., Coppin, P., Correa, P., Countryman, S., Cowen, D. F., Cross, R., Dappen, C., Dave, P., De Clercq, C., DeLaunay, J. J., Lรณpez, D. Delgado, Dembinski, H., Deoskar, K., Desai, A., Desiati, P., de Vries, K. D., de Wasseige, G., DeYoung, T., Diaz, A., Dรญaz-Vรฉlez, J. C., Dittmer, M., Dujmovic, H., DuVernois, M. A., Ehrhardt, T., Eller, P., Engel, R., Erpenbeck, H., Evans, J., Evenson, P. A., Fan, K. L., Fazely, A. R., Fedynitch, A., Feigl, N., Fiedlschuster, S., Fienberg, A. T., Finley, C., Fischer, L., Fox, D., Franckowiak, A., Friedman, E., Fritz, A., Fรผrst, P., Gaisser, T. K., Gallagher, J., Ganster, E., Garcia, A., Garrappa, S., Gerhardt, L., Ghadimi, A., Glaser, C., Glauch, T., Glรผsenkamp, T., Goehlke, N., Gonzalez, J. G., Goswami, S., Grant, D., Gray, S. J., Grรฉgoire, T., Griswold, S., Gรผnther, C., Gutjahr, P., Haack, C., Hallgren, A., Halliday, R., Halve, L., Halzen, F., Hamdaoui, H., Minh, M. Ha, Hanson, K., Hardin, J., Harnisch, A. A., Hatch, P., Haungs, A., Helbing, K., Hellrung, J., Henningsen, F., Heuermann, L., Hickford, S., Hill, C., Hill, G. C., Hoffman, K. D., Hoshina, K., Hou, W., Huber, T., Hultqvist, K., Hรผnnefeld, M., Hussain, R., Hymon, K., In, S., Iovine, N., Ishihara, A., Jansson, M., Japaridze, G. S., Jeong, M., Jin, M., Jones, B. J. P., Kang, D., Kang, W., Kang, X., Kappes, A., Kappesser, D., Kardum, L., Karg, T., Karl, M., Karle, A., Katz, U., Kauer, M., Kelley, J. L., Kheirandish, A., Kin, K., Kiryluk, J., Klein, S. R., Kochocki, A., Koirala, R., Kolanoski, H., Kontrimas, T., Kรถpke, L., Kopper, C., Koskinen, D. J., Koundal, P., Kovacevich, M., Kowalski, M., Kozynets, T., Krupczak, E., Kun, E., Kurahashi, N., Lad, N., Gualda, C. Lagunas, Larson, M. J., Lauber, F., Lazar, J. P., Lee, J. W., Leonard, K., Leszczyลska, A., Lincetto, M., Liu, Q. R., Liubarska, M., Lohfink, E., Love, C., Mariscal, C. J. Lozano, Lu, L., Lucarelli, F., Ludwig, A., Luszczak, W., Lyu, Y., Ma, W. Y., Madsen, J., Mahn, K. B. M., Makino, Y., Mancina, S., Sainte, W. Marie, Mariล, I. C., Marka, S., Marka, Z., Marsee, M., Martinez-Soler, I., Maruyama, R., McElroy, T., McNally, F., Mead, J. V., Meagher, K., Mechbal, S., Medina, A., Meier, M., Meighen-Berger, S., Merckx, Y., Micallef, J., Mockler, D., Montaruli, T., Moore, R. W., Morse, R., Moulai, M., Mukherjee, T., Naab, R., Nagai, R., Naumann, U., Nayerhoda, A., Necker, J., Neumann, M., Niederhausen, H., Nisa, M. U., Nowicki, S. C., Pollmann, A. Obertacke, Oehler, M., Oeyen, B., Olivas, A., Orsoe, R., Osborn, J., O'Sullivan, E., Pandya, H., Pankova, D. V., Park, N., Parker, G. K., Paudel, E. N., Paul, L., Heros, C. Pรฉrez de los, Peters, L., Petersen, T. C., Peterson, J., Philippen, S., Pieper, S., Pizzuto, A., Plum, M., Popovych, Y., Porcelli, A., Rodriguez, M. Prado, Pries, B., Procter-Murphy, R., Przybylski, G. T., Raab, C., Rack-Helleis, J., Rameez, M., Rawlins, K., Rechav, Z., Rehman, A., Reichherzer, P., Renzi, G., Resconi, E., Reusch, S., Rhode, W., Richman, M., Riedel, B., Roberts, E. J., Robertson, S., Rodan, S., Roellinghoff, G., Rongen, M., Rott, C., Ruhe, T., Ruohan, L., Ryckbosch, D., Cantu, D. Rysewyk, Safa, I., Saffer, J., Salazar-Gallegos, D., Sampathkumar, P., Herrera, S. E. Sanchez, Sandrock, A., Santander, M., Sarkar, S., Sarkar, S., Schaufel, M., Schieler, H., Schindler, S., Schlueter, B., Schmidt, T., Schneider, J., Schrรถder, F. G., Schumacher, L., Schwefer, G., Sclafani, S., Seckel, D., Seunarine, S., Sharma, A., Shefali, S., Shimizu, N., Silva, M., Skrzypek, B., Smithers, B., Snihur, R., Soedingrekso, J., Sรธgaard, A., Soldin, D., Spannfellner, C., Spiczak, G. M., Spiering, C., Stamatikos, M., Stanev, T., Stein, R., Stezelberger, T., Stรผrwald, T., Stuttard, T., Sullivan, G. W., Taboada, I., Ter-Antonyan, S., Thompson, W. G., Thwaites, J., Tilav, S., Tollefson, K., Tรถnnis, C., Toscano, S., Tosi, D., Trettin, A., Tung, C. F., Turcotte, R., Twagirayezu, J. P., Ty, B., Elorrieta, M. A. Unland, Upshaw, K., Valtonen-Mattila, N., Vandenbroucke, J., van Eijndhoven, N., Vannerom, D., van Santen, J., Vara, J., Veitch-Michaelis, J., Verpoest, S., Veske, D., Walck, C., Wang, W., Watson, T. B., Weaver, C., Weigel, P., Weindl, A., Weldert, J., Wendt, C., Werthebach, J., Weyrauch, M., Whitehorn, N., Wiebusch, C. H., Willey, N., Williams, D. R., Wolf, M., Wrede, G., Wulff, J., Xu, X. W., Yanez, J. P., Yildizci, E., Yoshida, S., Yu, S., Yuan, T., Zhang, Z., Zhelnin, P.
IceCube, a cubic-kilometer array of optical sensors built to detect atmospheric and astrophysical neutrinos between 1 GeV and 1 PeV, is deployed 1.45 km to 2.45 km below the surface of the ice sheet at the South Pole. The classification and reconstruction of events from the in-ice detectors play a central role in the analysis of data from IceCube. Reconstructing and classifying events is a challenge due to the irregular detector geometry, inhomogeneous scattering and absorption of light in the ice and, below 100 GeV, the relatively low number of signal photons produced per event. To address this challenge, it is possible to represent IceCube events as point cloud graphs and use a Graph Neural Network (GNN) as the classification and reconstruction method. The GNN is capable of distinguishing neutrino events from cosmic-ray backgrounds, classifying different neutrino event types, and reconstructing the deposited energy, direction and interaction vertex. Based on simulation, we provide a comparison in the 1-100 GeV energy range to the current state-of-the-art maximum likelihood techniques used in current IceCube analyses, including the effects of known systematic uncertainties. For neutrino event classification, the GNN increases the signal efficiency by 18% at a fixed false positive rate (FPR), compared to current IceCube methods. Alternatively, the GNN offers a reduction of the FPR by over a factor 8 (to below half a percent) at a fixed signal efficiency. For the reconstruction of energy, direction, and interaction vertex, the resolution improves by an average of 13%-20% compared to current maximum likelihood techniques in the energy range of 1-30 GeV. The GNN, when run on a GPU, is capable of processing IceCube events at a rate nearly double of the median IceCube trigger rate of 2.7 kHz, which opens the possibility of using low energy neutrinos in online searches for transient events.
DeepVARwT: Deep Learning for a VAR Model with Trend
The vector autoregressive (VAR) model has been used to describe the dependence within and across multiple time series. This is a model for stationary time series which can be extended to allow the presence of a deterministic trend in each series. Detrending the data either parametrically or nonparametrically before fitting the VAR model gives rise to more errors in the latter part. In this study, we propose a new approach called DeepVARwT that employs deep learning methodology for maximum likelihood estimation of the trend and the dependence structure at the same time. A Long Short-Term Memory (LSTM) network is used for this purpose. To ensure the stability of the model, we enforce the causality condition on the autoregressive coefficients using the transformation of Ansley & Kohn (1986). We provide a simulation study and an application to real data. In the simulation study, we use realistic trend functions generated from real data and compare the estimates with true function/parameter values. In the real data application, we compare the prediction performance of this model with state-of-the-art models in the literature.
Aggregating Crowdsourced and Automatic Judgments to Scale Up a Corpus of Anaphoric Reference for Fiction and Wikipedia Texts
Yu, Juntao, Paun, Silviu, Camilleri, Maris, Garcia, Paloma Carretero, Chamberlain, Jon, Kruschwitz, Udo, Poesio, Massimo
Although several datasets annotated for anaphoric reference/coreference exist, even the largest such datasets have limitations in terms of size, range of domains, coverage of anaphoric phenomena, and size of documents included. Yet, the approaches proposed to scale up anaphoric annotation haven't so far resulted in datasets overcoming these limitations. In this paper, we introduce a new release of a corpus for anaphoric reference labelled via a game-with-a-purpose. This new release is comparable in size to the largest existing corpora for anaphoric reference due in part to substantial activity by the players, in part thanks to the use of a new resolve-and-aggregate paradigm to 'complete' markable annotations through the combination of an anaphoric resolver and an aggregation method for anaphoric reference. The proposed method could be adopted to greatly speed up annotation time in other projects involving games-with-a-purpose. In addition, the corpus covers genres for which no comparable size datasets exist (Fiction and Wikipedia); it covers singletons and non-referring expressions; and it includes a substantial number of long documents (> 2K in length).
Synthetic Model Combination: An Instance-wise Approach to Unsupervised Ensemble Learning
Chan, Alex J., van der Schaar, Mihaela
Consider making a prediction over new test data without any opportunity to learn from a training set of labelled data - instead given access to a set of expert models and their predictions alongside some limited information about the dataset used to train them. In scenarios from finance to the medical sciences, and even consumer practice, stakeholders have developed models on private data they either cannot, or do not want to, share. Given the value and legislation surrounding personal information, it is not surprising that only the models, and not the data, will be released - the pertinent question becoming: how best to use these models? Previous work has focused on global model selection or ensembling, with the result of a single final model across the feature space. Machine learning models perform notoriously poorly on data outside their training domain however, and so we argue that when ensembling models the weightings for individual instances must reflect their respective domains - in other words models that are more likely to have seen information on that instance should have more attention paid to them. We introduce a method for such an instance-wise ensembling of models, including a novel representation learning step for handling sparse high-dimensional domains. Finally, we demonstrate the need and generalisability of our method on classical machine learning tasks as well as highlighting a real world use case in the pharmacological setting of vancomycin precision dosing.
FasterRisk: Fast and Accurate Interpretable Risk Scores
Liu, Jiachang, Zhong, Chudi, Li, Boxuan, Seltzer, Margo, Rudin, Cynthia
Over the last century, risk scores have been the most popular form of predictive model used in healthcare and criminal justice. Risk scores are sparse linear models with integer coefficients; often these models can be memorized or placed on an index card. Typically, risk scores have been created either without data or by rounding logistic regression coefficients, but these methods do not reliably produce high-quality risk scores. Recent work used mathematical programming, which is computationally slow. We introduce an approach for efficiently producing a collection of high-quality risk scores learned from data. Specifically, our approach produces a pool of almost-optimal sparse continuous solutions, each with a different support set, using a beam-search algorithm. Each of these continuous solutions is transformed into a separate risk score through a "star ray" search, where a range of multipliers are considered before rounding the coefficients sequentially to maintain low logistic loss. Our algorithm returns all of these high-quality risk scores for the user to consider. This method completes within minutes and can be valuable in a broad variety of applications.
Excess risk analysis for epistemic uncertainty with application to variational inference
Futami, Futoshi, Iwata, Tomoharu, Ueda, Naonori, Sato, Issei, Sugiyama, Masashi
Bayesian deep learning plays an important role especially for its ability evaluating epistemic uncertainty (EU). Due to computational complexity issues, approximation methods such as variational inference (VI) have been used in practice to obtain posterior distributions and their generalization abilities have been analyzed extensively, for example, by PAC-Bayesian theory; however, little analysis exists on EU, although many numerical experiments have been conducted on it. In this study, we analyze the EU of supervised learning in approximate Bayesian inference by focusing on its excess risk. First, we theoretically show the novel relations between generalization error and the widely used EU measurements, such as the variance and mutual information of predictive distribution, and derive their convergence behaviors. Next, we clarify how the objective function of VI regularizes the EU. With this analysis, we propose a new objective function for VI that directly controls the prediction performance and the EU based on the PAC-Bayesian theory. Numerical experiments show that our algorithm significantly improves the EU evaluation over the existing VI methods.
Generalization Bounds for Gradient Methods via Discrete and Continuous Prior
Luo, Xuanyuan, Bei, Luo, Li, Jian
Proving algorithm-dependent generalization error bounds for gradient-type optimization methods has attracted significant attention recently in learning theory. However, most existing trajectory-based analyses require either restrictive assumptions on the learning rate (e.g., fast decreasing learning rate), or continuous injected noise (such as the Gaussian noise in Langevin dynamics). In this paper, we introduce a new discrete data-dependent prior to the PAC-Bayesian framework, and prove a high probability generalization bound of order $O(\frac{1}{n}\cdot \sum_{t=1}^T(\gamma_t/\varepsilon_t)^2\left\|{\mathbf{g}_t}\right\|^2)$ for Floored GD (i.e. a version of gradient descent with precision level $\varepsilon_t$), where $n$ is the number of training samples, $\gamma_t$ is the learning rate at step $t$, $\mathbf{g}_t$ is roughly the difference of the gradient computed using all samples and that using only prior samples. $\left\|{\mathbf{g}_t}\right\|$ is upper bounded by and and typical much smaller than the gradient norm $\left\|{\nabla f(W_t)}\right\|$. We remark that our bound holds for nonconvex and nonsmooth scenarios. Moreover, our theoretical results provide numerically favorable upper bounds of testing errors (e.g., $0.037$ on MNIST). Using a similar technique, we can also obtain new generalization bounds for certain variants of SGD. Furthermore, we study the generalization bounds for gradient Langevin Dynamics (GLD). Using the same framework with a carefully constructed continuous prior, we show a new high probability generalization bound of order $O(\frac{1}{n} + \frac{L^2}{n^2}\sum_{t=1}^T(\gamma_t/\sigma_t)^2)$ for GLD. The new $1/n^2$ rate is due to the concentration of the difference between the gradient of training samples and that of the prior.
Weakly supervised causal representation learning
Brehmer, Johann, de Haan, Pim, Lippe, Phillip, Cohen, Taco
Learning high-level causal representations together with a causal model from unstructured low-level data such as pixels is impossible from observational data alone. We prove under mild assumptions that this representation is however identifiable in a weakly supervised setting. This involves a dataset with paired samples before and after random, unknown interventions, but no further labels. We then introduce implicit latent causal models, variational autoencoders that represent causal variables and causal structure without having to optimize an explicit discrete graph structure. On simple image data, including a novel dataset of simulated robotic manipulation, we demonstrate that such models can reliably identify the causal structure and disentangle causal variables.
Bayesian adaptive and interpretable functional regression for exposure profiles
Pollutant exposure during gestation is a known and adverse factor for birth and health outcomes. However, the links between prenatal air pollution exposures and educational outcomes are less clear, in particular the critical windows of susceptibility during pregnancy. Using a large cohort of students in North Carolina, we study the link between prenatal daily $\mbox{PM}_{2.5}$ exposure and 4th end-of-grade reading scores. We develop and apply a locally adaptive and highly scalable Bayesian regression model for scalar responses with functional and scalar predictors. The proposed model pairs a B-spline basis expansion with dynamic shrinkage priors to capture both smooth and rapidly-changing features in the regression surface. The model is accompanied by a new decision analysis approach for functional regression that extracts the critical windows of susceptibility and guides the model interpretations. These tools help to identify and address broad limitations with the interpretability of functional regression models. Simulation studies demonstrate more accurate point estimation, more precise uncertainty quantification, and far superior window selection than existing approaches. Leveraging the proposed modeling, computational, and decision analysis framework, we conclude that prenatal $\mbox{PM}_{2.5}$ exposure during early and late pregnancy is most adverse for 4th end-of-grade reading scores.