Scientific Discovery
Guiding Scientific Discovery with Explanations Using DEMUD
Wagstaff, Kiri L. (Jet Propulsion Laboratory) | Lanza, Nina L. ( Los Alamos National Laboratory ) | Thompson, David R. ( Jet Propulsion Laboratory ) | Dietterich, Thomas G. ( Oregon State University ) | Gilmore, Martha S ( Wesleyan University )
In the era of large scientific data sets, there is an urgent need for methods to automatically prioritize data for review. At the same time, for any automated method to be adopted by scientists, it must make decisions that they can understand and trust. In this paper, we propose Discovery through Eigenbasis Modeling of Uninteresting Data (DEMUD), which uses principal components modeling and reconstruction error to prioritize data. DEMUDโs major advance is to offer domain-specific explanations for its prioritizations. We evaluated DEMUDโs ability to quickly identify diverse items of interest and the value of the explanations it provides. We found that DEMUD performs as well or better than existing class discovery methods and provides, uniquely, the first explanations for why those items are of interest. Further, in collaborations with planetary scientists, we found that DEMUD (1) quickly identifies very rare items of scientific value, (2) maintains high diversity in its selections, and (3) provides explanations that greatly improve human classification accuracy.
Epistemology of Modeling and Simulation: How can we gain Knowledge from Simulations?
Tolk, Andreas, Diallo, Saikou Y., Padilla, Jose J., Gore, Ross
Epistemology is the branch of philosophy that deals with gaining knowledge. It is closely related to ontology. The branch that deals with questions like "What is real?" and "What do we know?" as it provides these components. When using modeling and simulation, we usually imply that we are doing so to either apply knowledge, in particular when we are using them for training and teaching, or that we want to gain new knowledge, for example when doing analysis or conducting virtual experiments. This paper looks at the history of science to give a context to better cope with the question, how we can gain knowledge from simulation. It addresses aspects of computability and the general underlying mathematics, and applies the findings to validation and verification and development of federations. As simulations are understood as computable executable hypotheses, validation can be understood as hypothesis testing and theory building. The mathematical framework allows furthermore addressing some challenges when developing federations and the potential introduction of contradictions when composing different theories, as they are represented by the federated simulation systems.
Shikake as Affordance and Curation in Chance Discovery
Abe, Akinori (Chiba University)
In this paper first I introduce curation and affordance in chance discovery. According to Matsumura's definition, a shikake is a trigger to start a certain action or to change person's mind and behaviour. As a result of the action, all or part of problem will be solved. A chance and shikake are in a certain sense similar. In addition, affrodance seems to play a significant role in shikakeology. From the point I will discuss the relationships between chance discovery and Shikakeology.
Preface
Bridewell, Will (Stanford University) | Gil, Yolanda (University of Southern California) | Hirsh, Haym (Rutgers University) | Dam, Kerstin Kleese van (Pacific Northwest National Laboratory) | Steinhaeuser, Karsten (University of Minnesota)
Addressing the ambitious research agendas put forward by many scientific disciplines requires meeting a multitude of challenges in intelligent systems, information sciences, and human-computer interaction. Many aspects of the scientific discovery process are often largely manual and could be automated, improved, or made more efficient. Better interfaces for collaboration, visualization, and understanding would significantly improve scientific practice. Scientific data, publications, and tools could be published in open formats with appropriate semantic descriptions and metadata annotations to improve sharing and dissemination. Opportunities for broader participation in well-defined scientific tasks enable human contributors to provide large amounts of data, annotations, or complex processing results that could not otherwise be obtained. Improvements and innovations across the spectrum of scientific processes and activities will have a profound impact on the rate of scientific discoveries.
Discovery Informatics: AI Opportunities in Scientific Discovery
Gil, Yolanda (University of Southern California) | Hirsh, Haym (Rutgers University)
Artificial Intelligence researchers have long sought to understand and replicate processes of scientific discovery. This article discusses Discovery Informatics as an emerging area of research that builds on that tradition and applies principles of intelligent computing and information systems to understand, automate, improve, and innovate processes of scientific discovery.
Hypothesis Testing in Speckled Data with Stochastic Distances
Nascimento, Abraรฃo D. C., Cintra, Renato J., Frery, Alejandro C.
Images obtained with coherent illumination, as is the case of sonar, ultrasound-B, laser and Synthetic Aperture Radar -- SAR, are affected by speckle noise which reduces the ability to extract information from the data. Specialized techniques are required to deal with such imagery, which has been modeled by the G0 distribution and under which regions with different degrees of roughness and mean brightness can be characterized by two parameters; a third parameter, the number of looks, is related to the overall signal-to-noise ratio. Assessing distances between samples is an important step in image analysis; they provide grounds of the separability and, therefore, of the performance of classification procedures. This work derives and compares eight stochastic distances and assesses the performance of hypothesis tests that employ them and maximum likelihood estimation. We conclude that tests based on the triangular distance have the closest empirical size to the theoretical one, while those based on the arithmetic-geometric distances have the best power. Since the power of tests based on the triangular distance is close to optimum, we conclude that the safest choice is using this distance for hypothesis testing, even when compared with classical distances as Kullback-Leibler and Bhattacharyya.
Hypothesis testing using pairwise distances and associated kernels (with Appendix)
Sejdinovic, Dino, Gretton, Arthur, Sriperumbudur, Bharath, Fukumizu, Kenji
We provide a unifying framework linking two classes of statistics used in two-sample and independence testing: on the one hand, the energy distances and distance covariances from the statistics literature; on the other, distances between embeddings of distributions to reproducing kernel Hilbert spaces (RKHS), as established in machine learning. The equivalence holds when energy distances are computed with semimetrics of negative type, in which case a kernel may be defined such that the RKHS distance between distributions corresponds exactly to the energy distance. We determine the class of probability distributions for which kernels induced by semimetrics are characteristic (that is, for which embeddings of the distributions to an RKHS are injective). Finally, we investigate the performance of this family of kernels in two-sample and independence tests: we show in particular that the energy distance most commonly employed in statistics is just one member of a parametric family of kernels, and that other choices from this family can yield more powerful tests.
Discovery of Invariants through Automated Theory Formation
Llano, Maria Teresa, Ireland, Andrew, Pease, Alison
Refinement is a powerful mechanism for mastering the complexities that arise when formally modelling systems. Refinement also brings with it additional proof obligations -- requiring a developer to discover properties relating to their design decisions. With the goal of reducing this burden, we have investigated how a general purpose theory formation tool, HR, can be used to automate the discovery of such properties within the context of Event-B. Here we develop a heuristic approach to the automatic discovery of invariants and report upon a series of experiments that we undertook in order to evaluate our approach. The set of heuristics developed provides systematic guidance in tailoring HR for a given Event-B development. These heuristics are based upon proof-failure analysis, and have given rise to some promising results.
Notes on a New Philosophy of Empirical Science
This book presents a methodology and philosophy of empirical science based on large scale lossless data compression. In this view a theory is scientific if it can be used to build a data compression program, and it is valuable if it can compress a standard benchmark database to a small size, taking into account the length of the compressor itself. This methodology therefore includes an Occam principle as well as a solution to the problem of demarcation. Because of the fundamental difficulty of lossless compression, this type of research must be empirical in nature: compression can only be achieved by discovering and characterizing empirical regularities in the data. Because of this, the philosophy provides a way to reformulate fields such as computer vision and computational linguistics as empirical sciences: the former by attempting to compress databases of natural images, the latter by attempting to compress large text databases. The book argues that the rigor and objectivity of the compression principle should set the stage for systematic progress in these fields. The argument is especially strong in the context of computer vision, which is plagued by chronic problems of evaluation. The book also considers the field of machine learning. Here the traditional approach requires that the models proposed to solve learning problems be extremely simple, in order to avoid overfitting. However, the world may contain intrinsically complex phenomena, which would require complex models to understand. The compression philosophy can justify complex models because of the large quantity of data being modeled (if the target database is 100 Gb, it is easy to justify a 10 Mb model). The complex models and abstractions learned on the basis of the raw data (images, language, etc) can then be reused to solve any specific learning problem, such as face recognition or machine translation.
A novel family of non-parametric cumulative based divergences for point processes
Seth, Sohan, Il, Park, Brockmeier, Austin, Semework, Mulugeta, Choi, John, Francis, Joseph, Principe, Jose
Hypothesis testing on point processes has several applications such as model fitting, plasticity detection, and non-stationarity detection. Standard tools for hypothesis testing include tests on mean firing rate and time varying rate function. However, these statistics do not fully describe a point process and thus the tests can be misleading. In this paper, we introduce a family of non-parametric divergence measures for hypothesis testing. We extend the traditional Kolmogorov--Smirnov and Cramer--von-Mises tests for point process via stratification. The proposed divergence measures compare the underlying probability structure and, thus, is zero if and only if the point processes are the same. This leads to a more robust test of hypothesis. We prove consistency and show that these measures can be efficiently estimated from data. We demonstrate an application of using the proposed divergence as a cost function to find optimally matched spike trains.