Scientific Discovery
Benevolent AI: Shaping the Future of Scientific Discovery
Hi there, could you tell us a little about Benevolent AI and your motivations as a company? Despite the huge growth of knowledge, scientific discovery has not changed for 50 years. It's impossible for humans alone to process all the information potentially available to advance scientific research. A new scientific paper is published every 30 seconds, there are 10,000 updates to PubMed every day. Consequently, only a small fraction of globally generated scientific information can form'useable' knowledge.
Universal Hypothesis Testing with Kernels: Asymptotically Optimal Tests for Goodness of Fit
Zhu, Shengyu, Chen, Biao, Yang, Pengfei, Chen, Zhitang
We characterize the asymptotic performance of nonparametric goodness of fit testing, otherwise known as the universal hypothesis testing that dates back to Hoeffding (1965). The exponential decay rate of the type-II error probability is used as the asymptotic performance metric, hence an optimal test achieves the maximum decay rate subject to a constant level constraint on the type-I error probability. We show that two classes of Maximum Mean Discrepancy (MMD) based tests attain this optimality on $\mathbb R^d$, while a Kernel Stein Discrepancy (KSD) based test achieves a weaker one under this criterion. In the finite sample regime, these tests have similar statistical performance in our experiments, while the KSD based test is more computationally efficient. Key to our approach are Sanov's theorem from large deviation theory and recent results on the weak convergence properties of the MMD and KSD.
Cyclica CEO Naheed Kurji Says AI Could Create a New Paradigm for Drug Development - Top Chinese CRO, Biopharma News, Drug Development News WXPRESS
Toronto-based Cyclica President and CEO Naheed Kurji acknowledges that artificial intelligence (AI) is a transformative technology, but he contends it is not the "silver bullet" for drug discovery and development. Instead, he says that AI together with cloud-based computing could serve as a catalyst for a new approach to drug development. Kurji emphasizes that it is important to create "a virtual drug discovery ecosystem where a number of companies who are expert in their space come together and present a more holistic solution than any individual one could do itself because there is no one silver bullet to this problem. The market is so big and there are so many issues, one company can't do it alone." Kurji leads a five-year-old company that has developed and validated a cloud-based platform, called Ligand Express, which uses biophysics, bioinformatics and AI to help pharmaceutical companies navigate the drug discovery pipeline by assessing the safety and efficacy of drugs. The integrated platform enables companies to screen potential small-molecule drugs against repositories of structurally-characterized proteins or'proteomes' to identify significant protein targets. The platform then leverages AI to determine the biological relevance of these targets, and systems biology data to link this information to particular biological pathways or diseases. Kurji says Cyclica's platform, broadly launched in November 2017 already is being used by some of the top 50 pharma companies globally.
Machine learning & data discovery: You can't analyze what you can't find
Corporate data teams are under intense pressure to leverage machine learning and other technologies to transform everything from customer engagement to supply chain management. But data science doesn't do you much good without good data. So if you want to successfully compete and innovate in today's data-driven marketplace, you better be able to put the right information resources into the right hands -- quickly, efficiently, and comprehensively.
The Vestibular Domain
Although not all researchers agree on the exact bounds of scientific discovery, theory formation is clearly at the core of the domain. Relevant AI research done in scientific discovery includes Kocabas (1992); Karp (1989); Prager, Belanger, and De Mori (1989); Kulkarni and Simon (1988); and Langley et al. (1987). I consider model-based discovery to be a diagnosis and design problem. More precisely, modelbased-theory refinement can be seen as a four-step process: (1) gather data, (2) compare the data to model-based predictions, (3) identify the sources of discrepancies between the predictions and the field data, and (4) fix these discrepancies by modifying the model. The first three steps are traditionally addressed by diagnosis systems, but the fourth step requires design techniques.
Toward Automated Discovery in the Biological Sciences
Knowledge discovery programs in the biological sciences require flexibility in the use of symbolic data and semantic information. Because of the volume of nonnumeric, as well as numeric, data, the programs must be able to explore a large space of possibly interesting relationships to discover those that are novel and interesting. Thus, the framework for the discovery program must facilitate proposing and selecting the next task to perform and performing the selected tasks. The framework we describe, called the agenda-and justificationbased framework, has several properties that are desirable in semiautonomous discovery systems: It provides a mechanism for estimating the plausibility of tasks, it uses heuristics to propose and perform tasks, and it facilitates the encoding of general discovery strategies and the use of background knowledge. The complexity of the data and the underlying mechanisms argue for providing computer assistance to biologists.
Machine Discovery of Chemical Reaction Pathways
A fundamental question in AI is what mechanisms suffice for computer programs to make scientific discoveries. My Ph.D. thesis (Valdés-Pérez 1990e) addresses this question by automating the following scientific task to a significant extent: Given observed data about a particular chemical reaction, discover the underlying set of reaction steps from starting materials to products, that is, elucidate the reaction pathway. My scientific contribution is to describe and interpret the design of a system that forms plausible explanatory hypotheses about dynamic processes in science and that proposes unseen entities in a manner justified by simplicity. Some byproducts of the thesis are several novel contributions to chemistry knowledge in addition to scientific tools of immediate use. Chapter 1 surveys previous work in machine discovery, focusing on work that involved assembling an extensive amount of knowledge particular to a domain.
The central thesis of my dissertation (Kocabas 1989)
I describe a formalism for organizing knowledge into such functional categories and some of its implementations. In this formalism, descriptive scientific knowledge is classified into seven categories. The categorization formalism allows complex propositions to be analyzed into their simple constituents; in turn, these constituents can be maintained in their categories. They can then be combined using a simple transformation function to form complex constructs such as frames and schemata. To demonstrate its viability, I implemented this knowledge organization in four computational models of scientific reasoning and discovery that operate in astronomy and particle physics, incorporating various methods of learning and theory revision.
Adaptive Active Hypothesis Testing under Limited Information
We consider the problem of active sequential hypothesis testing where a Bayesian decision maker must infer the true hypothesis from a set of hypotheses. The decision maker may choose for a set of actions, where the outcome of an action is corrupted by independent noise. In this paper we consider a special case where the decision maker has limited knowledge about the distribution of observations for each action, in that only a binary value is observed. Our objective is to infer the true hypothesis with low error, while minimizing the number of action sampled. Our main results include the derivation of a lower bound on sample size for our system under limited knowledge and the design of an active learning policy that matches this lower bound and outperforms similar known algorithms.
Baba Vanga, Bulgarian Mystic Known For 9/11 Prediction, Forecasts 2018 Scientific Discovery
Baba Vanga, a blind Bulgarian mystic who died in 1996, is said to have predicted a huge scientific discovery for 2018: a new form of energy on the planet Venus. With no planned missions to Venus this year, her prediction is not expected to come to fruition. More than 20 years after her death, people are waiting to see if Baba Vanga's prophecies for 2018 will come to pass. They reportedly include the Venus discovery, as well as China passing the United States in economic power, although it is unclear where these statements are coming from. Baba Vanga, whose real name was Vangelia Gushterova, was blinded as a child during a tornado.