Goto

Collaborating Authors

 Scientific Discovery


This Government Agency Is A Surprising Powerhouse In AI

#artificialintelligence

Among the many departments and agencies within the United States federal government, the US Department of Energy (DOE) stands out as one of the most science, technology, and innovation-focused. This should come as little surprise to those who know the DOE's storied history with its breakthrough labs, world-leading research institutions, and highly educated staff. Since World War II, the DOE has been at the forefront of most of the groundbreaking and world-changing revolutions in science and technology including the development and harnessing of nuclear energy, innovations in genomics including the DOE initiative Human Genome Project, work in high-performance computing, and many other research-oriented efforts. In fact, the DOE supports more research in the physical sciences than any other US federal agency, providing more than 40% of US funding in computing, physics, chemistry, materials science, and other area through a system of national laboratories including Lawrence Berkeley National Laboratory, Oak Ridge National Laboratory, Argonne National Laboratory, Ames Laboratory, Brookhaven National Laboratory, Los Alamos National Laboratory, Sandia National Labs, Lawrence Livermore National Laboratory, the SLAC National Accelerator Laboratory, and dozens more institutions. Until very recently, the DOE also ran the world's top two fastest supercomputers: Summit and Sierra.


Data is the new gold. This is how it can benefit everyone โ€“ while harming no one

#artificialintelligence

COVID-19 has dealt the world a twin crisis. We face not only our greatest global health shock but also our greatest economic shock in a century. With these dual crises comes a twin watershed moment. First, whether for school, work, health or keeping in touch with family and friends, we have realized the deep value of digital technologies. Second, the appetite for change (arguably a more challenging shift to achieve) has grown significantly.


Impact Biomedical Initiates Quantum, a New Frontier in Pharmaceutical Development

#artificialintelligence

Impact Biomedical, a wholly-owned subsidiary of SGX-listed Singapore eDevelopment, has announced the initiation of Quantum, a research program designed as a solution to the'patent cliff', the impending pharmaceutical threat. A patent cliff looms when patents for blockbuster drugs expire without being replaced with new drugs, and pharmaceutical companies experience an abrupt decrease in revenue, reducing overall pharmaceutical innovation globally, including crucial research into new methods to prevent and treat illnesses. Impact, through their strategic partner Global Research and Discovery Group Sciences (GRDG), has created a solution called Quantum, a new frontier in pharmaceutical development. Quantum is a new class of medicinal chemistry that uses advanced methods to boost efficacy and persistence of natural compounds and existing drugs while maintaining the safety profile of the original molecules. Instead of modifying functional groups, as is typically done presently in drug discovery, this new technique alters the behavior of molecules at the sub-molecular level.


5G and AI Power New Paradigms

#artificialintelligence

We talk a lot about speed and capacity when it comes to 5G. But some of its greatest potential lies in a capability we haven't seen before: Distributed intelligence. In the same way that smartphones ignited the app economy and changed how we live, the next ...generation of wireless networks will make AI applications accessible to any connected device.


The Lasso with general Gaussian designs with applications to hypothesis testing

arXiv.org Machine Learning

The Lasso is a method for high-dimensional regression, which is now commonly used when the number of covariates $p$ is of the same order or larger than the number of observations $n$. Classical asymptotic normality theory is not applicable for this model due to two fundamental reasons: $(1)$ The regularized risk is non-smooth; $(2)$ The distance between the estimator $\bf \widehat{\theta}$ and the true parameters vector $\bf \theta^\star$ cannot be neglected. As a consequence, standard perturbative arguments that are the traditional basis for asymptotic normality fail. On the other hand, the Lasso estimator can be precisely characterized in the regime in which both $n$ and $p$ are large, while $n/p$ is of order one. This characterization was first obtained in the case of standard Gaussian designs, and subsequently generalized to other high-dimensional estimation procedures. Here we extend the same characterization to Gaussian correlated designs with non-singular covariance structure. This characterization is expressed in terms of a simpler ``fixed design'' model. We establish non-asymptotic bounds on the distance between distributions of various quantities in the two models, which hold uniformly over signals $\bf \theta^\star$ in a suitable sparsity class, and values of the regularization parameter. As applications, we study the distribution of the debiased Lasso, and show that a degrees-of-freedom correction is necessary for computing valid confidence intervals.


Hypothesis Testing

#artificialintelligence

The confidence intervals are the type of estimate which give us an estimation of where the parameters are located. Nonetheless, when we have to make a decision we need a'yes' or'no' answer, to do so we will perform a test known as Hypothesis Testing. Steps in data-driven decision making.: A hypothesis is an idea that can be tested. For example, apples in London are expensive.


Inferences and Modal Vocabulary

arXiv.org Artificial Intelligence

Deduction is the one of the major forms of inferences and commonly used in formal logic. This kind of inference has the feature of monotonicity, which can be problematic. There are different types of inferences that are not monotonic, e.g. abductive inferences. The debate between advocates and critics of abduction as a useful instrument can be reconstructed along the issue, how an abductive inference warrants to pick out one hypothesis as the best one. But how can the goodness of an inference be assessed? Material inferences express good inferences based on the principle of material incompatibility. Material inferences are based on modal vocabulary, which enriches the logical expressivity of the inferential relations. This leads also to certain limits in the application of labeling in machine learning. I propose a modal interpretation of implications to express conceptual relations.


Data Scientist - IoT BigData Jobs

#artificialintelligence

DuPont has a rich history of scientific discovery that has enabled countless innovations and today, we're looking for more people, in more places, to collaborate with us to make life the best that it can be. DuPont Pioneer is aggressively building Big Data and Predictive Analytics capabilities in order to deliver improved services to our customers. We seek a strong data scientist with a background in math, statistics, machine learning and scientific computing to join our team. This is a critical position with the potential to make immediate, significant impact on our business. The successful candidate will have an extensive background in statistical computing and machine learning through courses or thesis/dissertation, and proven experience validating models against experimental data.


Optimal Statistical Hypothesis Testing for Social Choice

arXiv.org Artificial Intelligence

We address the following question in this paper: "What are the most robust statistical methods for social choice?'' By leveraging the theory of uniformly least favorable distributions in the Neyman-Pearson framework to finite models and randomized tests, we characterize uniformly most powerful (UMP) tests, which is a well-accepted statistical optimality w.r.t. robustness, for testing whether a given alternative is the winner under Mallows' model and under Condorcet's model, respectively.


Tangles: a new paradigm for clusters and types

arXiv.org Artificial Intelligence

Traditional clustering identifies groups of objects that share certain qualities. Tangles do the converse: they identify groups of qualities that often occur together. They can thereby discover, relate, and structure types: of behaviour, political views, texts, or viruses. If desired, tangles can also be used for direct clustering of objects. They offer a precise, quantitative paradigm suited particularly to fuzzy clusters, since they do not require any `hard' assignments of objects to the clusters they collectively form. This is a draft of the introductory chapter of a book I am preparing on the application of tangles in the empirical sciences. The purpose of posting this draft early is to give authors of tangle application papers a generic reference for the basic guiding principles underlying tangle applications outside mathematics, so that in their own papers they can concentrate on the ideas specific to their particular application rather than having to repeat the generic story each time. The text starts with three separate generic introductions to tangles in the natural sciences, in the social sciences, and in data science including machine learning. It then gives a short informal description of the abstract notion of tangles that encompasses all these potential applications.