Scientific Discovery
Applications of Hypothesis Testing part3(Advanced Statistics)
Abstract: In many scenarios such as genome-wide association studies where dependences between variables commonly exist, it is often of interest to infer the interaction effects in the model. However, testing pairwise interactions among millions of variables in complex and high-dimensional data suffers from low statistical power and huge computational cost. To address these challenges, we propose a two-stage testing procedure with false discovery rate (FDR) control, which is known as a less conservative multiple-testing correction. Theoretically, the difficulty in the FDR control dues to the data dependence among test statistics in two stages, and the fact that the number of hypothesis tests conducted in the second stage depends on the screening result in the first stage. By using the Cramรฉr type moderate deviation technique, we show that our procedure controls FDR at the desired level asymptotically in the generalized linear model (GLM), where the model is allowed to be misspecified.
Computer scientist, Data scientist or similar with a focus on knowledge management (f/m/x) - Data Discovery for Anonymised Health Data
The focus of the DLR Institute for Data Science in Jena is to find solutions for the major challenges of the digitalisation age. The research focuses on the areas of data extraction and mobilisation, data management and preparation, and data analysis and intelligence. The position is part of the BMBF project Avatar (anonymisation of personal health data by creating virtual avatars). Topics include, in particular, the semantic modelling of relevant metadata and data discovery. The overall goal of the project is providing anonymised health data for both academic and commercial research.
Mediamorphosis: How AI is enabling a new paradigm for work and play
Did you miss a session from MetaBeat 2022? Head over to the on-demand library for all of our featured sessions here. Text-to-image AI systems such as DALL-E 2, Imagen and Midjourney are growing in popularity and capability right now, offering creators a revolutionary new way to produce content. Generating images from text prompts is a radical new approach to art-making and creative expression. But it also gives us the first glimpse of a fundamental shift in how we can better communicate and collaborate with our machines.
ZeroC: A Neuro-Symbolic Model for Zero-shot Concept Recognition and Acquisition at Inference Time
Wu, Tailin, Tjandrasuwita, Megan, Wu, Zhengxuan, Yang, Xuelin, Liu, Kevin, Sosiฤ, Rok, Leskovec, Jure
Humans have the remarkable ability to recognize and acquire novel visual concepts in a zero-shot manner. Given a high-level, symbolic description of a novel concept in terms of previously learned visual concepts and their relations, humans can recognize novel concepts without seeing any examples. Moreover, they can acquire new concepts by parsing and communicating symbolic structures using learned visual concepts and relations. Endowing these capabilities in machines is pivotal in improving their generalization capability at inference time. In this work, we introduce Zero-shot Concept Recognition and Acquisition (ZeroC), a neuro-symbolic architecture that can recognize and acquire novel concepts in a zero-shot way. ZeroC represents concepts as graphs of constituent concept models (as nodes) and their relations (as edges). To allow inference time composition, we employ energy-based models (EBMs) to model concepts and relations. We design ZeroC architecture so that it allows a one-to-one mapping between a symbolic graph structure of a concept and its corresponding EBM, which for the first time, allows acquiring new concepts, communicating its graph structure, and applying it to classification and detection tasks (even across domains) at inference time. We introduce algorithms for learning and inference with ZeroC. We evaluate ZeroC on a challenging grid-world dataset which is designed to probe zero-shot concept recognition and acquisition, and demonstrate its capability.
Riemannian geometry as a unifying theory for robot motion learning and control
Jaquier, Noรฉmie, Asfour, Tamim
Riemannian geometry is a mathematical field which has been the cornerstone of revolutionary scientific discoveries such as the theory of general relativity. Despite early uses in robot design and recent applications for exploiting data with specific geometries, it mostly remains overlooked in robotics. With this blue sky paper, we argue that Riemannian geometry provides the most suitable tools to analyze and generate well-coordinated, energy-efficient motions of robots with many degrees of freedom. Via preliminary solutions and novel research directions, we discuss how Riemannian geometry may be leveraged to design and combine physically-meaningful synergies for robotics, and how this theory also opens the door to coupling motion synergies with perceptual inputs.
[100%OFF] Scanning & Discovery Techniques For Penstesters
Udemy is the biggest website in the world that offer courses in many categories, all the skills that you would be looking for are offered in Udemy, including languages, design, marketing and a lot of other categories, so when you ever want to buy a courses and pay for a new skills, Udemy would be the best forum for you. You can find payment courses, 100 free courses From Udemy and coupons also, more than 12 categories are offered, and that what makes sure you will find the domain and the skill you are looking for. Our duty is to search for 100 off courses and free coupons. Nmap is an indispensable tool that all techies should know well. It is used by all good ethical hackers, penetration testers, systems administrators, and anyone in fact who wants to discovery more about the security of a network and its hosts.
The Secret Microscope That Sparked a Scientific Revolution
While he was examining algae from a nearby lake through his homemade microscope, a creature "with green and very glittering little scales," which he estimated to be a thousand times smaller than a mite, had darted across his vision. Two years later, on October 9, 1676, he followed up with another report so extraordinary that microbiologists today refer to it simply as "Letter 18": Van Leeuwenhoek (lay-u-when-hoke) had looked everywhere and found what he called animalcules (Latin for "little animals") in everything. He found them in the bellies of other animals, his food, his own mouth, and other people's mouths. When he noticed a set of remarkably rancid teeth, he asked the owner for a sample of his plaque, put it beneath his lens, and witnessed "an inconceivably great number of little animalcules" moving "so nimbly among one another, that the whole stuff seemed alive." After a particularly uncomfortable evening, which he blamed on a fatty meal of hot smoked beef, he examined his own stool beneath his lens and saw animalcules that were "somewhat longer than broad, and their belly, which was flat-like, furnished with sundry little paws"--a clear description of what we now know as the parasite giardia. With his observations of these fast, fat, and sundry-pawed creatures, Van Leeuwenhoek became the first person to ever see a microorganism--a discovery of almost incalculable significance to human health and our understanding of life on this planet.
Active Few-Shot Classification: a New Paradigm for Data-Scarce Learning Settings
Abdali, Aymane, Gripon, Vincent, Drumetz, Lucas, Boguslawski, Bartosz
We consider a novel formulation of the problem of Active Few-Shot Classification (AFSC) where the objective is to classify a small, initially unlabeled, dataset given a very restrained labeling budget. This problem can be seen as a rival paradigm to classical Transductive Few-Shot Classification (TFSC), as both these approaches are applicable in similar conditions. We first propose a methodology that combines statistical inference, and an original two-tier active learning strategy that fits well into this framework. We then adapt several standard vision benchmarks from the field of TFSC. Our experiments show the potential benefits of AFSC can be substantial, with gains in average weighted accuracy of up to 10% compared to state-of-the-art TFSC methods for the same labeling budget. We believe this new paradigm could lead to new developments and standards in data-scarce learning settings.
Trust Calibration as a Function of the Evolution of Uncertainty in Knowledge Generation: A Survey
User trust is a crucial consideration in designing robust visual analytics systems that can guide users to reasonably sound conclusions despite inevitable biases and other uncertainties introduced by the human, the machine, and the data sources which paint the canvas upon which knowledge emerges. A multitude of factors emerge upon studied consideration which introduce considerable complexity and exacerbate our understanding of how trust relationships evolve in visual analytics systems, much as they do in intelligent sociotechnical systems. A visual analytics system, however, does not by its nature provoke exactly the same phenomena as its simpler cousins, nor are the phenomena necessarily of the same exact kind. Regardless, both application domains present the same root causes from which the need for trustworthiness arises: Uncertainty and the assumption of risk. In addition, visual analytics systems, even more than the intelligent systems which (traditionally) tend to be closed to direct human input and direction during processing, are influenced by a multitude of cognitive biases that further exacerbate an accounting of the uncertainties that may afflict the user's confidence, and ultimately trust in the system. In this article we argue that accounting for the propagation of uncertainty from data sources all the way through extraction of information and hypothesis testing is necessary to understand how user trust in a visual analytics system evolves over its lifecycle, and that the analyst's selection of visualization parameters affords us a simple means to capture the interactions between uncertainty and cognitive bias as a function of the attributes of the search tasks the analyst executes while evaluating explanations. We sample a broad cross-section of the literature from visual analytics, human cognitive theory, and uncertainty, and attempt to synthesize a useful perspective.
A Case for Dataset Specific Profiling
Data-driven science is an emerging paradigm where scientific discoveries depend on the execution of computational AI models against rich, discipline-specific datasets. With modern machine learning frameworks, anyone can develop and execute computational models that reveal concepts hidden in the data that could enable scientific applications. For important and widely used datasets, computing the performance of every computational model that can run against a dataset is cost prohibitive in terms of cloud resources. Benchmarking approaches used in practice use representative datasets to infer performance without actually executing models. While practicable, these approaches limit extensive dataset profiling to a few datasets and introduce bias that favors models suited for representative datasets. As a result, each dataset's unique characteristics are left unexplored and subpar models are selected based on inference from generalized datasets. This necessitates a new paradigm that introduces dataset profiling into the model selection process. To demonstrate the need for dataset-specific profiling, we answer two questions:(1) Can scientific datasets significantly permute the rank order of computational models compared to widely used representative datasets? (2) If so, could lightweight model execution improve benchmarking accuracy? Taken together, the answers to these questions lay the foundation for a new dataset-aware benchmarking paradigm.