Government
Unstructured Data Architect- Federal
BE PART OF BUILDING THE FUTURE. What do NASA and emerging space companies have in common with COVID vaccine R&D teams or with Roblox and the Metaverse? The answer is data, -- all fast moving, fast growing industries rely on data for a competitive edge in their industries. And the most advanced companies are realizing the full data advantage by partnering with Pure Storage. Pure's vision is to redefine the storage experience and empower innovators by simplifying how people consume and interact with data.
Staff Fellow - Senior Data Engineer
INTRODUCTION: The Center for Devices and Radiological Health (CDRH or Center), as the scientific and regulatory component of the U.S. Food and Drug Administration (FDA) charged with facilitating and ensuring medical device innovation, safety, and effectiveness, and advancing regulatory science is now accepting applications for a Staff Fellow (Senior Data Engineer) to serve in the Office of Clinical Evidence and Analysis (OCEA or Office). OCEA oversees the application of modern artificial intelligence tools, including machine learning and deep learning methodologies, that can be evaluated, piloted, and implemented at scale by CDRH to help the Center evaluate clinical evidence and other device-related data to more efficiently conduct regulatory review activities in support of the Center's mission. POSITION SUMMARY: OCEA is seeking a Staff Fellow (Senior Data Engineer) to develop and build the infrastructure and tools for exciting new technologies. The Staff Fellow will design and build Amazon Web Services (AWS) applications and data pipeline services that work at scale to support new data analytics products and also enable development workflows. This will require the automated deployment of microservices and data tools for high-throughput stream processing.
Supervised and Unsupervised Learning for Data Science (Unsupervised and Semi-Supervised Learning): Berry, Michael W., Mohamed, Azlinah, Yap, Bee Wah: 9783030224776: Amazon.com: Books
Professor Michael W. Berry is a Full Professor in the Departments of Electrical Engineering and Computer Science (EECS) and Mathematics at the University of Tennessee, Knoxville. He served as Interim Department Head of Computer Science from January 2004 to June 2007, and as Associate Head in the Department of Electrical Engineering and Computer Science from July 2007 to July 2012. He worked in the Communications Product Division of IBM in Raleigh, NC for about 1 year before accepting a research staff position in the Center for Supercomputing Research and Development at the University of Illinois at Urbana-Champaign. In 1990, he received a PhD in Computer Science from the University of Illinois at Urbana-Champaign. He has published well over 150 peer-refereed journal and conference publications and book chapters.
The Caltech Fish Counting Dataset: A Benchmark for Multiple-Object Tracking and Counting
Kay, Justin, Kulits, Peter, Stathatos, Suzanne, Deng, Siqi, Young, Erik, Beery, Sara, Van Horn, Grant, Perona, Pietro
We present the Caltech Fish Counting Dataset (CFC), a large-scale dataset for detecting, tracking, and counting fish in sonar videos. We identify sonar videos as a rich source of data for advancing low signal-to-noise computer vision applications and tackling domain generalization in multiple-object tracking (MOT) and counting. In comparison to existing MOT and counting datasets, which are largely restricted to videos of people and vehicles in cities, CFC is sourced from a natural-world domain where targets are not easily resolvable and appearance features cannot be easily leveraged for target re-identification. With over half a million annotations in over 1,500 videos sourced from seven different sonar cameras, CFC allows researchers to train MOT and counting algorithms and evaluate generalization performance at unseen test locations. We perform extensive baseline experiments and identify key challenges and opportunities for advancing the state of the art in generalization in MOT and counting.
EVHA: Explainable Vision System for Hardware Testing and Assurance -- An Overview
Hasan, Md Mahfuz Al, Mostafiz, Mohammad Tahsin, Le, Thomas An, Julia, Jake, Vashistha, Nidish, Taheri, Shayan, Asadizanjani, Navid
Due to the ever-growing demands for electronic chips in different sectors the semiconductor companies have been mandated to offshore their manufacturing processes. This unwanted matter has made security and trustworthiness of their fabricated chips concerning and caused creation of hardware attacks. In this condition, different entities in the semiconductor supply chain can act maliciously and execute an attack on the design computing layers, from devices to systems. Our attack is a hardware Trojan that is inserted during mask generation/fabrication in an untrusted foundry. The Trojan leaves a footprint in the fabricated through addition, deletion, or change of design cells. In order to tackle this problem, we propose Explainable Vision System for Hardware Testing and Assurance (EVHA) in this work that can detect the smallest possible change to a design in a low-cost, accurate, and fast manner. The inputs to this system are Scanning Electron Microscopy (SEM) images acquired from the Integrated Circuits (ICs) under examination. The system output is determination of IC status in terms of having any defect and/or hardware Trojan through addition, deletion, or change in the design cells at the cell-level. This article provides an overview on the design, development, implementation, and analysis of our defense system.
A-SFS: Semi-supervised Feature Selection based on Multi-task Self-supervision
Qiu, Zhifeng, Zeng, Wanxin, Liao, Dahua, Gui, Ning
Feature selection is an important process in machine learning. It builds an interpretable and robust model by selecting the features that contribute the most to the prediction target. However, most mature feature selection algorithms, including supervised and semi-supervised, fail to fully exploit the complex potential structure between features. We believe that these structures are very important for the feature selection process, especially when labels are lacking and data is noisy. To this end, we innovatively introduce a deep learning-based self-supervised mechanism into feature selection problems, namely batch-Attention-based Self-supervision Feature Selection(A-SFS). Firstly, a multi-task self-supervised autoencoder is designed to uncover the hidden structure among features with the support of two pretext tasks. Guided by the integrated information from the multi-self-supervised learning model, a batch-attention mechanism is designed to generate feature weights according to batch-based feature selection patterns to alleviate the impacts introduced by a handful of noisy data. This method is compared to 14 major strong benchmarks, including LightGBM and XGBoost. Experimental results show that A-SFS achieves the highest accuracy in most datasets. Furthermore, this design significantly reduces the reliance on labels, with only 1/10 labeled data needed to achieve the same performance as those state of art baselines. Results show that A-SFS is also most robust to the noisy and missing data.
Multi-parametric Analysis for Mixed Integer Linear Programming: An Application to Transmission Planning and Congestion Control
Liu, Jian, Bo, Rui, Wang, Siyuan
Enhancing existing transmission lines is a useful tool to combat transmission congestion and guarantee transmission security with increasing demand and boosting the renewable energy source. This study concerns the selection of lines whose capacity should be expanded and by how much from the perspective of independent system operator (ISO) to minimize the system cost with the consideration of transmission line constraints and electricity generation and demand balance conditions, and incorporating ramp-up and startup ramp rates, shutdown ramp rates, ramp-down rate limits and minimum up and minimum down times. For that purpose, we develop the ISO unit commitment and economic dispatch model and show it as a right-hand side uncertainty multiple parametric analysis for the mixed integer linear programming (MILP) problem. We first relax the binary variable to continuous variables and employ the Lagrange method and Karush-Kuhn-Tucker conditions to obtain optimal solutions (optimal decision variables and objective function) and critical regions associated with active and inactive constraints. Further, we extend the traditional branch and bound method for the large-scale MILP problem by determining the upper bound of the problem at each node, then comparing the difference between the upper and lower bounds and reaching the approximate optimal solution within the decision makers' tolerated error range. In additional, the objective function's first derivative on the parameters of each line is used to inform the selection of lines to ease congestion and maximize social welfare. Finally, the amount of capacity upgrade will be chosen by balancing the cost-reduction rate of the objective function on parameters and the cost of the line upgrade. Our findings are supported by numerical simulation and provide transmission line planners with decision-making guidance.
Finding and Following Optimal Trajectories for an Overactuated Floating Robotic Platform
Bredenbeck, Anton, Vyas, Shubham, Suter, Willem, Zwick, Martin, Borrmann, Dorit, Olivares-Mendez, Miguel, Nüchter, Andreas
The recent increase in yearly spacecraft launches and the high number of planned launches have raised questions about maintaining accessibility to space for all interested parties. A key to sustaining the future of space-flight is the ability to service malfunctioning - and actively remove dysfunctional spacecraft from orbit. Robotic platforms that autonomously perform these tasks are a topic of ongoing research and thus must undergo thorough testing before launch. For representative system-level testing, the European Space Agency (ESA) uses, among other things, the Orbital Robotics and GNC Lab (ORGL), a flat-floor facility where air-bearing based platforms exhibit free-floating behavior in three Degrees of Freedom (DoF). This work introduces a representative simulation of a free-floating platform in the testing environment and a software framework for controller development. Finally, this work proposes a controller within that framework for finding and following optimal trajectories between arbitrary states, which is evaluated in simulation and reality.
FactGraph: Evaluating Factuality in Summarization with Semantic Graph Representations
Ribeiro, Leonardo F. R., Liu, Mengwen, Gurevych, Iryna, Dreyer, Markus, Bansal, Mohit
Despite recent improvements in abstractive summarization, most current approaches generate summaries that are not factually consistent with the source document, severely restricting their trust and usage in real-world applications. Recent works have shown promising improvements in factuality error identification using text or dependency arc entailments; however, they do not consider the entire semantic graph simultaneously. To this end, we propose FactGraph, a method that decomposes the document and the summary into structured meaning representations (MR), which are more suitable for factuality evaluation. MRs describe core semantic concepts and their relations, aggregating the main content in both document and summary in a canonical form, and reducing data sparsity. FactGraph encodes such graphs using a graph encoder augmented with structure-aware adapters to capture interactions among the concepts based on the graph connectivity, along with text representations using an adapter-based text encoder. Experiments on different benchmarks for evaluating factuality show that FactGraph outperforms previous approaches by up to 15%. Furthermore, FactGraph improves performance on identifying content verifiability errors and better captures subsentence-level factual inconsistencies.
Investigation of a Data Split Strategy Involving the Time Axis in Adverse Event Prediction Using Machine Learning
Morita, Katsuhisa, Mizuno, Tadahaya, Kusuhara, Hiroyuki
Adverse events are a serious issue in drug development and many prediction methods using machine learning have been developed. The random split cross-validation is the de facto standard for model building and evaluation in machine learning, but care should be taken in adverse event prediction because this approach does not match to the real-world situation. The time split, which uses the time axis, is considered suitable for real-world prediction. However, the differences in model performance obtained using the time and random splits are not clear due to the lack of the comparable studies. To understand the differences, we compared the model performance between the time and random splits using nine types of compound information as input, eight adverse events as targets, and six machine learning algorithms. The random split showed higher area under the curve values than did the time split for six of eight targets. The chemical spaces of the training and test datasets of the time split were similar, suggesting that the concept of applicability domain is insufficient to explain the differences derived from the splitting. The area under the curve differences were smaller for the protein interaction than for the other datasets. Subsequent detailed analyses suggested the danger of confounding in the use of knowledge-based information in the time split. These findings indicate the importance of understanding the differences between the time and random splits in adverse event prediction and strongly suggest that appropriate use of the splitting strategies and interpretation of results are necessary for the real-world prediction of adverse events. We provide analysis code and datasets used in the present study (https://github.com/mizuno-group/AE_prediction).