Government
What Do Killer Robots Dream Of?
The cinematic depictions are pretty clear. From I, Robot to The Terminator series to The Matrix, humans, either wittingly or unwittingly, manage to clash with the machinery that had previously served them. It's a narrative that, like some of our best mythos, puts us at the center of the action and more often than not shows the supremacy of human ingenuity under pressure. Because it's Hollywood, and Hollywood specializes in fictions. Well, at least part of it is a fiction.
Bayesian Few-Shot Classification with One-vs-Each P\'olya-Gamma Augmented Gaussian Processes
Few-shot classification (FSC), the task of adapting a classifier to unseen classes given a small labeled dataset, is an important step on the path toward human-like machine learning. Bayesian methods are well-suited to tackling the fundamental issue of overfitting in the few-shot scenario because they allow practitioners to specify prior beliefs and update those beliefs in light of observed data. Contemporary approaches to Bayesian few-shot classification maintain a posterior distribution over model parameters, which is slow and requires storage that scales with model size. Instead, we propose a Gaussian process classifier based on a novel combination of P\'olya-gamma augmentation and the one-vs-each softmax approximation that allows us to efficiently marginalize over functions rather than model parameters. We demonstrate improved accuracy and uncertainty quantification on both standard few-shot classification benchmarks and few-shot domain transfer tasks.
An Empirical Characterization of Fair Machine Learning For Clinical Risk Prediction
Pfohl, Stephen R., Foryciarz, Agata, Shah, Nigam H.
The use of machine learning to guide clinical decision making has the potential to worsen existing health disparities. Several recent works frame the problem as that of algorithmic fairness, a framework that has attracted considerable attention and criticism. However, the appropriateness of this framework is unclear due to both ethical as well as technical considerations, the latter of which include trade-offs between measures of fairness and model performance that are not well-understood for predictive models of clinical outcomes. To inform the ongoing debate, we conduct an empirical study to characterize the impact of penalizing group fairness violations on an array of measures of model performance and group fairness. We repeat the analyses across multiple observational healthcare databases, clinical outcomes, and sensitive attributes. We find that procedures that penalize differences between the distributions of predictions across groups induce nearly-universal degradation of multiple performance metrics within groups. On examining the secondary impact of these procedures, we observe heterogeneity of the effect of these procedures on measures of fairness in calibration and ranking across experimental conditions. Beyond the reported trade-offs, we emphasize that analyses of algorithmic fairness in healthcare lack the contextual grounding and causal awareness necessary to reason about the mechanisms that lead to health disparities, as well as about the potential of algorithmic fairness methods to counteract those mechanisms. In light of these limitations, we encourage researchers building predictive models for clinical use to step outside the algorithmic fairness frame and engage critically with the broader sociotechnical context surrounding the use of machine learning in healthcare.
ProtTrans: Towards Cracking the Language of Life's Code Through Self-Supervised Deep Learning and High Performance Computing
Elnaggar, Ahmed, Heinzinger, Michael, Dallago, Christian, Rihawi, Ghalia, Wang, Yu, Jones, Llion, Gibbs, Tom, Feher, Tamas, Angerer, Christoph, Steinegger, Martin, Bhowmik, Debsindhu, Rost, Burkhard
Computational biology and bioinformatics provide vast data gold-mines from protein sequences, ideal for Language Models (LMs) taken from Natural Language Processing (NLP). These LMs reach for new prediction frontiers at low inference costs. Here, we trained two auto-regressive language models (Transformer-XL, XLNet) and two auto-encoder models (Bert, Albert) on data from UniRef and BFD containing up to 393 billion amino acids (words) from 2.1 billion protein sequences (22- and 112-times the entire English Wikipedia). The LMs were trained on the Summit supercomputer at Oak Ridge National Laboratory (ORNL), using 936 nodes (total 5616 GPUs) and one TPU Pod (V3-512 or V3-1024). We validated the advantage of up-scaling LMs to larger models supported by bigger data by predicting secondary structure (3-states: Q3=76-84, 8-states: Q8=65-73), sub-cellular localization for 10 cellular compartments (Q10=74) and whether a protein is membrane-bound or water-soluble (Q2=89). Dimensionality reduction revealed that the LM-embeddings from unlabeled data (only protein sequences) captured important biophysical properties governing protein shape. This implied learning some of the grammar of the language of life realized in protein sequences. The successful up-scaling of protein LMs through HPC to larger data sets slightly reduced the gap between models trained on evolutionary information and LMs. The official GitHub repository: https://github.com/agemagician/ProtTrans
Robust Causal Inference Under Covariate Shift via Worst-Case Subpopulation Treatment Effects
Jeong, Sookyo, Namkoong, Hongseok
We propose the worst-case treatment effect (WTE) across all subpopulations of a given size, a conservative notion of topline treatment effect. Compared to the average treatment effect (ATE), whose validity relies on the covariate distribution of collected data, WTE is robust to unanticipated covariate shifts, and positive findings guarantee uniformly valid treatment effects over subpopulations. We develop a semiparametrically efficient estimator for the WTE, leveraging machine learning-based estimates of the heterogeneous treatment effect and propensity score. By virtue of satisfying a key (Neyman) orthogonality property, our estimator enjoys central limit behavior---oracle rates with true nuisance parameters---even when estimates of nuisance parameters converge at slower rates. For both randomized trials and observational studies, we establish a semiparametric efficiency bound, proving that our estimator achieves the optimal asymptotic variance. On real datasets where robustness to covariate shift is of core concern, we illustrate the non-robustness of ATE under even mild distributional shift, and demonstrate that the WTE guards against brittle findings that are invalidated by unanticipated covariate shifts.
The Hidden Agendas of Masks, Distancing, and Tracing - Vaxxter
The maker of the N95 respirator mask designed it for the mining and construction industry. It filters out 95 percent of the visible airborne dust particles, keeping them from entering the lungs. It has vents so the wearers can exhale. Demolition crews use them, concrete laborers use them, so do workers who sand wood floors, gypsum board walls, and plaster ceilings. Workers in other trades, such as electrical, plumbing, rough carpentry, and wood finishers rarely wear masks.
What No One Will Tell You About Robots
Human fascination with robots has long been fused with fear. The first widespread use of the term came a century ago in a Czech play about robots manufactured to serve and work for people. The bots turn on their masters. That plot has played out in fiction countless times since. Meanwhile, the real world has created ever more advanced versions of mechanical servants.
Differences Between Europe and the United States on AI/Digital Policy: Comment Response to Roundtable Discussion on AI - Pierre-Antoine Gourraud, 2020
For AI policy, there are significant differences between Europe and the United States. The General Data Protection Regulation, which applies not only to EU companies but also to all American companies with European customers, is more protective than health insurance portability and accountability act for individual health data. Its Article 22 stipulates that citizens cannot be submitted to medical decisions generated by an automated source. For the creation and implementation of national health databases, European companies have an advantage over the United States because of their small sizes, single-payer systems, and existing national cohorts. For instance, France is in the process of developing a national health data platform (Health Data Hub [HDH]), as part of the Healthcare Law of July 14, 2019.1 It has its origins in the report presented by Cedric Villani to the French government in March 2018.2
What No One Will Tell You About Robots
Human fascination with robots has long been fused with fear. The first widespread use of the term came a century ago in a Czech play about robots manufactured to serve and work for people. The bots turn on their masters. That plot has played out in fiction countless times since. Meanwhile, the real world has created ever more advanced versions of mechanical servants.
Contact Tracing with AI Poses Personal Privacy Tradeoffs - AI Trends
Efforts in contact tracing to try to control the spread of the Covid-19 virus had been going on before Google and Apple in early April announced their partnership on contact tracing technology. However, the two tech giants have proposed a way to share data while keeping user privacy central to the design. Recent news out of Singapore may point the way to how this is likely to go, pointing in the direction of the surveillance state. The pursuit of effective contact tracing embodies a confluence of issues around AI and surveillance, data privacy and public safety, and the roles of government and industry. Most contact tracing apps installed on smartphones use Bluetooth radio technology to record when other phones with the same app are detected nearby When a user shows symptoms or tests positive for Covid-19, alerts can be sent to all those in proximity over the previous week or two, along with suggestions for how to respond.