Government
Agreeing to Disagree: Annotating Offensive Language Datasets with Annotators' Disagreement
Leonardelli, Elisa, Menini, Stefano, Aprosio, Alessio Palmero, Guerini, Marco, Tonelli, Sara
Since state-of-the-art approaches to offensive language detection rely on supervised learning, it is crucial to quickly adapt them to the continuously evolving scenario of social media. While several approaches have been proposed to tackle the problem from an algorithmic perspective, so to reduce the need for annotated data, less attention has been paid to the quality of these data. Following a trend that has emerged recently, we focus on the level of agreement among annotators while selecting data to create offensive language datasets, a task involving a high level of subjectivity. Our study comprises the creation of three novel datasets of English tweets covering different topics and having five crowd-sourced judgments each. We also present an extensive set of experiments showing that selecting training and test data according to different levels of annotators' agreement has a strong effect on classifiers performance and robustness. Our findings are further validated in cross-domain experiments and studied using a popular benchmark dataset. We show that such hard cases, where low agreement is present, are not necessarily due to poor-quality annotation and we advocate for a higher presence of ambiguous cases in future datasets, particularly in test sets, to better account for the different points of view expressed online.
Heterogeneous Distributed Lag Models to Estimate Personalized Effects of Maternal Exposures to Air Pollution
Mork, Daniel, Kioumourtzoglou, Marianthi-Anna, Weisskopf, Marc, Coull, Brent A, Wilson, Ander
Children's health studies support an association between maternal environmental exposures and children's birth and health outcomes. A common goal in such studies is to identify critical windows of susceptibility -- periods during gestation with increased association between maternal exposures and a future outcome. The associations and timings of critical windows are likely heterogeneous across different levels of individual, family, and neighborhood characteristics. However, the few studies that have considered effect modification were limited to a few pre-specified subgroups. We propose a statistical learning method to estimate critical windows at the individual level and identify important characteristics that induce heterogeneity. The proposed approach uses distributed lag models (DLMs) modified by Bayesian additive regression trees to account for effect heterogeneity based on a potentially high-dimensional set of modifying factors. We show in a simulation study that our model can identify both critical windows and modifiers responsible for DLM heterogeneity. We estimate the relationship between weekly exposures to fine particulate matter during gestation and birth weight in an administrative Colorado birth cohort. We identify maternal body mass index (BMI), age, Hispanic designation, and education as modifiers of the distributed lag effects and find non-Hispanics with increased BMI to be a susceptible population.
slimTrain -- A Stochastic Approximation Method for Training Separable Deep Neural Networks
Newman, Elizabeth, Chung, Julianne, Chung, Matthias, Ruthotto, Lars
Deep neural networks (DNNs) have shown their success as high-dimensional function approximators in many applications; however, training DNNs can be challenging in general. DNN training is commonly phrased as a stochastic optimization problem whose challenges include non-convexity, non-smoothness, insufficient regularization, and complicated data distributions. Hence, the performance of DNNs on a given task depends crucially on tuning hyperparameters, especially learning rates and regularization parameters. In the absence of theoretical guidelines or prior experience on similar tasks, this requires solving many training problems, which can be time-consuming and demanding on computational resources. This can limit the applicability of DNNs to problems with non-standard, complex, and scarce datasets, e.g., those arising in many scientific applications. To remedy the challenges of DNN training, we propose slimTrain, a stochastic optimization method for training DNNs with reduced sensitivity to the choice hyperparameters and fast initial convergence. The central idea of slimTrain is to exploit the separability inherent in many DNN architectures; that is, we separate the DNN into a nonlinear feature extractor followed by a linear model. This separability allows us to leverage recent advances made for solving large-scale, linear, ill-posed inverse problems. Crucially, for the linear weights, slimTrain does not require a learning rate and automatically adapts the regularization parameter. Since our method operates on mini-batches, its computational overhead per iteration is modest. In our numerical experiments, slimTrain outperforms existing DNN training methods with the recommended hyperparameter settings and reduces the sensitivity of DNN training to the remaining hyperparameters.
Learning from Small Samples: Transformation-Invariant SVMs with Composition and Locality at Multiple Scales
Liu, Tao, Kumar, P. R., Liu, Xi
Motivated by the problem of learning when the number of training samples is small, this paper shows how to incorporate into support-vector machines (SVMs) those properties that have made convolutional neural networks (CNNs) successful. Particularly important is the ability to incorporate domain knowledge of invariances, e.g., translational invariance of images. Kernels based on the \textit{minimum} distance over a group of transformations, which corresponds to defining similarity as the \textit{best} over the possible transformations, are not generally positive definite. Perhaps it is for this reason that they have neither previously been experimentally tested for their performance nor studied theoretically. Instead, previous attempts have employed kernels based on the \textit{average} distance over a group of transformations, which are trivially positive definite, but which generally yield both poor margins as well as poor performance, as we show. We address this lacuna and show that positive definiteness indeed holds \textit{with high probability} for kernels based on the minimum distance in the small training sample set regime of interest, and that they do yield the best results in that regime. Another important property of CNNs is their ability to incorporate local features at multiple spatial scales, e.g., through max pooling. A third important property is their ability to provide the benefits of composition through the architecture of multiple layers. We show how these additional properties can also be embedded into SVMs. We verify through experiments on widely available image sets that the resulting SVMs do provide superior accuracy in comparison to well-established deep neural network (DNN) benchmarks for small sample sizes.
ONR at 75: Virtual Anniversary Event to Highlight Future of Naval Power
On Thursday, Sept. 30, from 10:30 a.m. to 11:50 a.m., senior naval and congressional leaders will participate in a special Office of Naval Research (ONR)-sponsored virtual event to discuss "The Future of Warfare." Held in honor of ONR's 75th anniversary, the event is titled "ONR at 75: Reimagine Naval Power." It will feature remarks from the Chief of Naval Operations Adm. Michael Gilday, and from the Acting Assistant Secretary of the Navy for Research, Development and Acquisition, the Hon. A panel discussion will follow, featuring two members of the U.S. Congress-Rep. The panel, titled "The Future of Warfare," will be led by Chief of Naval Research Rear Adm. Lorin C. Selby.
Misinformation Is About to Get So Much Worse
For years now, artificial intelligence has been hailed as both a savior and a destroyer. The technology really can make our lives easier, letting us summon our phones with a "Hey, Siri" and (more importantly) assisting doctors on the operating table. But as any science-fiction reader knows, AI is not an unmitigated good: It can be prone to the same racial biases as humans are, and, as is the case with self-driving cars, it can be forced to make murky split-second decisions that determine who lives and who dies. Like it or not, AI is only going to become an even more omnipresent force: We're in a "watershed moment" for the technology, says Eric Schmidt, the former Google CEO. Schmidt is a longtime fixture in a tech industry that seems to constantly be in a state of upheaval. He was the first software manager at Sun Microsystems, in the 1980s, and the CEO of the former software giant Novell in the '90s. He joined Google as CEO in 2001, then was the company's executive chairman from 2011 until 2017. Since leaving Google, Schmidt has made AI his focus: In 2018, he wrote in The Atlantic about the need to prepare for the AI boom, along with his co-authors Henry Kissinger, the former secretary of state, and the MIT dean Daniel Huttenlocher. The trio have followed up that story with The Age of AI, a book about how AI will transform how we experience the world, coming out in November.
Movers and Shakers news roundup
The end of June saw the announcement that Peter Thomas would be stepping into the chief clinical information officer (CCIO) role at Moorfields Eye Hospital NHS Foundation Trust in August. The consultant joined the trust in 2017 and took a special interest in machine learning and artificial intelligence, pioneering the hospital's use of digital medicine and telemedicine. In his capacity as CCIO, Thomas will be in charge of raising awareness of clinical informatics as an important element in safe, high-quality patient care. He said: "I am delighted to be offered the role of chief clinical information officer at Moorfields and I hope to use this opportunity to use digital medicine in innovative ways to help our patients receive the best care possible." University Hospitals of Leicester NHS Trust has announced Richard Mitchell will take up the position as the trust's chief executive from Autumn 2021.
When Using AI in Enterprises, Balancing Innovation and Privacy Is Critical
While the U.S. is making strides in the advancement of AI use cases across industries, we have a long way to go before AI technologies are commonplace and truly ingrained in our daily life. What are the missing pieces? Better data access and improved data sharing. As our ability to address point applications and solutions with AI technology matures, we will need a greater ability to share data and insights while being able to draw conclusions across problem domains. Cooperation between individuals from government, research, higher education and the private sector to make greater data sharing feasible will drive acceleration of new use cases while balancing the need for data privacy.
News - Research in Germany
Which issues do the parties highlight with regard to artificial intelligence? The electoral platforms mention AI mainly in connection with the economy, foreign policy and the area of education and research. The proposals are mostly framed in the context of the competitiveness of German and European companies. The need for better international cooperation on AI and the issue of whether the technology should be used in military intelligence are mentioned with similar frequency. Another important topic is research funding. In general, positive paradigms outweigh the statements with neutral or negative connotations.
Left-wing activist Michael Moore says US defense should focus on climate, white supremacists, Covid vaccines
Left-wing activist Michael Moore claimed Sunday that U.S. defense policy and spending should be refocused from military involvement in other countries to instead fight climate change, White supremacy and the coronavirus pandemic. Appearing on MSNBC, Moore suggested U.S. armed forces weren't "the good guys" because of their military involvement in countries like Syria and Somalia, and instead claimed he wanted to be known for doing "the good things," like building wells in poor villages rather than provide funding to Israel's Iron Dome defense system. "As we speak tonight the U.S. still has 2,500 troops in Iraq. Last week the U.S. military admitted that a drone strike in Kabul that killed ten innocent civilians was an American drone strike. Seven of them were kids. We talk about ending the war but we're still going to carry on with drones, we're still going to carry on with these, quote, over-the-horizon operations. You know, we are still at war," host Mehdi Hasan said after playing a clip of President Joe Biden touting the U.S. not being at war for the first time in 20 years.