Pacific Ocean
'I'm afraid': critics of anti-cheating technology for students hit by lawsuits
In 2020, a Canadian university employee named Ian Linkletter became increasingly alarmed by a new kind of technology that was exploding in use with the pandemic. It was meant to detect cheating by college and high-school students taking tests at home, and claimed to work by watching students' movements and analyzing sounds around them through their webcams and microphones to automatically flag suspicious behavior. So Linkletter accessed a section of the website of one of the anti-cheating companies, named Proctorio, intended only for instructors and administrators. He shared what he found on social media. Now Linkletter, who became a prominent critic of the technology, has been sued by the company. But he is not the only one.
Physically Constrained Generative Adversarial Networks for Improving Precipitation Fields from Earth System Models
Hess, Philipp, Drรผke, Markus, Petri, Stefan, Strnad, Felix M., Boers, Niklas
Precipitation results from complex processes across many scales, making its accurate simulation in Earth system models (ESMs) challenging. Existing post-processing methods can improve ESM simulations locally, but cannot correct errors in modelled spatial patterns. Here we propose a framework based on physically constrained generative adversarial networks (GANs) to improve local distributions and spatial structure simultaneously. We apply our approach to the computationally efficient ESM CM2Mc-LPJmL. Our method outperforms existing ones in correcting local distributions, and leads to strongly improved spatial patterns especially regarding the intermittency of daily precipitation. Notably, a double-peaked Intertropical Convergence Zone, a common problem in ESMs, is removed. Enforcing a physical constraint to preserve global precipitation sums, the GAN can generalize to future climate scenarios unseen during training. Feature attribution shows that the GAN identifies regions where the ESM exhibits strong biases. Our method constitutes a general framework for correcting ESM variables and enables realistic simulations at a fraction of the computational costs.
A differentiable short-time Fourier transform with respect to the window length
Leiber, Maxime, Barrau, Axel, Marnissi, Yosra, Abboud, Dany
In this paper, we revisit the use of spectrograms in neural networks, by making the window length a continuous parameter optimizable by gradient descent instead of an empirically tuned integer-valued hyperparameter. The contribution is mostly theoretical at this point, but plugging the modified STFT into any existing neural network is straightforward. We first define a differentiable version of the STFT in the case where local bins centers are fixed and independent of the window length parameter. We then discuss the more difficult case where the window length affects the position and number of bins. We illustrate the benefits of this new tool on an estimation and a classification problems, showing it can be of interest not only to neural networks but to any STFT-based signal processing algorithm.
AtmoDist: Self-supervised Representation Learning for Atmospheric Dynamics
Hoffmann, Sebastian, Lessig, Christian
Representation learning has proven to be a powerful methodology in a wide variety of machine learning applications. For atmospheric dynamics, however, it has so far not been considered, arguably due to the lack of large-scale, labeled datasets that could be used for training. In this work, we show that the difficulty is benign and introduce a self-supervised learning task that defines a categorical loss for a wide variety of unlabeled atmospheric datasets. Specifically, we train a neural network on the simple yet intricate task of predicting the temporal distance between atmospheric fields from distinct but nearby times. We demonstrate that training with this task on ERA5 reanalysis leads to internal representations capturing intrinsic aspects of atmospheric dynamics. We do so by introducing a data-driven distance metric for atmospheric states. When employed as a loss function in other machine learning applications, this Atmodist distance leads to improved results compared to the classical $\ell_2$-loss. For example, for downscaling one obtains higher resolution fields that match the true statistics more closely than previous approaches and for the interpolation of missing or occluded data the AtmoDist distance leads to results that contain more realistic fine scale features. Since it is derived from observational data, AtmoDist also provides a novel perspective on atmospheric predictability.
The Development of a Labelled te reo M\=aori-English Bilingual Database for Language Technology
James, Jesin, Shields, Isabella, Yogarajan, Vithya, Keegan, Peter J., Watson, Catherine, Jones, Peter-Lucas, Mahelona, Keoni
Te reo M\=aori (referred to as M\=aori), New Zealand's indigenous language, is under-resourced in language technology. M\=aori speakers are bilingual, where M\=aori is code-switched with English. Unfortunately, there are minimal resources available for M\=aori language technology, language detection and code-switch detection between M\=aori-English pair. Both English and M\=aori use Roman-derived orthography making rule-based systems for detecting language and code-switching restrictive. Most M\=aori language detection is done manually by language experts. This research builds a M\=aori-English bilingual database of 66,016,807 words with word-level language annotation. The New Zealand Parliament Hansard debates reports were used to build the database. The language labels are assigned using language-specific rules and expert manual annotations. Words with the same spelling, but different meanings, exist for M\=aori and English. These words could not be categorised as M\=aori or English based on word-level language rules. Hence, manual annotations were necessary. An analysis reporting the various aspects of the database such as metadata, year-wise analysis, frequently occurring words, sentence length and N-grams is also reported. The database developed here is a valuable tool for future language and speech technology development for Aotearoa New Zealand. The methodology followed to label the database can also be followed by other low-resourced language pairs.
Carefully choose the baseline: Lessons learned from applying XAI attribution methods for regression tasks in geoscience
Mamalakis, Antonios, Barnes, Elizabeth A., Ebert-Uphoff, Imme
Methods of eXplainable Artificial Intelligence (XAI) are used in geoscientific applications to gain insights into the decision-making strategy of Neural Networks (NNs) highlighting which features in the input contribute the most to a NN prediction. Here, we discuss our lesson learned that the task of attributing a prediction to the input does not have a single solution. Instead, the attribution results and their interpretation depend greatly on the considered baseline (sometimes referred to as reference point) that the XAI method utilizes; a fact that has been overlooked so far in the literature. This baseline can be chosen by the user or it is set by construction in the method s algorithm, often without the user being aware of that choice. We highlight that different baselines can lead to different insights for different science questions and, thus, should be chosen accordingly. To illustrate the impact of the baseline, we use a large ensemble of historical and future climate simulations forced with the SSP3-7.0 scenario and train a fully connected NN to predict the ensemble- and global-mean temperature (i.e., the forced global warming signal) given an annual temperature map from an individual ensemble member. We then use various XAI methods and different baselines to attribute the network predictions to the input. We show that attributions differ substantially when considering different baselines, as they correspond to answering different science questions. We conclude by discussing some important implications and considerations about the use of baselines in XAI research.
Learning-based estimation of in-situ wind speed from underwater acoustics
Zambra, Matteo, Cazau, Dorian, Farrugia, Nicolas, Gensse, Alexandre, Pensieri, Sara, Bozzano, Roberto, Fablet, Ronan
Wind speed retrieval at sea surface is of primary importance for scientific and operational applications. Besides weather models, in-situ measurements and remote sensing technologies, especially satellite sensors, provide complementary means to monitor wind speed. As sea surface winds produce sounds that propagate underwater, underwater acoustics recordings can also deliver fine-grained wind-related information. Whereas model-driven schemes, especially data assimilation approaches, are the state-of-the-art schemes to address inverse problems in geoscience, machine learning techniques become more and more appealing to fully exploit the potential of observation datasets. Here, we introduce a deep learning approach for the retrieval of wind speed time series from underwater acoustics possibly complemented by other data sources such as weather model reanalyses. Our approach bridges data assimilation and learning-based frameworks to benefit both from prior physical knowledge and computational efficiency. Numerical experiments on real data demonstrate that we outperform the state-of-the-art data-driven methods with a relative gain up to 16% in terms of RMSE. Interestingly, these results support the relevance of the time dynamics of underwater acoustic data to better inform the time evolution of wind speed. They also show that multimodal data, here underwater acoustics data combined with ECMWF reanalysis data, may further improve the reconstruction performance, including the robustness with respect to missing underwater acoustics data.
OK Google, get me a Coke: AI giant demos soda-fetching robots
MOUNTAIN VIEW, Calif., Aug 16 (Reuters) - Alphabet Inc's (GOOGL.O) Google is combining the eyes and arms of physical robots with the knowledge and conversation skills of virtual chatbots to help its employees fetch soda and chips from breakrooms with ease. The mechanical waiters, shown in action to reporters last week, embody an artificial intelligence breakthrough that paves the way for multipurpose robots as easy to control as ones that perform single, structured tasks such as vacuuming or standing guard. Google robots are not ready for sale. They perform only a few dozen simple actions, and the company has not yet embedded them with the "OK, Google" summoning feature familiar to consumers. While Google says it is pursuing development responsibly, adoption could ultimately stall over concerns such as robots becoming surveillance machines, or being equipped with chat technology that can give offensive responses, as Meta Platforms Inc (META.O) and others have experienced in recent years.
Simulation of Atlantic Hurricane Tracks and Features: A Deep Learning Approach
Bose, Rikhi, Pintar, Adam L., Simiu, Emil
The objective of this paper is to employ machine learning (ML) and deep learning (DL) techniques to obtain from input data (storm features) available in or derived from the HURDAT2 database models capable of simulating important hurricane properties such as landfall location and wind speed that are consistent with historical records. In pursuit of this objective, a trajectory model providing the storm center in terms of longitude and latitude, and intensity models providing the central pressure and maximum 1-$min$ wind speed at 10 $m$ elevation were created. The trajectory and intensity models are coupled and must be advanced together, six hours at a time, as the features that serve as inputs to the models at any given step depend on predictions at the previous time steps. Once a synthetic storm database is generated, properties of interest, such as the frequencies of large wind speeds may be extracted from any part of the simulation domain. The coupling of the trajectory and intensity models obviates the need for an intensity decay inland of the coastline. Prediction results are compared to historical data, and the efficacy of the storm simulation models is demonstrated for three examples: New Orleans, Miami and Cape Hatteras.