Africa
SEN12TS -- Largest land cover classification dataset ?!
Land cover classification (or semantic segmentation in the CV context), is one of the most important applications of machine / deep learning models in remote sensing image analysis. There are numerous benchmark datasets with different features, designed and published for LULC classification task. Although radar-derived and optical imagery are widely available at similar timescales and spatial resolutions, some issues make their combined processing more complicated. These issues include coregistration between satellite missions, processing of SAR imagery to correct for ground geometry and incidence angle; and the most important one, lack of reliable labeled ground truth pixels appropriate for research purposes. Here, I'm going to introduce SEN12TS; a very large satellite image dataset (1.69 TB in storage!), designed by University of Colombia and Descartes Lab, specifically for land cover classification.
AI-Based Algorithmic Trading Can Disrupt Bitcoin Market
AI technology has had a huge impact on the direction of many industries. The traditional financial sector has been one of the most heavily affected. The market for AI in finance is exploding. Allied Research projects it will grow from under $4 billion in 2020 to over $64 billion in 2030. These forecasts pertain to the expected proliferation of AI in the traditional financial sector.
A scalable pipeline for COVID-19: the case study of Germany, Czechia and Poland
Abdussalam, Wildan, Mertel, Adam, Fan, Kai, Schรผler, Lennart, Schlechte-Weลnicz, Weronika, Calabrese, Justin M.
Throughout the coronavirus disease 2019 (COVID-19) pandemic, decision makers have relied on forecasting models to determine and implement non-pharmaceutical interventions (NPI). In building the forecasting models, continuously updated datasets from various stakeholders including developers, analysts, and testers are required to provide precise predictions. Here we report the design of a scalable pipeline which serves as a data synchronization to support inter-country top-down spatiotemporal observations and forecasting models of COVID-19, named the where2test, for Germany, Czechia and Poland. We have built an operational data store (ODS) using PostgreSQL to continuously consolidate datasets from multiple data sources, perform collaborative work, facilitate high performance data analysis, and trace changes. The ODS has been built not only to store the COVID-19 data from Germany, Czechia, and Poland but also other areas. Employing the dimensional fact model, a schema of metadata is capable of synchronizing the various structures of data from those regions, and is scalable to the entire world. Next, the ODS is populated using batch Extract, Transfer, and Load (ETL) jobs. The SQL queries are subsequently created to reduce the need for pre-processing data for users. The data can then support not only forecasting using a version-controlled Arima-Holt model and other analyses to support decision making, but also risk calculator and optimisation apps. The data synchronization runs at a daily interval, which is displayed at https://www.where2test.de.
Tensor Decomposition based Personalized Federated Learning
Wang, Qing, Jin, Jing, Liu, Xiaofeng, Zong, Huixuan, Shao, Yunfeng, Li, Yinchuan
Federated learning (FL) is a new distributed machine learning framework that can achieve reliably collaborative training without collecting users' private data. However, due to FL's frequent communication and average aggregation strategy, they experience challenges scaling to statistical diversity data and large-scale models. In this paper, we propose a personalized FL framework, named Tensor Decomposition based Personalized Federated learning (TDPFed), in which we design a novel tensorized local model with tensorized linear layers and convolutional layers to reduce the communication cost. TDPFed uses a bi-level loss function to decouple personalized model optimization from the global model learning by controlling the gap between the personalized model and the tensorized local model. Moreover, an effective distributed learning strategy and two different model aggregation strategies are well designed for the proposed TDPFed framework. Theoretical convergence analysis and thorough experiments demonstrate that our proposed TDPFed framework achieves state-of-the-art performance while reducing the communication cost.
Ab-initio quantum chemistry with neural-network wavefunctions
Hermann, Jan, Spencer, James, Choo, Kenny, Mezzacapo, Antonio, Foulkes, W. M. C., Pfau, David, Carleo, Giuseppe, Noรฉ, Frank
Machine learning and specifically deep-learning methods have outperformed human capabilities in many pattern recognition and data processing problems, in game playing, and now also play an increasingly important role in scientific discovery. A key application of machine learning in the molecular sciences is to learn potential energy surfaces or force fields from ab-initio solutions of the electronic Schr\"odinger equation using datasets obtained with density functional theory, coupled cluster, or other quantum chemistry methods. Here we review a recent and complementary approach: using machine learning to aid the direct solution of quantum chemistry problems from first principles. Specifically, we focus on quantum Monte Carlo (QMC) methods that use neural network ansatz functions in order to solve the electronic Schr\"odinger equation, both in first and second quantization, computing ground and excited states, and generalizing over multiple nuclear configurations. Compared to existing quantum chemistry methods, these new deep QMC methods have the potential to generate highly accurate solutions of the Schr\"odinger equation at relatively modest computational cost.
What Do NLP Researchers Believe? Results of the NLP Community Metasurvey
Michael, Julian, Holtzman, Ari, Parrish, Alicia, Mueller, Aaron, Wang, Alex, Chen, Angelica, Madaan, Divyam, Nangia, Nikita, Pang, Richard Yuanzhe, Phang, Jason, Bowman, Samuel R.
We present the results of the NLP Community Metasurvey. Run from May to June 2022, the survey elicited opinions on controversial issues, including industry influence in the field, concerns about AGI, and ethics. Our results put concrete numbers to several controversies: For example, respondents are split almost exactly in half on questions about the importance of artificial general intelligence, whether language models understand language, and the necessity of linguistic structure and inductive bias for solving NLP problems. In addition, the survey posed meta-questions, asking respondents to predict the distribution of survey responses. This allows us not only to gain insight on the spectrum of beliefs held by NLP researchers, but also to uncover false sociological beliefs where the community's predictions don't match reality. We find such mismatches on a wide range of issues. Among other results, the community greatly overestimates its own belief in the usefulness of benchmarks and the potential for scaling to solve real-world problems, while underestimating its own belief in the importance of linguistic structure, inductive bias, and interdisciplinary science.
An End-to-End OCR Framework for Robust Arabic-Handwriting Recognition using a Novel Transformers-based Model and an Innovative 270 Million-Words Multi-Font Corpus of Classical Arabic with Diacritics
Mostafa, Aly, Mohamed, Omar, Ashraf, Ali, Elbehery, Ahmed, Jamal, Salma, Salah, Anas, Ghoneim, Amr S.
This research is the second phase in a series of investigations on developing an Optical Character Recognition (OCR) of Arabic historical documents and examining how different modeling procedures interact with the problem. The first research studied the effect of Transformers on our custom-built Arabic dataset. One of the downsides of the first research was the size of the training data, a mere 15000 images from our 30 million images, due to lack of resources. Also, we add an image enhancement layer, time and space optimization, and Post-Correction layer to aid the model in predicting the correct word for the correct context. Notably, we propose an end-to-end text recognition approach using Vision Transformers as an encoder, namely BEIT, and vanilla Transformer as a decoder, eliminating CNNs for feature extraction and reducing the model's complexity. The experiments show that our end-to-end model outperforms Convolutions Backbones. The model attained a CER of 4.46%.
Towards Higher-order Topological Consistency for Unsupervised Network Alignment
Sun, Qingqiang, Lin, Xuemin, Zhang, Ying, Zhang, Wenjie, Chen, Chaoqi
--Network alignment task, which aims to identify corresponding nodes in different networks, is of great significance for many subsequent applications. Without the need for labeled anchor links, unsupervised alignment methods have been attracting more and more attention. However, the topological consistency assumptions defined by existing methods are generally low-order and less accurate because only the edge-indiscriminative topological pattern is considered, which is especially risky in an unsupervised setting. T o reposition the focus of the alignment process from low-order to higher-order topological consistency, in this paper, we propose a fully unsupervised network alignment framework named HTC. The proposed higher-order topological consistency is formulated based on edge orbits, which is merged into the information aggregation process of a graph convolutional network so that the alignment consistencies are transformed into the similarity of node embeddings. Furthermore, the encoder is trained to be multi-orbit-aware and then be refined to identify more trusted anchor links. Node correspondence is comprehensively evaluated by integrating all different orders of consistency. In addition to sound theoretical analysis, the superiority of the proposed method is also empirically demonstrated through extensive experimental evaluation. On three pairs of real-world datasets and two pairs of synthetic datasets, our HTC consistently outperforms a wide variety of unsupervised and supervised methods with the least or comparable time consumption. It also exhibits robustness to structural noise as a result of our multi-orbit-aware training mechanism. Network alignment task, which aims to identify entity correspondence across different networks, is usually the very first step of many downstream analyzing tasks. For instance, recognizing the same user on different social networks can facilitate friend suggestion, item recommendation, personalized advertisement [1]-[5]. Similar scenarios also exist widely in other fields, such as protein network analysis [6], knowledge discovery [7], etc. Identifying corresponding nodes across different networks is an extremely hard task, even for humans. Manually labelling correspondence can be prohibitively challenging, expensive (in human efforts, time, and money costs), and tedious [8]. Due to such obstacles, in some cases, it may be impractical to get access to sufficient labels for training well-performed supervised or even semi-supervised models [4], [9]. By contrast, unsupervised models can be trained without the need for labeled data, which is more flexible and practical in real-world application scenarios. Thus, unsupervised alignment methods have been drawing a surge of interest recently [10]-[12].
Large-N dynamics of the spiked tensor model with random initial conditions
Non-convex multidimensional optimization and the related problem of finding the global minimum in rough landscapes are crucial challenges of modern science. Such problems were extensively studied in the context of the spin-glass systems [1], and found applications in biology [2], finance [3], and data science [4]. Here, we focus on a task motivated by data science and consider the model of the signal recovering from a noisy high-dimensional data tensor - the spiked tensor model (tensor PCA) [5, 6, 7].
Investigating data partitioning strategies for crosslinguistic low-resource ASR evaluation
Liu, Zoey, Spence, Justin, Prud'hommeaux, Emily
Many automatic speech recognition (ASR) data sets include a single pre-defined test set consisting of one or more speakers whose speech never appears in the training set. This "hold-speaker(s)-out" data partitioning strategy, however, may not be ideal for data sets in which the number of speakers is very small. This study investigates ten different data split methods for five languages with minimal ASR training resources. We find that (1) model performance varies greatly depending on which speaker is selected for testing; (2) the average word error rate (WER) across all held-out speakers is comparable not only to the average WER over multiple random splits but also to any given individual random split; (3) WER is also generally comparable when the data is split heuristically or adversarially; (4) utterance duration and intensity are comparatively more predictive factors of variability regardless of the data split. These results suggest that the widely used hold-speakers-out approach to ASR data partitioning can yield results that do not reflect model performance on unseen data or speakers. Random splits can yield more reliable and generalizable estimates when facing data sparsity.