Statistical Learning
Behavior Structformer: Learning Players Representations with Structured Tokenization
Smirnov, Oleg, Polisi, Labinot
The landmark Transformer [9] model has demonstrated impressive performance across a wide range of scenarios, extending well beyond the realm of Natural Language Processing. The potential of multi-head self-attention method lies in the ability to pick up a signal from any data modality, provided it exhibits a spatial (e.g., sequential) structure and is appropriately preprocessed into discrete tokens for model consumption. However, in practice, the convergence rate of Transformer models in default configurations is considered unsatisfactory. This issue can be mitigated by incorporating prior domain knowledge and inductive biases during the tokenization phase, making the data more suitable for processing by the algorithm. In the field of Computer Vision, the Hybrid Vision Transformers approach [2] has shown that leveraging a pre-trained convolutional backbone as a feature extractor leads to faster convergence and improved downstream performance. Similar observations have been made in customer modeling for personalization [7], where purchase and non-purchase actions were pre-embedded before processing with a BERT-like model. In the healthcare sector, a sequence of electronic health records was preprocessed based on domain expert knowledge to be further consumed by a Transformerbased model with an objective to predict the next medical code [5]. Inspired by these advances, we propose a method for modeling in-game player behavior data that employs a structured approach to convert tracking events into dense tokens. We benchmark and compare the proposed approach against the tabular and semi-structured baselines.
AGBD: A Global-scale Biomass Dataset
Sialelli, Ghjulia, Peters, Torben, Wegner, Jan D., Schindler, Konrad
Accurate estimates of Above Ground Biomass (AGB) are essential in addressing two of humanity's biggest challenges, climate change and biodiversity loss. Existing datasets for AGB estimation from satellite imagery are limited. Either they focus on specific, local regions at high resolution, or they offer global coverage at low resolution. There is a need for a machine learning-ready, globally representative, high-resolution benchmark. Our findings indicate significant variability in biomass estimates across different vegetation types, emphasizing the necessity for a dataset that accurately captures global diversity. To address these gaps, we introduce a comprehensive new dataset that is globally distributed, covers a range of vegetation types, and spans several years. This dataset combines AGB reference data from the GEDI mission with data from Sentinel-2 and PALSAR-2 imagery. Additionally, it includes pre-processed high-level features such as a dense canopy height map, an elevation map, and a land-cover classification map. We also produce a dense, high-resolution (10m) map of AGB predictions for the entire area covered by the dataset. Rigorously tested, our dataset is accompanied by several benchmark models and is publicly available. It can be easily accessed using a single line of code, offering a solid basis for efforts towards global AGB estimation. The GitHub repository github.com/ghjuliasialelli/AGBD serves as a one-stop shop for all code and data.
A Combination Model Based on Sequential General Variational Mode Decomposition Method for Time Series Prediction
Chen, Wei, Yang, Yuanyuan, Liu, Jianyu
For example, combining ARIMA with various decomposition algorithms such as Empirical Mode Decomposition (EMD) and Variational Mode Decomposition (VMD) for predicting complex time series; For example, using an improved ARMA model for stock market forecasting. However, the above models need to be built on the basis of stable sequence data, and usually require testing and preprocessing of the original data, which may lead to the loss of some hidden information, especially in big data samples, and this disadvantage is easily magnified. With the development of computer technology, intelligent models represented by artificial neural networks (ANNs) are gradually emerging. This type of model is good at handling incomplete, fuzzy, uncertain, or irregular data, and has a good fit to nonlinear relationships. Shallow neural networks represented by backpropagation neural networks (BPNN) and shallow machine learning represented by support vector machines (SVM) are also widely used in financial market prediction. However, shallow neural networks do not consider the temporal nature of data, and financial time series often have certain long-term dependencies. Therefore, recurrent neural networks (RNNs) with memory function have become the latest choice. The output of RNN at a certain moment can be used as input to feedback to neurons again, and this cascade structure is very suitable for time series data, which can preserve the dependency relationships in the data.
Adaptively Learning to Select-Rank in Online Platforms
Wang, Jingyuan, Dong, Perry, Jin, Ying, Zhan, Ruohan, Zhou, Zhengyuan
Ranking algorithms are fundamental to various online platforms across e-commerce sites to content streaming services. Our research addresses the challenge of adaptively ranking items from a candidate pool for heterogeneous users, a key component in personalizing user experience. We develop a user response model that considers diverse user preferences and the varying effects of item positions, aiming to optimize overall user satisfaction with the ranked list. We frame this problem within a contextual bandits framework, with each ranked list as an action. Our approach incorporates an upper confidence bound to adjust predicted user satisfaction scores and selects the ranking action that maximizes these adjusted scores, efficiently solved via maximum weight imperfect matching. We demonstrate that our algorithm achieves a cumulative regret bound of $O(d\sqrt{NKT})$ for ranking $K$ out of $N$ items in a $d$-dimensional context space over $T$ rounds, under the assumption that user responses follow a generalized linear model. This regret alleviates dependence on the ambient action space, whose cardinality grows exponentially with $N$ and $K$ (thus rendering direct application of existing adaptive learning algorithms -- such as UCB or Thompson sampling -- infeasible). Experiments conducted on both simulated and real-world datasets demonstrate our algorithm outperforms the baseline.
VERA: Generating Visual Explanations of Two-Dimensional Embeddings via Region Annotation
Poličar, Pavlin G., Zupan, Blaž
Two-dimensional embeddings obtained from dimensionality reduction techniques, such as MDS, t-SNE, and UMAP, are widely used across various disciplines to visualize high-dimensional data. These visualizations provide a valuable tool for exploratory data analysis, allowing researchers to visually identify clusters, outliers, and other interesting patterns in the data. However, interpreting the resulting visualizations can be challenging, as it often requires additional manual inspection to understand the differences between data points in different regions of the embedding space. To address this issue, we propose Visual Explanations via Region Annotation (VERA), an automatic embedding-annotation approach that generates visual explanations for any two-dimensional embedding. VERA produces informative explanations that characterize distinct regions in the embedding space, allowing users to gain an overview of the embedding landscape at a glance. Unlike most existing approaches, which typically require some degree of manual user intervention, VERA produces static explanations, automatically identifying and selecting the most informative visual explanations to show to the user. We illustrate the usage of VERA on a real-world data set and validate the utility of our approach with a comparative user study. Our results demonstrate that the explanations generated by VERA are as useful as fully-fledged interactive tools on typical exploratory data analysis tasks but require significantly less time and effort from the user.
Continuous Geometry-Aware Graph Diffusion via Hyperbolic Neural PDE
Liu, Jiaxu, Yi, Xinping, Wu, Sihao, Yin, Xiangyu, Zhang, Tianle, Huang, Xiaowei, Jin, Shi
While Hyperbolic Graph Neural Network (HGNN) has recently emerged as a powerful tool dealing with hierarchical graph data, the limitations of scalability and efficiency hinder itself from generalizing to deep models. In this paper, by envisioning depth as a continuous-time embedding evolution, we decouple the HGNN and reframe the information propagation as a partial differential equation, letting node-wise attention undertake the role of diffusivity within the Hyperbolic Neural PDE (HPDE). By introducing theoretical principles \textit{e.g.,} field and flow, gradient, divergence, and diffusivity on a non-Euclidean manifold for HPDE integration, we discuss both implicit and explicit discretization schemes to formulate numerical HPDE solvers. Further, we propose the Hyperbolic Graph Diffusion Equation (HGDE) -- a flexible vector flow function that can be integrated to obtain expressive hyperbolic node embeddings. By analyzing potential energy decay of embeddings, we demonstrate that HGDE is capable of modeling both low- and high-order proximity with the benefit of local-global diffusivity functions. Experiments on node classification and link prediction and image-text classification tasks verify the superiority of the proposed method, which consistently outperforms various competitive models by a significant margin.
Concept Formation and Alignment in Language Models: Bridging Statistical Patterns in Latent Space to Concept Taxonomy
Khatir, Mehrdad, Reddy, Chandan K.
This paper explores the concept formation and alignment within the realm of language models (LMs). We propose a mechanism for identifying concepts and their hierarchical organization within the semantic representations learned by various LMs, encompassing a spectrum from early models like Glove to the transformer-based language models like ALBERT and T5. Our approach leverages the inherent structure present in the semantic embeddings generated by these models to extract a taxonomy of concepts and their hierarchical relationships. This investigation sheds light on how LMs develop conceptual understanding and opens doors to further research to improve their ability to reason and leverage real-world knowledge. We further conducted experiments and observed the possibility of isolating these extracted conceptual representations from the reasoning modules of the transformer-based LMs. The observed concept formation along with the isolation of conceptual representations from the reasoning modules can enable targeted token engineering to open the door for potential applications in knowledge transfer, explainable AI, and the development of more modular and conceptually grounded language models.
Learning Divergence Fields for Shift-Robust Graph Representations
Wu, Qitian, Nie, Fan, Yang, Chenxiao, Yan, Junchi
Real-world data generation often involves certain geometries (e.g., graphs) that induce instance-level interdependence. This characteristic makes the generalization of learning models more difficult due to the intricate interdependent patterns that impact data-generative distributions and can vary from training to testing. In this work, we propose a geometric diffusion model with learnable divergence fields for the challenging generalization problem with interdependent data. We generalize the diffusion equation with stochastic diffusivity at each time step, which aims to capture the multi-faceted information flows among interdependent data. Furthermore, we derive a new learning objective through causal inference, which can guide the model to learn generalizable patterns of interdependence that are insensitive across domains. Regarding practical implementation, we introduce three model instantiations that can be considered as the generalized versions of GCN, GAT, and Transformers, respectively, which possess advanced robustness against distribution shifts. We demonstrate their promising efficacy for out-of-distribution generalization on diverse real-world datasets.
Predicting Polymer Properties Based on Multimodal Multitask Pretraining
Wang, Fanmeng, Guo, Wentao, Cheng, Minjie, Yuan, Shen, Xu, Hongteng, Gao, Zhifeng
In the past few decades, polymers, high-molecular-weight compounds formed by bonding numerous identical or similar monomers covalently, have played an essential role in various scientific fields. In this context, accurate prediction of their properties is becoming increasingly crucial. Typically, the properties of a polymer, such as plasticity, conductivity, bio-compatibility, and so on, are highly correlated with its 3D structure. However, current methods for predicting polymer properties heavily rely on information from polymer SMILES sequences (P-SMILES strings) while ignoring crucial 3D structural information, leading to sub-optimal performance. In this work, we propose MMPolymer, a novel multimodal multitask pretraining framework incorporating both polymer 1D sequential information and 3D structural information to enhance downstream polymer property prediction tasks. Besides, to overcome the limited availability of polymer 3D data, we further propose the "Star Substitution" strategy to extract 3D structural information effectively. During pretraining, MMPolymer not only predicts masked tokens and recovers 3D coordinates but also achieves the cross-modal alignment of latent representation. Subsequently, we further fine-tune the pretrained MMPolymer for downstream polymer property prediction tasks in the supervised learning paradigm. Experimental results demonstrate that MMPolymer achieves state-of-the-art performance in various polymer property prediction tasks. Moreover, leveraging the pretrained MMPolymer and using only one modality (either P-SMILES string or 3D conformation) during fine-tuning can also surpass existing polymer property prediction methods, highlighting the exceptional capability of MMPolymer in polymer feature extraction and utilization. Our online platform for polymer property prediction is available at https://app.bohrium.dp.tech/mmpolymer.
PANDORA: Deep graph learning based COVID-19 infection risk level forecasting
Yu, Shuo, Xia, Feng, Wang, Yueru, Li, Shihao, Febrinanto, Falih, Chetty, Madhu
COVID-19 as a global pandemic causes a massive disruption to social stability that threatens human life and the economy. Policymakers and all elements of society must deliver measurable actions based on the pandemic's severity to minimize the detrimental impact of COVID-19. A proper forecasting system is arguably important to provide an early signal of the risk of COVID-19 infection so that the authorities are ready to protect the people from the worst. However, making a good forecasting model for infection risks in different cities or regions is not an easy task, because it has a lot of influential factors that are difficult to be identified manually. To address the current limitations, we propose a deep graph learning model, called PANDORA, to predict the infection risks of COVID-19, by considering all essential factors and integrating them into a geographical network. The framework uses geographical position relations and transportation frequency as higher-order structural properties formulated by higher-order network structures (i.e., network motifs). Moreover, four significant node attributes (i.e., multiple features of a particular area, including climate, medical condition, economy, and human mobility) are also considered. We propose three different aggregators to better aggregate node attributes and structural features, namely, Hadamard, Summation, and Connection. Experimental results over real data show that PANDORA outperforms the baseline method with higher accuracy and faster convergence speed, no matter which aggregator is chosen. We believe that PANDORA using deep graph learning provides a promising approach to get superior performance in infection risk level forecasting and help humans battle the COVID-19 crisis.