Statistical Learning
An Exploratory Study on Utilising the Web of Linked Data for Product Data Mining
The Linked Open Data practice has led to a significant growth of structured data on the Web in the last decade. Such structured data describe real-world entities in a machine-readable way, and have created an unprecedented opportunity for research in the field of Natural Language Processing. However, there is a lack of studies on how such data can be used, for what kind of tasks, and to what extent they can be useful for these tasks. This work focuses on the e-commerce domain to explore methods of utilising such structured data to create language resources that may be used for product classification and linking. We process billions of structured data points in the form of RDF n-quads, to create multi-million words of product-related corpora that are later used in three different ways for creating of language resources: training word embedding models, continued pre-training of BERT-like language models, and training Machine Translation models that are used as a proxy to generate product-related keywords. Our evaluation on an extensive set of benchmarks shows word embeddings to be the most reliable and consistent method to improve the accuracy on both tasks (with up to 6.9 percentage points in macro-average F1 on some datasets). The other two methods however, are not as useful. Our analysis shows that this could be due to a number of reasons, including the biased domain representation in the structured data and lack of vocabulary coverage. We share our datasets and discuss how our lessons learned could be taken forward to inform future research in this direction.
Efficient Algorithms For Fair Clustering with a New Fairness Notion
Gupta, Shivam, Ghalme, Ganesh, Krishnan, Narayanan C., Jain, Shweta
We revisit the problem of fair clustering, first introduced by Chierichetti et al., that requires each protected attribute to have approximately equal representation in every cluster; i.e., a balance property. Existing solutions to fair clustering are either not scalable or do not achieve an optimal trade-off between clustering objective and fairness. In this paper, we propose a new notion of fairness, which we call $tau$-fair fairness, that strictly generalizes the balance property and enables a fine-grained efficiency vs. fairness trade-off. Furthermore, we show that simple greedy round-robin based algorithms achieve this trade-off efficiently. Under a more general setting of multi-valued protected attributes, we rigorously analyze the theoretical properties of the our algorithms. Our experimental results suggest that the proposed solution outperforms all the state-of-the-art algorithms and works exceptionally well even for a large number of clusters.
Large-Scale Learning with Fourier Features and Tensor Decompositions
Wesel, Frederiek, Batselier, Kim
Random Fourier features provide a way to tackle large-scale machine learning problems with kernel methods. Their slow Monte Carlo convergence rate has motivated the research of deterministic Fourier features whose approximation error decreases exponentially with the number of frequencies. However, due to their tensor product structure these methods suffer heavily from the curse of dimensionality, limiting their applicability to two or three-dimensional scenarios. In our approach we overcome said curse of dimensionality by exploiting the tensor product structure of deterministic Fourier features, which enables us to represent the model parameters as a low-rank tensor decomposition. We derive a monotonically converging block coordinate descent algorithm with linear complexity in both the sample size and the dimensionality of the inputs for a regularized squared loss function, allowing to learn a parsimonious model in decomposed form using deterministic Fourier features. We demonstrate by means of numerical experiments how our low-rank tensor approach obtains the same performance of the corresponding nonparametric model, consistently outperforming random Fourier features.
LightAutoML: AutoML Solution for a Large Financial Services Ecosystem
Vakhrushev, Anton, Ryzhkov, Alexander, Savchenko, Maxim, Simakov, Dmitry, Damdinov, Rinchin, Tuzhilin, Alexander
In particular, our ecosystem has the satisfying the set of idiosyncratic requirements that this ecosystem following set of requirements: has for AutoML solutions. Our framework was piloted and deployed in numerous applications and performed at the level of - AutoML system should be able to work with different types the experienced data scientists while building high-quality ML of data collected from hundreds of different information models significantly faster than these data scientists. We also compare systems and often changes more rapidly than these systems the performance of our system with various general-purpose can be fully documented using metadata and painstakingly open source AutoML solutions and show that it performs better for preprocessed by data scientists for the ML tasks using ETL most of the ecosystem and OpenML problems. We also present the tools.
Relating the Partial Dependence Plot and Permutation Feature Importance to the Data Generating Process
Molnar, Christoph, Freiesleben, Timo, König, Gunnar, Casalicchio, Giuseppe, Wright, Marvin N., Bischl, Bernd
Scientists and practitioners increasingly rely on machine learning to model data and draw conclusions. Compared to statistical modeling approaches, machine learning makes fewer explicit assumptions about data structures, such as linearity. However, their model parameters usually cannot be easily related to the data generating process. To learn about the modeled relationships, partial dependence (PD) plots and permutation feature importance (PFI) are often used as interpretation methods. However, PD and PFI lack a theory that relates them to the data generating process. We formalize PD and PFI as statistical estimators of ground truth estimands rooted in the data generating process. We show that PD and PFI estimates deviate from this ground truth due to statistical biases, model variance and Monte Carlo approximation errors. To account for model variance in PD and PFI estimation, we propose the learner-PD and the learner-PFI based on model refits, and propose corrected variance and confidence interval estimators.
Sample Noise Impact on Active Learning
Abraham, Alexandre, Dreyfus-Schmidt, Léo
This work explores the effect of noisy sample selection in active learning strategies. We show on both synthetic problems and real-life use-cases that knowledge of the sample noise can significantly improve the performance of active learning strategies. Building on prior work, we propose a robust sampler, Incremental Weighted K-Means that brings significant improvement on the synthetic tasks but only a marginal uplift on real-life ones. We hope that the questions raised in this paper are of interest to the community and could open new paths for active learning research.
Statistical Estimation and Inference via Local SGD in Federated Learning
Li, Xiang, Liang, Jiadong, Chang, Xiangyu, Zhang, Zhihua
Federated Learning is a novel distributed computing paradigm for collaboratively training a global model from data that remote clients hold [McMahan et al., 2017]. The clients can only cooperate with a central server (e.g., service provider) to train the global model without sharing local datasets. Thus, federated learning can protect sensitive information that data often contain, such as personal identity information and state of health information, from unauthorized access of service providers. The challenge arises when limited data access together with memory constraints, communication budget, and computation restrictions make the traditional statistical estimation and inference methods [Li et al., 2020b, Fan et al., 2021] no longer applicable in the federated learning scenario. A typical federated learning system considers a pool of K clients, in which the k-th client has a local dataset consisting of i.i.d.
Coordinating Narratives and the Capitol Riots on Parler
Ng, Lynnette Hui Xian, Cruickshank, Iain, Carley, Kathleen M.
Coordinated disinformation campaigns are used to influence social media users, potentially leading to offline violence. In this study, we introduce a general methodology to uncover coordinated messaging through analysis of user parleys on Parler. The proposed method constructs a user-to-user coordination network graph induced by a user-to-text graph and a text-to-text similarity graph. The text-to-text graph is constructed based on the textual similarity of Parler posts. We study three influential groups of users in the 6 January 2020 Capitol riots and detect networks of coordinated user clusters that are all posting similar textual content in support of different disinformation narratives related to the U.S. 2020 elections.
Artificial Intelligence in Dry Eye Disease
Storås, Andrea M., Strümke, Inga, Riegler, Michael A., Grauslund, Jakob, Hammer, Hugo L., Yazidi, Anis, Halvorsen, Pål, Gundersen, Kjell G., Utheim, Tor P., Jackson, Catherine
Dry eye disease (DED) has a prevalence of between 5 and 50\%, depending on the diagnostic criteria used and population under study. However, it remains one of the most underdiagnosed and undertreated conditions in ophthalmology. Many tests used in the diagnosis of DED rely on an experienced observer for image interpretation, which may be considered subjective and result in variation in diagnosis. Since artificial intelligence (AI) systems are capable of advanced problem solving, use of such techniques could lead to more objective diagnosis. Although the term `AI' is commonly used, recent success in its applications to medicine is mainly due to advancements in the sub-field of machine learning, which has been used to automatically classify images and predict medical outcomes. Powerful machine learning techniques have been harnessed to understand nuances in patient data and medical images, aiming for consistent diagnosis and stratification of disease severity. This is the first literature review on the use of AI in DED. We provide a brief introduction to AI, report its current use in DED research and its potential for application in the clinic. Our review found that AI has been employed in a wide range of DED clinical tests and research applications, primarily for interpretation of interferometry, slit-lamp and meibography images. While initial results are promising, much work is still needed on model development, clinical testing and standardisation.
On-target Adaptation
Wang, Dequan, Liu, Shaoteng, Ebrahimi, Sayna, Shelhamer, Evan, Darrell, Trevor
Domain adaptation seeks to mitigate the shift between training on the \emph{source} domain and testing on the \emph{target} domain. Most adaptation methods rely on the source data by joint optimization over source data and target data. Source-free methods replace the source data with a source model by fine-tuning it on target. Either way, the majority of the parameter updates for the model representation and the classifier are derived from the source, and not the target. However, target accuracy is the goal, and so we argue for optimizing as much as possible on the target data. We show significant improvement by on-target adaptation, which learns the representation purely from target data while taking only the source predictions for supervision. In the long-tailed classification setting, we show further improvement by on-target class distribution learning, which learns the (im)balance of classes from target data.