Statistical Learning
Deep MMD Gradient Flow without adversarial training
Galashov, Alexandre, de Bortoli, Valentin, Gretton, Arthur
One challenge that arises when applying these models in practice is that the Stein score (that is, the gradient log We propose a gradient flow procedure for generative of the current noisy density) becomes ill-behaved near the modeling by transporting particles from an data distribution (Yang et al., 2023): the diffusion process initial source distribution to a target distribution, needs to be slowed down at this point, which incurs a large where the gradient field on the particles is given number of sampling steps near the data distribution. Indeed, by a noise-adaptive Wasserstein Gradient of the if the manifold hypothesis holds (Tenenbaum et al., 2000; Maximum Mean Discrepancy (MMD). The noiseadaptive Fefferman et al., 2016; Brown et al., 2022) and the data MMD is trained on data distributions corrupted is supported on a lower dimensional space, it is expected by increasing levels of noise, obtained via that the score will explode for noise levels close to zero, a forward diffusion process, as commonly used to ensure that the backward process concentrates on this in denoising diffusion probabilistic models. The lower dimensional manifold (Bortoli, 2023; Pidstrigach, result is a generalization of MMD Gradient Flow, 2022; Chen et al., 2022). While strategies exist to mitigate which we call Diffusion-MMD-Gradient Flow or these issues, they trade-off the quality of the output against DMMD. The divergence training procedure is inference speed, see for instance (Song et al., 2023; Xu et al., related to discriminator training in Generative Adversarial 2023; Sauer et al., 2023). Networks (GAN), but does not require adversarial training. We obtain competitive empirical Generative Adversarial Networks (GANs) (Goodfellow performance in unconditional image generation et al., 2014) represent an alternative popular generative modelling on CIFAR10, MNIST, CELEB-A (64 x64) framework (Brock et al., 2019; Karras et al., 2020a).
Opportunities for Persian Digital Humanities Research with Artificial Intelligence Language Models; Case Study: Forough Farrokhzad
Meymandi, Arash Rasti, Hosseini, Zahra, Davari, Sina, Moshiri, Abolfazl, Rahimi-Golkhandan, Shabnam, Namdar, Khashayar, Feizi, Nikta, Tavakoli-Targhi, Mohamad, Khalvati, Farzad
This study explores the integration of advanced Natural Language Processing (NLP) and Artificial Intelligence (AI) techniques to analyze and interpret Persian literature, focusing on the poetry of Forough Farrokhzad. Utilizing computational methods, we aim to unveil thematic, stylistic, and linguistic patterns in Persian poetry. Specifically, the study employs AI models including transformer-based language models for clustering of the poems in an unsupervised framework. This research underscores the potential of AI in enhancing our understanding of Persian literary heritage, with Forough Farrokhzad's work providing a comprehensive case study. This approach not only contributes to the field of Persian Digital Humanities but also sets a precedent for future research in Persian literary studies using computational techniques.
Disttack: Graph Adversarial Attacks Toward Distributed GNN Training
Zhang, Yuxiang, Liu, Xin, Wu, Meng, Yan, Wei, Yan, Mingyu, Ye, Xiaochun, Fan, Dongrui
Graph Neural Networks (GNNs) have emerged as potent models for graph learning. Distributing the training process across multiple computing nodes is the most promising solution to address the challenges of ever-growing real-world graphs. However, current adversarial attack methods on GNNs neglect the characteristics and applications of the distributed scenario, leading to suboptimal performance and inefficiency in attacking distributed GNN training. In this study, we introduce Disttack, the first framework of adversarial attacks for distributed GNN training that leverages the characteristics of frequent gradient updates in a distributed system. Specifically, Disttack corrupts distributed GNN training by injecting adversarial attacks into one single computing node. The attacked subgraphs are precisely perturbed to induce an abnormal gradient ascent in backpropagation, disrupting gradient synchronization between computing nodes and thus leading to a significant performance decline of the trained GNN. We evaluate Disttack on four large real-world graphs by attacking five widely adopted GNNs. Compared with the state-of-the-art attack method, experimental results demonstrate that Disttack amplifies the model accuracy degradation by 2.75 and achieves speedup by 17.33 on average while maintaining unnoticeability. Keywords: Graph Neural Network Distributed Training Adversarial Attack.
Aspect-oriented Consumer Health Answer Summarization
Chaturvedi, Rochana, Bhattacharya, Abari, Yadav, Shweta
Community Question-Answering (CQA) forums have revolutionized how people seek information, especially those related to their healthcare needs, placing their trust in the collective wisdom of the public. However, there can be several answers in response to a single query, which makes it hard to grasp the key information related to the specific health concern. Typically, CQA forums feature a single top-voted answer as a representative summary for each query. However, a single answer overlooks the alternative solutions and other information frequently offered in other responses. Our research focuses on aspect-based summarization of health answers to address this limitation. Summarization of responses under different aspects such as suggestions, information, personal experiences, and questions can enhance the usability of the platforms. We formalize a multi-stage annotation guideline and contribute a unique dataset comprising aspect-based human-written health answer summaries. We build an automated multi-faceted answer summarization pipeline with this dataset based on task-specific fine-tuning of several state-of-the-art models. The pipeline leverages question similarity to retrieve relevant answer sentences, subsequently classifying them into the appropriate aspect type. Following this, we employ several recent abstractive summarization models to generate aspect-based summaries. Finally, we present a comprehensive human analysis and find that our summaries rank high in capturing relevant content and a wide range of solutions.
CTRL: Continuous-Time Representation Learning on Temporal Heterogeneous Information Network
Li, Chenglin, Xie, Yuanzhen, Yu, Chenyun, Cheng, Lei, Hu, Bo, Li, Zang, Niu, Di
Inductive representation learning on temporal heterogeneous graphs is crucial for scalable deep learning on heterogeneous information networks (HINs) which are time-varying, such as citation networks. However, most existing approaches are not inductive and thus cannot handle new nodes or edges. Moreover, previous temporal graph embedding methods are often trained with the temporal link prediction task to simulate the link formation process of temporal graphs, while ignoring the evolution of high-order topological structures on temporal graphs. To fill these gaps, we propose a Continuous-Time Representation Learning (CTRL) model on temporal HINs. To preserve heterogeneous node features and temporal structures, CTRL integrates three parts in a single layer, they are 1) a \emph{heterogeneous attention} unit that measures the semantic correlation between nodes, 2) a \emph{edge-based Hawkes process} to capture temporal influence between heterogeneous nodes, and 3) \emph{dynamic centrality} that indicates the dynamic importance of a node. We train the CTRL model with a future event (a subgraph) prediction task to capture the evolution of the high-order network structure. Extensive experiments have been conducted on three benchmark datasets. The results demonstrate that our model significantly boosts performance and outperforms various state-of-the-art approaches. Ablation studies are conducted to demonstrate the effectiveness of the model design.
Work Smarter...Not Harder: Efficient Minimization of Dependency Length in SOV Languages
Ranjan, Sidharth, von der Malsburg, Titus
Dependency length minimization is a universally observed quantitative property of natural languages. However, the extent of dependency length minimization, and the cognitive mechanisms through which the language processor achieves this minimization remain unclear. This research offers mechanistic insights by postulating that moving a short preverbal constituent next to the main verb explains preverbal constituent ordering decisions better than global minimization of dependency length in SOV languages. This approach constitutes a least-effort strategy because it's just one operation but simultaneously reduces the length of all preverbal dependencies linked to the main verb. We corroborate this strategy using large-scale corpus evidence across all seven SOV languages that are prominently represented in the Universal Dependency Treebank. These findings align with the concept of bounded rationality, where decision-making is influenced by 'quick-yet-economical' heuristics rather than exhaustive searches for optimal solutions. Overall, this work sheds light on the role of bounded rationality in linguistic decision-making and language evolution.
CardioGenAI: A Machine Learning-Based Framework for Re-Engineering Drugs for Reduced hERG Liability
Kyro, Gregory W., Martin, Matthew T., Watt, Eric D., Batista, Victor S.
The link between in vitro hERG ion channel inhibition and subsequent in vivo QT interval prolongation, a critical risk factor for the development of arrythmias such as Torsade de Pointes, is so well established that in vitro hERG activity alone is often sufficient to end the development of an otherwise promising drug candidate. It is therefore of tremendous interest to develop advanced methods for identifying hERG-active compounds in the early stages of drug development, as well as for proposing redesigned compounds with reduced hERG liability and preserved on-target potency. In this work, we present CardioGenAI, a machine learning-based framework for re-engineering both developmental and commercially available drugs for reduced hERG activity while preserving their pharmacological activity. The framework incorporates novel state-of-the-art discriminative models for predicting hERG channel activity, as well as activity against the voltage-gated NaV1.5 and CaV1.2 channels due to their potential implications in modulating the arrhythmogenic potential induced by hERG channel blockade. We applied the complete framework to pimozide, an FDA-approved antipsychotic agent that demonstrates high affinity to the hERG channel, and generated 100 refined candidates. Remarkably, among the candidates is fluspirilene, a compound which is of the same class of drugs (diphenylmethanes) as pimozide and therefore has similar pharmacological activity, yet exhibits over 700-fold weaker binding to hERG. We envision that this method can effectively be applied to developmental compounds exhibiting hERG liabilities to provide a means of rescuing drug development programs that have stalled due to hERG-related safety concerns. Additionally, the discriminative models can also serve independently as effective components of a virtual screening pipeline. We have made all of our software open-source.
Yet Another Representation of Binary Decision Trees: A Mathematical Demonstration
A decision tree looks like a simple computational graph without cycles, where only the leaf nodes specify the output values and the non-terminals specify their tests or split conditions. From the numerical perspective, we express decision trees in the language of computational graph. We explicitly parameterize the test phase, traversal phase and prediction phase of decision trees based on the bitvectors of non-terminal nodes. As shown later, the decision tree is a shallow binary network in some sense. Especially, we introduce the bitvector matrix to implement the tree traversal in numerical approach, where the core is to convert the logical `AND' operation to arithmetic operations. And we apply this numerical representation to extend and unify diverse decision trees in concept.
Computational analysis of the language of pain: a systematic review
Nunes, Diogo A. P., Ferreira-Gomes, Joana, Neto, Fani, de Matos, David Martins
Objectives: This study aims to systematically review the literature on the computational processing of the language of pain, or pain narratives, whether generated by patients or physicians, identifying current trends and challenges. Methods: Following the PRISMA guidelines, a comprehensive literature search was conducted to select relevant studies on the computational processing of the language of pain and answer pre-defined research questions. Data extraction and synthesis were performed to categorize selected studies according to their primary purpose and outcome, patient and pain population, textual data, computational methodology, and outcome targets. Results: Physician-generated language of pain, specifically from clinical notes, was the most used data. Tasks included patient diagnosis and triaging, identification of pain mentions, treatment response prediction, biomedical entity extraction, correlation of linguistic features with clinical states, and lexico-semantic analysis of pain narratives. Only one study included previous linguistic knowledge on pain utterances in their experimental setup. Most studies targeted their outcomes for physicians, either directly as clinical tools or as indirect knowledge. The least targeted stage of clinical pain care was self-management, in which patients are most involved. Affective and sociocultural dimensions were the least studied domains. Only one study measured how physician performance on clinical tasks improved with the inclusion of the proposed algorithm. Discussion: This review found that future research should focus on analyzing patient-generated language of pain, developing patient-centered resources for self-management and patient-empowerment, exploring affective and sociocultural aspects of pain, and measuring improvements in physician performance when aided by the proposed tools.
Sharp analysis of out-of-distribution error for "importance-weighted" estimators in the overparameterized regime
Lai, Kuo-Wei, Muthukumar, Vidya
Overparameterized models are ubiquitous in machine learning theory and practice today because of their state-of-the-art generalization guarantees (in the sense of low test error) even while perfectly fitting the training data [30, 7]. However, this "good generalization" property does not extend to test data that is distributed differently from training data, termed out-of-distribution (OOD) data [20, 21, 29]. A particularly acute scenario arises when the data is drawn as a mixture from multiple groups (each with a different distribution) and some groups are very under-represented in training data [2]. Under such models, the worst-group generalization error can be significantly degraded with respect to the average generalization error on all groups [1, 27, 21, 20]. The effect of distribution shift on generalization has been sharply characterized in a worst-case/minimax sense, e.g.