kcl
Towards Understanding the Mechanism of Contrastive Learning via Similarity Structure: A Theoretical Analysis
Waida, Hiroki, Wada, Yuichiro, Andéol, Léo, Nakagawa, Takumi, Zhang, Yuhui, Kanamori, Takafumi
Contrastive learning is an efficient approach to self-supervised representation learning. Although recent studies have made progress in the theoretical understanding of contrastive learning, the investigation of how to characterize the clusters of the learned representations is still limited. In this paper, we aim to elucidate the characterization from theoretical perspectives. To this end, we consider a kernel-based contrastive learning framework termed Kernel Contrastive Learning (KCL), where kernel functions play an important role when applying our theoretical results to other frameworks. We introduce a formulation of the similarity structure of learned representations by utilizing a statistical dependency viewpoint. We investigate the theoretical properties of the kernel-based contrastive loss via this formulation. We first prove that the formulation characterizes the structure of representations learned with the kernel-based contrastive learning framework. We show a new upper bound of the classification error of a downstream task, which explains that our theory is consistent with the empirical success of contrastive learning. We also establish a generalization error bound of KCL. Finally, we show a guarantee for the generalization ability of KCL to the downstream classification task via a surrogate bound.
Complementary Labels Learning with Augmented Classes
Li, Zhongnian, Zhang, Jian, Xu, Mengting, Xu, Xinzheng, Zhang, Daoqiang
Complementary Labels Learning (CLL) arises in many real-world tasks such as private questions classification and online learning, which aims to alleviate the annotation cost compared with standard supervised learning. Unfortunately, most previous CLL algorithms were in a stable environment rather than an open and dynamic scenarios, where data collected from unseen augmented classes in the training process might emerge in the testing phase. In this paper, we propose a novel problem setting called Complementary Labels Learning with Augmented Classes (CLLAC), which brings the challenge that classifiers trained by complementary labels should not only be able to classify the instances from observed classes accurately, but also recognize the instance from the Augmented Classes in the testing phase. Specifically, by using unlabeled data, we propose an unbiased estimator of classification risk for CLLAC, which is guaranteed to be provably consistent. Moreover, we provide generalization error bound for proposed method which shows that the optimal parametric convergence rate is achieved for estimation error. Finally, the experimental results on several benchmark datasets verify the effectiveness of the proposed method.
Watch the Neighbors: A Unified K-Nearest Neighbor Contrastive Learning Framework for OOD Intent Discovery
Mou, Yutao, He, Keqing, Wang, Pei, Wu, Yanan, Wang, Jingang, Wu, Wei, Xu, Weiran
Discovering out-of-domain (OOD) intent is important for developing new skills in task-oriented dialogue systems. The key challenges lie in how to transfer prior in-domain (IND) knowledge to OOD clustering, as well as jointly learn OOD representations and cluster assignments. Previous methods suffer from in-domain overfitting problem, and there is a natural gap between representation learning and clustering objectives. In this paper, we propose a unified K-nearest neighbor contrastive learning framework to discover OOD intents. Specifically, for IND pre-training stage, we propose a KCL objective to learn inter-class discriminative features, while maintaining intra-class diversity, which alleviates the in-domain overfitting problem. For OOD clustering stage, we propose a KCC method to form compact clusters by mining true hard negative samples, which bridges the gap between clustering and representation learning. Extensive experiments on three benchmark datasets show that our method achieves substantial improvements over the state-of-the-art methods.
Molecular Contrastive Learning with Chemical Element Knowledge Graph
Fang, Yin, Zhang, Qiang, Yang, Haihong, Zhuang, Xiang, Deng, Shumin, Zhang, Wen, Qin, Ming, Chen, Zhuo, Fan, Xiaohui, Chen, Huajun
Molecular representation learning contributes to multiple downstream tasks such as molecular property prediction and drug design. To properly represent molecules, graph contrastive learning is a promising paradigm as it utilizes self-supervision signals and has no requirements for human annotations. However, prior works fail to incorporate fundamental domain knowledge into graph semantics and thus ignore the correlations between atoms that have common attributes but are not directly connected by bonds. To address these issues, we construct a Chemical Element Knowledge Graph (KG) to summarize microscopic associations between elements and propose a novel Knowledge-enhanced Contrastive Learning (KCL) framework for molecular representation learning. KCL framework consists of three modules. The first module, knowledge-guided graph augmentation, augments the original molecular graph based on the Chemical Element KG. The second module, knowledge-aware graph representation, extracts molecular representations with a common graph encoder for the original molecular graph and a Knowledge-aware Message Passing Neural Network (KMPNN) to encode complex information in the augmented molecular graph. The final module is a contrastive objective, where we maximize agreement between these two views of molecular graphs. Extensive experiments demonstrated that KCL obtained superior performances against state-of-the-art baselines on eight molecular datasets. Visualization experiments properly interpret what KCL has learned from atoms and attributes in the augmented molecular graphs. Our codes and data are available in supplementary materials.
King's College London to deliver healthcare AI model
King's College London (KCL) is partnering up with two companies to deliver an artificial intelligence model in the healthcare and life sciences sector. KCL is joining forces with Owkin, a company that develops AI algorithms for cancer centres and pharmaceutical companies, and American technology company, NVIDIA, to provide Federated Learning, a framework for AI. Federated learning is a machine learning technique that trains an algorithm across multiple decentralised servers holding local data samples, without exchanging their data samples. Owkin aims to demonstrate that the Federating Learning model is safer for patients, and statistically equivalent to the traditional pooled model for analysis. KCL will use Owkin's Federated Learning software and NVIDIA's EGX Intelligent Edge Computing platform to develop research, clinical and operational improvements across a large number of clinical pathways, with cancer, heart failure, dementia and stroke likely areas of early focus.
Collaboration established to deliver federated learning in life sciences
Owkin, which is developing federated learning and AI technologies to advance medical research, has announced a collaboration with technology company NVIDIA and King's College London (KCL) to deliver federated learning in the healthcare and life sciences sector. It will initially connect four of London's teaching hospitals before expanding throughout the UK, and will offer AI services with the aim of accelerating research and improving clinical practice in a wide range of therapeutic areas, including cancer, heart failure and neurodegenerative disease. Owkin's co-founder and chief scientific officer, Gilles Wainrib, said: "This partnership brings together the best players in life science & healthcare, machine learning and data centre infrastructure. NVIDIA's platforms create the ideal and flexible footprint for hospitals to invest in machine learning. King's College London has assembled the engineering, medical and data science talent, the high-quality patient data, and the governance framework in the AI4VBH Centre, that will show the world the future of healthcare analytics and the power of machine learning. Together we will be enabling the formation of a decentralised dataset that will generate enormous value for research and clinical practice. "Owkin hopes to demonstrate that a Federating Learning architecture is safer for patients, and statistically equivalent to the traditional pooled model for analysis.
Multi-class Classification without Multi-class Labels
Hsu, Yen-Chang, Lv, Zhaoyang, Schlosser, Joel, Odom, Phillip, Kira, Zsolt
This work presents a new strategy for multi-class classification that requires no class-specific labels, but instead leverages pairwise similarity between examples, which is a weaker form of annotation. The proposed method, meta classification learning, optimizes a binary classifier for pairwise similarity prediction and through this process learns a multi-class classifier as a submodule. We formulate this approach, present a probabilistic graphical model for it, and derive a surprisingly simple loss function that can be used to learn neural network-based models. We then demonstrate that this same framework generalizes to the supervised, unsupervised cross-task, and semi-supervised settings. Our method is evaluated against state of the art in all three learning paradigms and shows a superior or comparable accuracy, providing evidence that learning multi-class classification without multi-class labels is a viable learning option.
TES HireWire
The Division of Health and Social Care Research (www.kcl.ac.uk/hscr) is seeking to appoint a Research Associate in Health Informatics to join the RobotReviewer project (www.robotreviewer.net), The project is a UK/US collaboration funded by the US National Institutes of Health/National Library of Medicine. Candidates should hold a PhD (in computer science, statistics, artificial intelligence, life sciences, or a related subject) or substantial equivalent experience (working in a research capacity in industry with a focus on the same subject areas). The successful candidate will participate in all aspects of the research, including developing and evaluating novel machine learning algorithms, writing research articles for publication, and contributing to the development of our open source software. S/he will be proficient in one or more general purpose programming languages, and ideally will have experience of using our current toolset (Python, Scikit-learn, Keras, and Theano).