Goto

Collaborating Authors

 enzyme function prediction


Multimodal Quantum Vision Transformer for Enzyme Commission Classification from Biochemical Representations

arXiv.org Artificial Intelligence

--Accurately predicting enzyme functionality remains one of the major challenges in computational biology, particularly for enzymes with limited structural annotations or sequence homology. We present a novel multimodal Quantum Machine Learning (QML) framework that enhances Enzyme Commission (EC) classification by integrating four complementary biochemical modalities: protein sequence embeddings, quantum-derived electronic descriptors, molecular graph structures, and 2D molecular image representations. Quantum Vision Transformer (QVT) backbone equipped with modality-specific encoders and a unified cross-attention fusion module. Experimental results demonstrate that our multimodal QVT model achieves a top-1 accuracy of 85.1%, outperforming sequence-only baselines by a substantial margin and achieving better performance results compared to other QML models. Enzymes play a pivotal role in virtually every aspect of biological chemistry, acting as highly specialized catalysts that drive metabolic pathways, DNA replication, cell signaling, and other essential processes in living organisms [1], [2]. Consequently, the ability to accurately predict enzyme function has immense implications, facilitating the discovery of novel biocatalysts, guiding metabolic engineering efforts, and accelerating drug development.


Interpretable Enzyme Function Prediction via Residue-Level Detection

arXiv.org Artificial Intelligence

Predicting multiple functions labeled with Enzyme Commission (EC) numbers from the enzyme sequence is of great significance but remains a challenge due to its sparse multi-label classification nature, i.e., each enzyme is typically associated with only a few labels out of more than 6000 possible EC numbers. However, existing machine learning algorithms generally learn a fixed global representation for each enzyme to classify all functions, thereby they lack interpretability and the fine-grained information of some function-specific local residue fragments may be overwhelmed. Here we present an attention-based framework, namely ProtDETR (Protein Detection Transformer), by casting enzyme function prediction as a detection problem. It uses a set of learnable functional queries to adaptatively extract different local representations from the sequence of residue-level features for predicting different EC numbers. ProtDETR not only significantly outperforms existing deep learning-based enzyme function prediction methods, but also provides a new interpretable perspective on automatically detecting different local regions for identifying different functions through cross-attentions between queries and residue-level features. The development of genome sequencing technologies has unveiled a vast collection of protein sequences, but detailed functional annotations are only available for a very small number of them [2]. Evaluating the functions of protein sequences via wet experiments is time-consuming, labor-intensive, and expensive, underscoring the critical need for computational methods to predict protein functions. This is particularly acute in the study of enzymes, which catalyze various biological reactions and are central to understanding metabolic processes. For the most widely-used EC number classification scheme, each class of enzyme function is assigned an EC number, which is a four-level hierarchy reflecting the intricate organization of enzyme functions.


Enzyme Function Prediction from Amino Acid Sequence: How AI is Leading the Way - CBIRT

#artificialintelligence

An innovative artificial intelligence program called CLEAN (contrastive learning–enabled enzyme annotation) has the ability to predict enzyme activities based on their amino acid sequences, even if the enzymes are unfamiliar or inadequately understood. The researchers have reported that CLEAN has surpassed the most advanced tools in terms of precision, consistency, and sensitivity. However, a deeper understanding of enzymes and their roles would be beneficial in a number of disciplines, including genetics, chemistry, pharmaceuticals, medicine, and industrial materials. The scientists are using the protein language to forecast their performance, similar to how ChatGPT uses written language data to generate predictive phrases. Almost all scientists desire to comprehend the purpose of a protein as soon as they encounter a new protein sequence.