pharmacophore
OMTRA: A Multi-Task Generative Model for Structure-Based Drug Design
Dunn, Ian, Toft, Liv, Katz, Tyler, Gupta, Juhi, Shah, Riya, Hettiarachchi, Ramith, Koes, David R.
Structure-based drug design (SBDD) focuses on designing small-molecule ligands that bind to specific protein pockets. Computational methods are integral in modern SBDD workflows and often make use of virtual screening methods via docking or pharmacophore search. Modern generative modeling approaches have focused on improving novel ligand discovery by enabling de novo design. In this work, we recognize that these tasks share a common structure and can therefore be represented as different instantiations of a consistent generative modeling framework. We propose a unified approach in OMTRA, a multi-modal flow matching model that flexibly performs many tasks relevant to SBDD, including some with no analogue in conventional workflows. Additionally, we curate a dataset of 500M 3D molecular conformers, complementing protein-ligand data and expanding the chemical diversity available for training. OMTRA obtains state of the art performance on pocket-conditioned de novo design and docking; however, the effects of large-scale pretraining and multi-task training are modest. All code, trained models, and dataset for reproducing this work are available at https://github.com/gnina/OMTRA
Pharmacophore-constrained de novo drug design with diffusion bridge
Wang, Conghao, Mu, Yuguang, Rajapakse, Jagath C.
Computer-aided drug design (CADD) plays a crucial role in the modern drug discovery procedure. However, conventional CADD approaches such as virtual screening are undertaken to search for the candidates with optimal molecular properties in a vast chemistry library. Although accelerated by the high-throughput technology, this process can still be time-consuming and costly [1] since the relationship between chemical structures and the molecular property of interest is obscure. De novo design, on the other hand, models the chemical space of molecular structures and properties and seeks for the optimal candidates in a directed manner [2] instead of enumerating every possibility, thus facilitating the drug discovery process. Moreover, the flourishing of deep generative models in various domains such as large language models and image synthesis has endowed us an opportunity of applying deep learning to improving de novo drug design algorithms. Generative models including variational autoencoder (VAE) [3], generative adversarial networks (GAN) [4] and denoising diffusion probabilistic models (DDPM) [5], have been successfully adapted for molecular design. Initially, researchers tend to represent drugs with linear notations such as Simplified Molecular-Input Line-Entry System (SMILES) [6] due to its simplicity. Then long-short term memory (LSTM) networks were readily applied to encoding the SMILES notations, and VAE and GAN algorithms were utilized for generation [7, 8, 9, 10]. Such methods, however, suffered from low chemistry validity of generated molecules since the structural information is neglected in SMILES notations.
ShEPhERD: Diffusing shape, electrostatics, and pharmacophores for bioisosteric drug design
Adams, Keir, Abeywardane, Kento, Fromer, Jenna, Coley, Connor W.
Engineering molecules to exhibit precise 3D intermolecular interactions with their environment forms the basis of chemical design. In ligand-based drug design, bioisosteric analogues of known bioactive hits are often identified by virtually screening chemical libraries with shape, electrostatic, and pharmacophore similarity scoring functions. We instead hypothesize that a generative model which learns the joint distribution over 3D molecular structures and their interaction profiles may facilitate 3D interaction-aware chemical design. We specifically design ShEPhERD, an SE(3)-equivariant diffusion model which jointly diffuses/denoises 3D molecular graphs and representations of their shapes, electrostatic potential surfaces, and (directional) pharmacophores to/from Gaussian noise. Inspired by traditional ligand discovery, we compose 3D similarity scoring functions to assess ShEPhERD's ability to conditionally generate novel molecules with desired interaction profiles. We demonstrate ShEPhERD's potential for impact via exemplary drug design tasks including natural product ligand hopping, protein-blind bioactive hit diversification, and bioisosteric fragment merging.
SynthFormer: Equivariant Pharmacophore-based Generation of Molecules for Ligand-Based Drug Design
Jocys, Zygimantas, Willems, Henriette M. G., Farrahi, Katayoun
Drug discovery is a complex and resource-intensive process, with significant time and cost investments required to bring new medicines to patients. Recent advancements in generative machine learning (ML) methods offer promising avenues to accelerate early-stage drug discovery by efficiently exploring chemical space. This paper addresses the gap between in silico generative approaches and practical in vitro methodologies, highlighting the need for their integration to optimize molecule discovery. We introduce SynthFormer, a novel ML model that utilizes a 3D equivariant encoder for pharmacophores to generate fully synthesizable molecules, constructed as synthetic trees. Unlike previous methods, SynthFormer incorporates 3D information and provides synthetic paths, enhancing its ability to produce molecules with good docking scores across various proteins. Our contributions include a new methodology for efficient chemical space exploration using 3D information, a novel architecture called Synthformer for translating 3D pharmacophore representations into molecules, and a meaningful embedding space that organizes reagents for drug discovery optimization. Synthformer generates molecules that dock well and enables effective late-stage optimization restricted by synthesis paths.
PharmacoMatch: Efficient 3D Pharmacophore Screening through Neural Subgraph Matching
Rose, Daniel, Wieder, Oliver, Seidel, Thomas, Langer, Thierry
The increasing size of screening libraries poses a significant challenge for the development of virtual screening methods for drug discovery, necessitating a re-evaluation of traditional approaches in the era of big data. Although 3D pharmacophore screening remains a prevalent technique, its application to very large datasets is limited by the computational cost associated with matching query pharmacophores to database ligands. In this study, we introduce PharmacoMatch, a novel contrastive learning approach based on neural subgraph matching. Our method reinterprets pharmacophore screening as an approximate subgraph matching problem and enables efficient querying of conformational databases by encoding query-target relationships in the embedding space. We conduct comprehensive evaluations of the learned representations and benchmark our method on virtual screening datasets in a zero-shot setting. Our findings demonstrate significantly shorter runtimes for pharmacophore matching, offering a promising speed-up for screening very large datasets.
Compositional Deep Probabilistic Models of DNA Encoded Libraries
Chen, Benson, Sultan, Mohammad M., Karaletsos, Theofanis
DNA-Encoded Library (DEL) has proven to be a powerful tool that utilizes combinatorially constructed small molecules to facilitate highly-efficient screening assays. These selection experiments, involving multiple stages of washing, elution, and identification of potent binders via unique DNA barcodes, often generate complex data. This complexity can potentially mask the underlying signals, necessitating the application of computational tools such as machine learning to uncover valuable insights. We introduce a compositional deep probabilistic model of DEL data, DEL-Compose, which decomposes molecular representations into their mono-synthon, di-synthon, and tri-synthon building blocks and capitalizes on the inherent hierarchical structure of these molecules by modeling latent reactions between embedded synthons. Additionally, we investigate methods to improve the observation models for DEL count data such as integrating covariate factors to more effectively account for data noise. Across two popular public benchmark datasets (CA-IX and HRP), our model demonstrates strong performance compared to count baselines, enriches the correct pharmacophores, and offers valuable insights via its intrinsic interpretable structure, thereby providing a robust tool for the analysis of DEL data.
Interpretable Deep Learning in Drug Discovery
Preuer, Kristina, Klambauer, Günter, Rippmann, Friedrich, Hochreiter, Sepp, Unterthiner, Thomas
Without any means of interpretation, neural networks that predict molecular properties and bioactivities are merely black boxes. We will unravel these black boxes and will demonstrate approaches to understand the learned representations which are hidden inside these models. We show how single neurons can be interpreted as classifiers which determine the presence or absence of pharmacophore- or toxicophore-like structures, thereby generating new insights and relevant knowledge for chemistry, pharmacology and biochemistry. We further discuss how these novel pharmacophores/toxicophores can be determined from the network by identifying the most relevant components of a compound for the prediction of the network. Additionally, we propose a method which can be used to extract new pharmacophores from a model and will show that these extracted structures are consistent with literature findings. We envision that having access to such interpretable knowledge is a crucial aid in the development and design of new pharmaceutically active molecules, and helps to investigate and understand failures and successes of current methods.