MARATTO

article · BMC Bioinformatics

Advancing drug–target interaction prediction: a comprehensive graph-based approach integrating knowledge graph embedding and ProtBert pretraining

202336 citationsOpen accessUniversity of Tunis El Manar

In plain language

Validating drug-target interactions through laboratory experiments is costly and time-consuming, meaning only a small fraction are verified empirically. To help accelerate drug discovery, computational approaches are increasingly used to predict these interactions. A computational framework named DTIOG combines knowledge graph embeddings with contextual data derived from protein sequences using a pretrained model known as ProtBERT. The system calculates embedding vectors through knowledge graph techniques and evaluates target similarity alongside representations drawn from chemical structure strings of drugs and amino acid sequences of proteins. When tested on benchmark datasets covering enzymes, ion channels, and G-protein-coupled receptors, the method demonstrated superior predictive accuracy compared to existing algorithms across multiple classifiers and similarity measures. Case study results indicate the tool can assist in identifying previously unknown drug-target interactions.

Key takeaways

  • DTIOG combines knowledge graph embeddings with ProtBERT protein sequence representations to predict drug-target interactions.
  • The framework integrates structural data from chemical SMILES strings and amino acid sequences alongside graph-based embeddings.
  • In evaluations on enzyme, ion channel, and G-protein-coupled receptor datasets, the method outperformed existing algorithms.
  • Case study findings indicate the method can serve as a tool for identifying new drug-target interactions.

Why it matters

Discovering new medicines requires identifying which biological targets interact with specific drug compounds, but testing every combination in a laboratory is prohibitively slow and expensive. Accurate computational prediction tools help researchers filter and prioritise high-potential interactions before physical testing. This reduces the time and experimental resources needed during early pharmaceutical development, helping to make the overall drug discovery pipeline more efficient.

Commercialisation angle

The framework offers an algorithmic tool for pharmaceutical companies, biotechnology firms, and academic screening programmes seeking to shortlist candidate drug-target interactions. Because the method has been assessed computationally on benchmark datasets and an analytical case study, it is an early-stage discovery tool that requires subsequent laboratory validation before direct deployment in commercial drug pipelines.

AI-generated from the published abstract. Always read the original work before citing.

Abstract

BACKGROUND: The pharmaceutical field faces a significant challenge in validating drug target interactions (DTIs) due to the time and cost involved, leading to only a fraction being experimentally verified. To expedite drug discovery, accurate computational methods are essential for predicting potential interactions. Recently, machine learning techniques, particularly graph-based methods, have gained prominence. These methods utilize networks of drugs and targets, employing knowledge graph embedding (KGE) to represent structured information from knowledge graphs in a continuous vector space. This phenomenon highlights the growing inclination to utilize graph topologies as a means to improve the precision of predicting DTIs, hence addressing the pressing requirement for effective computational methodologies in the field of drug discovery. RESULTS: The present study presents a novel approach called DTIOG for the prediction of DTIs. The methodology employed in this study involves the utilization of a KGE strategy, together with the incorporation of contextual information obtained from protein sequences. More specifically, the study makes use of Protein Bidirectional Encoder Representations from Transformers (ProtBERT) for this purpose. DTIOG utilizes a two-step process to compute embedding vectors using KGE techniques. Additionally, it employs ProtBERT to determine target-target similarity. Different similarity measures, such as Cosine similarity or Euclidean distance, are utilized in the prediction procedure. In addition to the contextual embedding, the proposed unique approach incorporates local representations obtained from the Simplified Molecular Input Line Entry Specification (SMILES) of drugs and the amino acid sequences of protein targets. CONCLUSIONS: The effectiveness of the proposed approach was assessed through extensive experimentation on datasets pertaining to Enzymes, Ion Channels, and G-protein-coupled Receptors. The remarkable efficacy of DTIOG was showcased through the utilization of diverse similarity measures in order to calculate the similarities between drugs and targets. The combination of these factors, along with the incorporation of various classifiers, enabled the model to outperform existing algorithms in its ability to predict DTIs. The consistent observation of this advantage across all datasets underlines the robustness and accuracy of DTIOG in the domain of DTIs. Additionally, our case study suggests that the DTIOG can serve as a valuable tool for discovering new DTIs.

Research topics

  • Computational Drug Discovery Methods
  • Bioinformatics and Genomic Networks
  • Biomedical Text Mining and Ontologies

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1186/s12859-023-05593-6

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.