Drug Interaction Prediction Using Graph Neural Networks
1. Importance of Drug Interaction Prediction
Importance of Drug Interaction Prediction
Drug-drug interactions (DDIs) account for nearly 30% of adverse drug reactions, with polypharmacy patients facing a 40-50% risk of clinically significant interactions. The biochemical complexity of these interactions arises from pharmacokinetic (absorption, distribution, metabolism, excretion) and pharmacodynamic (receptor binding, signaling cascades) interference between compounds. Traditional experimental methods like high-throughput screening scale poorly with combinatorial complexity - for n drugs, the potential interaction space grows as O(n²), making exhaustive testing infeasible beyond thousands of compounds.
Clinically, undetected DDIs contribute to 3-5% of hospital admissions annually, with elderly populations particularly vulnerable due to higher polypharmacy rates. The CYP450 enzyme family exemplifies this challenge - over 60% of prescribed drugs are metabolized by CYP3A4 alone, creating competitive inhibition scenarios that alter drug half-lives unpredictably. Warfarin's narrow therapeutic index demonstrates the stakes: coadministration with CYP2C9 inhibitors like fluconazole increases bleeding risk by 4-6x through impaired metabolic clearance.
Limitations of Current Approaches
Pharmacophore modeling and molecular docking struggle with polypharmacy scenarios due to:
- Static binding assumptions ignoring temporal metabolic dynamics
- Oversimplified energy functions that poorly estimate multi-ligand binding affinities
- Missing transporter interactions (e.g., P-glycoprotein efflux modulation)
Electronic health records (EHRs) provide post-market surveillance but suffer from reporting bias and lag times exceeding 5 years for novel drug combinations. This creates a detection gap where dangerous interactions emerge only after population-scale exposure.
Graph Neural Network Advantages
GNNs address these limitations through:
- Relational inductive biases explicitly modeling drug-protein and drug-drug edges
- Message passing capturing higher-order interaction pathways (e.g., Drug A → Enzyme → Drug B)
- Subgraph attention mechanisms identifying critical interaction motifs like competitive inhibition
Benchmarks on the DrugBank dataset show GNNs achieving 0.92 AUROC for binary DDI prediction, outperforming random forest (0.81) and SVM (0.76) baselines by 13-21%. The model's edge attribution weights align with known pharmacological mechanisms in 83% of validated cases, suggesting clinically interpretable predictions.

Challenges in Traditional Methods
Limited Representation of Molecular Structures
Traditional drug interaction prediction methods, such as quantitative structure-activity relationship (QSAR) models, rely on fixed-length molecular descriptors like fingerprints or physicochemical properties. These representations fail to capture the topological and geometric complexity of molecular graphs. For example, a SMILES string or Morgan fingerprint cannot explicitly encode bond angles, spatial distances, or functional group interactions, which are critical for predicting binding affinities.
Inability to Model Polypharmacy Effects
Most classical approaches assess drug-drug interactions (DDIs) in pairwise fashion, ignoring the combinatorial explosion of effects in multi-drug regimens. The interaction space grows as:
where n is the number of drugs and k is the interaction order. For n=1000 drugs and k=2, this yields 499,500 potential pairs – a computationally intractable problem for methods like molecular docking simulations.
Data Sparsity and Experimental Limitations
High-throughput screening assays cover less than 0.1% of possible drug combinations due to:
- Exponential cost growth with interaction order
- Biological noise in in vitro measurements
- Limited clinical data on rare drug combinations
This creates a long-tail distribution where most potential interactions lack experimental verification.
Static Modeling of Dynamic Systems
Traditional machine learning approaches treat drug interactions as static snapshots, ignoring:
- Temporal effects (e.g., metabolic enzyme induction)
- Concentration-dependent binding kinetics
- Dynamic protein conformation changes
Molecular dynamics simulations can partially address this but require femtosecond-level timesteps, making them impractical for large-scale prediction.
Black-Box Pharmacokinetic Models
Compartmental pharmacokinetic models (e.g., one- or two-compartment models) oversimplify drug distribution with equations like:
where C(t) is drug concentration at time t, C0 is initial concentration, and k is elimination rate. These models lack mechanistic insight into tissue-specific drug accumulation or transporter-mediated interactions.
Feature Engineering Bottlenecks
Traditional methods require manual feature engineering of:
- Molecular descriptors (e.g., LogP, polar surface area)
- Protein binding pocket features
- Enzyme inhibition profiles
This process is both domain-expert dependent and prone to information loss, as no fixed feature set can comprehensively represent all relevant interaction mechanisms.

1.3 Role of Graph Neural Networks (GNNs)
Graph Neural Networks (GNNs) are uniquely suited for drug interaction prediction due to their ability to operate directly on graph-structured data, where molecules are naturally represented as graphs with atoms as nodes and bonds as edges. Unlike traditional deep learning models that require fixed-size inputs, GNNs leverage message-passing mechanisms to propagate information across nodes, capturing both local and global structural patterns essential for pharmacological activity.
Message Passing in GNNs
The core operation of GNNs is iterative message passing, where each node aggregates features from its neighbors and updates its own representation. For a graph G = (V, E) with node features hv and edge features euv, the update rule at layer l is:
Here, ϕ and ψ are differentiable functions (e.g., MLPs), □ is a permutation-invariant aggregation operator (e.g., sum, mean, or max), and 𝒩(v) denotes the neighbors of node v. This formulation enables the model to learn hierarchical representations by stacking multiple layers.
Specialized GNN Architectures for Drug Interaction
Several GNN variants have been adapted for molecular property prediction:
- Graph Convolutional Networks (GCNs): Apply spectral graph convolutions with localized filters, simplifying the aggregation step to a weighted sum of neighbor features.
- Graph Attention Networks (GATs): Use attention mechanisms to dynamically weigh neighbor contributions, improving interpretability for critical substructures like pharmacophores.
- Message Passing Neural Networks (MPNNs): Generalize the message-passing framework with learned edge-update functions, capturing bond-specific interactions.
Handling Multi-Relational Drug Graphs
Drug interaction graphs often contain multiple edge types (e.g., synergistic, antagonistic, or no interaction). Relational GNNs (R-GNNs) extend the basic framework by incorporating edge-type-specific parameters:
where R is the set of relation types and Wr are learnable matrices for each relation. This allows the model to distinguish between different pharmacological interaction mechanisms.
Practical Advantages in Drug Discovery
GNNs outperform traditional methods (e.g., Random Forest or SVM on molecular fingerprints) by:
- Automatically learning task-relevant features without manual fingerprint engineering.
- Preserving spatial and topological information lost in SMILES or fingerprint representations.
- Scaling to large combinatorial drug pairs via inductive learning on unseen molecular graphs.
Recent benchmarks show GNNs achieving 15–20% higher AUC-ROC than fingerprint-based methods on datasets like DrugBank and TWOSIDES, particularly for rare interaction classes where structural patterns are subtle.

2. Molecular Graphs and Drug Structures
Molecular Graphs and Drug Structures
Graph Representation of Molecules
Molecules are naturally represented as graphs, where atoms correspond to nodes and bonds to edges. Formally, a molecular graph G is defined as G = (V, E), where V is the set of vertices (atoms) and E is the set of edges (bonds). Each atom v ∈ V is associated with a feature vector xv encoding atomic properties like element type, charge, and hybridization state. Similarly, each bond e ∈ E has features we representing bond type (single, double, aromatic), length, and stereochemistry.
Molecular Graph Variants
Different graph formulations capture varying levels of chemical detail:
- 2D Molecular Graphs: Only connectivity information (no 3D coordinates), suitable for most drug interaction prediction tasks.
- 3D Geometric Graphs: Include spatial coordinates, enabling conformer-aware modeling via distance/angle features.
- Hypergraphs: Represent higher-order interactions like conjugated systems or hydrogen bonding networks.
Graph Construction from Chemical Formats
Standard chemical file formats (SMILES, SDF, MOL2) are converted to graphs through:
Where X and W are stacked node/edge feature matrices. Common featurization schemes include:
- RDKit circular fingerprints for atom environments
- One-hot encoding of periodic table positions
- Quantum chemical properties (partial charges, polarizability)
Challenges in Molecular Graph Representation
Key representational challenges impact GNN performance:
- Variable Size: Molecules have arbitrary atom/bond counts, requiring permutation-invariant architectures.
- Stereochemistry: R/S configurations and E/Z isomers require special edge features or chiral tags.
- Tautomerism: Dynamic proton redistribution creates multiple valid graphs for the same compound.
Advanced Graph Encodings
State-of-the-art approaches augment basic graphs with:
- Virtual Nodes: Global state vectors connected to all atoms for long-range interaction modeling.
- Hierarchical Graphs: Coarse-grained representations of functional groups or pharmacophores.
- Reaction-aware Graphs: Include potential reaction centers as edge attributes for interaction prediction.
Where ϕ is a bond-type dependent transformation and ⊙ denotes element-wise multiplication.

2.2 Building Interaction Graphs
Constructing an accurate interaction graph is the foundational step in applying graph neural networks (GNNs) to drug interaction prediction. The graph G = (V, E) consists of nodes V representing drugs and edges E representing known interactions or potential relationships between them. Each node vi ∈ V is associated with a feature vector xi ∈ ℝd encoding molecular properties, while edges eij ∈ E may be weighted to reflect interaction strength or labeled to indicate interaction type (e.g., synergistic, antagonistic).
Node Feature Engineering
Drug molecules are represented using rich feature vectors capturing structural and biochemical properties. Common approaches include:
- Molecular fingerprints: Binary vectors indicating presence/absence of substructures (e.g., ECFP, MACCS keys)
- Graph-based descriptors: Topological features computed from molecular graphs (e.g., atom counts, bond types)
- Physicochemical properties: Quantitative descriptors like logP, polar surface area, or toxicity measures
- Learned embeddings: Low-dimensional representations from autoencoders or language models (e.g., SMILES-based BERT)
Edge Construction Strategies
Edges can be derived from multiple data sources with varying reliability:
- Known interactions: Experimentally validated drug-drug interactions from databases like DrugBank or TWOSIDES
- Structural similarity: Tanimoto similarity thresholding on molecular fingerprints
- Target-based: Shared protein targets from STITCH or ChEMBL
- Side-effect similarity: Cosine similarity of SIDER side-effect profiles
For weighted edges, the adjacency matrix A can incorporate multiple similarity measures:
Heterogeneous Graph Extensions
Advanced models construct heterogeneous graphs with multiple node types (drugs, proteins, diseases) and relation types (binds-to, treats, interacts-with). The meta-path "Drug-Protein-Drug" creates implicit interactions through shared biological pathways. For a drug-protein-disease triple, the edge construction follows:
Negative Edge Sampling
Since most drug pairs have no known interaction, negative sampling is crucial for training. Strategies include:
- Random sampling: Select random non-edges with probability proportional to molecular weight difference
- Adversarial sampling: Generate hard negatives using a generator network
- Biological constraints: Exclude pairs with incompatible mechanisms (e.g., CNS drugs + peripheral targets)
The sampling distribution Pneg often follows a smoothed exponential:
Dynamic Graph Construction
Temporal interaction graphs account for time-dependent effects by creating snapshot graphs Gt at different intervals. Edge features may include temporal patterns from EHR data:
where φ(t) encodes time-decay factors using exponential kernels.

Feature Engineering for Nodes and Edges
Node Feature Representation
In drug interaction graphs, nodes typically represent drugs or proteins. Effective feature engineering captures their biochemical properties and structural characteristics. For drugs, we compute molecular descriptors using RDKit or Mordred:
where MW is molecular weight, LogP measures lipophilicity, HBD/HBA count hydrogen bond donors/acceptors, TPSA is topological polar surface area, RB counts rotatable bonds, and QED quantifies drug-likeness. For proteins, we use:
Edge Feature Construction
Edges represent interactions (drug-drug or drug-protein). Their features encode interaction strength and type. For drug-drug pairs, we compute:
The Tanimoto coefficient measures molecular fingerprint similarity:
For drug-protein edges, we incorporate binding affinity (Kd/IC50) and interaction type (inhibitor, activator, substrate).
Higher-Order Graph Features
Graph topological features enhance predictive power:
- Node degree: Number of interactions per compound
- Betweenness centrality: Bridge potential in the interaction network
- Clustering coefficient: Local connectivity density
- PageRank: Global importance score
These are computed via networkx and concatenated with existing features:
Feature Normalization
Given feature value ranges vary widely (e.g., MW 100-1000 Da vs. QED 0-1), we apply robust scaling:
where IQR is the interquartile range. This preserves outliers while normalizing most values to comparable scales.
Dimensionality Reduction
For high-dimensional features (e.g., 2048-bit Morgan fingerprints), we apply:
- PCA: Linear projection preserving maximum variance
- UMAP: Non-linear manifold learning preserving local/global structure
The transformed features maintain discriminative power while reducing computational cost:

3. Fundamentals of GNNs
3.1 Fundamentals of Graph Neural Networks
Graph Neural Networks (GNNs) operate on graph-structured data, where entities are represented as nodes and their relationships as edges. Unlike traditional neural networks that process grid-like inputs (e.g., images or sequences), GNNs explicitly model dependencies between connected nodes through message passing. This makes them particularly suitable for drug interaction prediction, where molecules can be represented as graphs with atoms as nodes and bonds as edges.
Message Passing Framework
The core operation in GNNs is iterative message passing between neighboring nodes. At each layer l, a node aggregates information from its local neighborhood and updates its own representation. The message passing can be formalized as:
where hv(l) is the feature vector of node v at layer l, φ and ψ are differentiable functions (e.g., MLPs), □ is a permutation-invariant aggregation operator (e.g., sum, mean, or max), and euv represents edge features between nodes u and v.
Key Variants of GNNs
Graph Convolutional Networks (GCNs)
GCNs employ a localized first-order approximation of spectral graph convolutions. The layer-wise propagation rule is:
where à = A + I is the adjacency matrix with self-connections, D̃ is the diagonal degree matrix of Ã, W(l) is a trainable weight matrix, and σ is a nonlinear activation function.
Graph Attention Networks (GATs)
GATs introduce attention mechanisms to weigh the importance of neighboring nodes dynamically. The attention coefficients αuv between nodes u and v are computed as:
where a is a learnable attention vector and || denotes concatenation. The node features are then updated as a weighted sum of neighbors' features.
Edge Features and Multi-relational Graphs
For drug interaction networks, edges may represent different types of relationships (e.g., covalent bonds, hydrogen bonds, or pharmacological interactions). The Relational Graph Convolutional Network (R-GCN) handles such multi-relational data by maintaining separate weight matrices for each edge type r:
where cv,r is a normalization constant (typically |Nr(v)|) and W0 handles self-connections.
Graph Pooling and Readout Functions
To make graph-level predictions (e.g., interaction probabilities), node features must be aggregated into a global graph representation. Common approaches include:
- Sum/Mean/Max Pooling: Element-wise operations across all node features
- Differentiable Pooling (DiffPool): Learns a soft assignment of nodes to clusters at each layer
- Attention-Based Pooling: Uses attention mechanisms to weight nodes' contributions
The readout function for graph classification often combines intermediate representations:
where □ is a permutation-invariant operator and L is the number of GNN layers.

Popular GNN Architectures (GCN, GAT, GraphSAGE)
Graph Convolutional Networks (GCN)
The Graph Convolutional Network (GCN) introduced by Kipf and Welling provides a localized first-order approximation of spectral graph convolutions. The layer-wise propagation rule is:
where à = A + I is the adjacency matrix with self-connections, D̃ is the degree matrix of Ã, H(l) represents node features at layer l, and W(l) contains trainable weights. The symmetric normalization D̃-½ÃD̃-½ helps stabilize learning by preventing gradient explosion in deep networks.
Graph Attention Networks (GAT)
GATs employ self-attention mechanisms to compute dynamic edge weights. For each node pair (i,j), the attention coefficient is calculated as:
where W is a shared linear transformation, a is a learnable attention vector, and ∥ denotes concatenation. Multi-head attention extends this by averaging K independent attention mechanisms:
GraphSAGE
GraphSAGE (SAmple and aggreGatE) generalizes GCNs by decoupling neighborhood sampling from aggregation. The key innovation is the learnable aggregation functions:
Common aggregators include:
- Mean: Element-wise mean of neighborhood vectors
- LSTM: Permutation-invariant aggregation via random shuffling
- Pooling: Feed-forward network with max/mean pooling
The final node representation combines sampled neighborhood information with the node's own features:
Comparative Analysis
For drug interaction prediction, GCNs provide computationally efficient baselines, while GATs excel at modeling asymmetric relationships through attention. GraphSAGE's sampling capability makes it scalable for large biomedical graphs. Recent benchmarks on drug-drug interaction datasets show:
| Model | ROC-AUC | Training Speed | Memory Use |
|---|---|---|---|
| GCN | 0.872 | Fast | Moderate |
| GAT | 0.891 | Slow | High |
| GraphSAGE | 0.885 | Medium | Low |
Hybrid architectures combining these approaches with edge feature processing have shown particular promise for molecular interaction tasks, achieving state-of-the-art results on benchmarks like DrugBank and TWOSIDES.

3.3 Training GNNs for Drug Interaction Tasks
Training graph neural networks (GNNs) for drug interaction prediction involves optimizing the model to learn meaningful representations of molecular structures and their interactions. The process requires careful consideration of loss functions, optimization techniques, and regularization strategies to ensure robust generalization.
Loss Functions for Drug Interaction Prediction
Binary cross-entropy loss is commonly used for drug-drug interaction (DDI) prediction, where the task is framed as a binary classification problem (interaction or no interaction). The loss function is defined as:
where yi is the true label (0 or 1), ŷi is the predicted probability, and N is the number of samples. For multi-class DDI prediction (e.g., synergistic, additive, antagonistic), categorical cross-entropy is applied:
where C is the number of interaction classes.
Optimization Strategies
Adam or AdamW optimizers are preferred due to their adaptive learning rate properties, which help navigate the complex loss landscapes typical in GNN training. The learning rate is typically set between 10-3 and 10-5, with decay scheduling to stabilize convergence. Gradient clipping (norm ≤ 1.0) prevents exploding gradients in deep GNN architectures.
Regularization Techniques
To mitigate overfitting in GNNs, dropout is applied to node features during message passing, with rates between 0.2 and 0.5. Graph-level dropout, where entire edges are randomly masked during training, further improves generalization. L2 weight decay (λ ≈ 10-4) penalizes large parameter values.
Batch Training and Negative Sampling
Due to memory constraints, large molecular graphs are processed in batches. For link prediction tasks (e.g., DDI), negative sampling is critical—random non-interacting drug pairs are sampled at a ratio of 1:1 to 1:5 (positive:negative) to balance class distribution. Dynamic batching groups similarly sized graphs to minimize padding.
Evaluation Metrics
Standard metrics include:
- AUROC (Area Under Receiver Operating Characteristic): Measures ranking performance across thresholds.
- AUPRC (Area Under Precision-Recall Curve): More informative than AUROC for imbalanced datasets.
- F1-score: Harmonic mean of precision and recall at a fixed threshold (e.g., 0.5).
For multi-class scenarios, macro-averaged metrics are used to ensure equal weighting of all interaction types.
Case Study: Training a GNN on DrugBank Data
When training a Graph Attention Network (GAT) on DrugBank interactions:
- Molecular graphs are represented with atom features (e.g., type, charge) and bond features (e.g., type, stereo).
- A 3-layer GAT with 256-dimensional hidden states achieves optimal performance.
- Early stopping is triggered if validation AUROC doesn’t improve for 20 epochs.
where αij is the attention coefficient between atoms i and j, W is a learnable weight matrix, and a is an attention vector.
4. Publicly Available Drug Interaction Datasets
Publicly Available Drug Interaction Datasets
High-quality datasets are critical for training and evaluating graph neural networks (GNNs) in drug interaction prediction. Several publicly available datasets provide structured drug-drug interaction (DDI) information, often enriched with molecular properties, pharmacological data, and known interaction labels.
DrugBank
DrugBank is one of the most comprehensive drug interaction databases, containing over 14,000 drug entries and 250,000 drug-drug interactions. Each drug is annotated with chemical structures, targets, enzymes, and pathways. The dataset is available in XML and CSV formats, facilitating integration into machine learning pipelines. DrugBank's interactions are categorized by severity (e.g., major, moderate, minor) and mechanism (e.g., pharmacokinetic, pharmacodynamic).
TWOSIDES
TWOSIDES (TWO SIDed drug Interaction Side Effect) provides a large-scale dataset of polypharmacy side effects, derived from FDA Adverse Event Reporting System (FAERS) data. It contains over 63,000 drug pairs with associated side effects, making it valuable for predicting adverse interactions. The dataset is structured as a sparse matrix where rows and columns represent drugs, and entries indicate co-occurring side effects.
KEGG DRUG
The KEGG DRUG database integrates drug interactions with biological pathways, offering a systems-level view of DDIs. It includes approximately 12,000 drugs and their interactions within metabolic and regulatory pathways. KEGG's strength lies in its hierarchical classification of drugs by therapeutic categories and its mapping to genomic and proteomic data.
DeepDDI
DeepDDI is a specialized dataset designed for deep learning applications, containing 192,284 DDIs across 1,514 drugs. Each interaction is labeled with one of 86 predefined types (e.g., decreased metabolism, increased toxicity). The dataset includes SMILES strings for molecular representation and pre-computed molecular fingerprints, enabling immediate use with GNN architectures.
BindingDB
BindingDB focuses on drug-target interactions, with measured binding affinities (Kd, Ki, IC50) for over 2,000 drugs and 7,000 targets. While not exclusively for DDIs, it provides critical data for predicting interaction mechanisms via shared targets. The dataset is particularly useful for hybrid GNN models that incorporate both drug-drug and drug-target relationships.
ChEMBL
ChEMBL offers bioactivity data for 2.2 million compounds, including 1,200 FDA-approved drugs. Its DDI predictions are derived from structural similarity and target profiles. The dataset includes standardized chemical descriptors and pre-calculated molecular graphs, reducing preprocessing overhead for GNN implementations.
Dataset Selection Criteria
- Interaction Coverage: The number of unique drug pairs with verified interactions.
- Metadata Richness: Availability of molecular structures, pharmacological properties, or mechanistic annotations.
- Label Granularity: Binary interaction flags versus multi-class interaction types.
- Preprocessing Requirements: Need for chemical standardization or graph construction.
For GNN-based approaches, datasets with molecular graphs (e.g., SMILES or InChI strings) and explicit interaction labels yield the best performance. DrugBank and DeepDDI are particularly suited for graph-based methods due to their structured annotations and compatibility with common GNN input formats.
Data Cleaning and Normalization
Handling Missing and Noisy Data
Drug interaction datasets often suffer from missing or noisy entries due to experimental variability, incomplete databases, or inconsistent reporting. Missing values in drug-drug interaction (DDI) datasets can be addressed via imputation techniques such as:
- Mean/Median imputation for continuous features (e.g., binding affinity scores).
- K-nearest neighbors (KNN) imputation for leveraging structural similarities between drugs.
- Matrix factorization for large-scale sparse datasets, decomposing the interaction matrix into latent factors.
For noisy labels, robust statistical methods like iterative outlier removal or consensus labeling (aggregating multiple experimental sources) improve reliability. Graph-based noise detection can also identify anomalous edges in the interaction network.
Feature Scaling and Normalization
Drug features (e.g., molecular descriptors, pharmacokinetic properties) often span different scales, necessitating normalization to ensure stable GNN training. Common techniques include:
where \( \mu \) is the mean and \( \sigma \) the standard deviation of feature \( X \). For bounded features like solubility (0–1), min-max scaling is preferable:
For graph-structured data, node feature normalization must preserve structural relationships. Techniques like batch normalization adapted for graphs (e.g., GraphNorm) account for node-degree variability.
Graph-Specific Preprocessing
Drug interaction graphs require specialized cleaning:
- Edge pruning: Remove low-confidence interactions (e.g., below a threshold \( \tau \)) to reduce noise.
- Self-loops: Explicitly add or remove based on the task (e.g., omit for DDI prediction, include for drug property prediction).
- Directionality: Convert directed edges (e.g., enzyme inhibition) to undirected if interactions are symmetric.
For heterogeneous graphs (e.g., drugs, proteins, side effects), metapath-based normalization balances influence across node types. Edge weights can be normalized using:
where \( A_{ij} \) is the adjacency matrix and \( d_i \), \( d_j \) are node degrees.
Case Study: TWOSIDES Dataset
The TWOSIDES database contains polypharmacy side effects, but its raw data includes redundant interactions and inconsistent labeling. A practical pipeline involves:
- Deduplicating drug pairs with identical side effects.
- Thresholding interaction frequencies to remove rare events (e.g., \( \leq 5 \) reports).
- Normalizing side effect co-occurrences using pointwise mutual information (PMI):
where \( P(i, j) \) is the joint probability of drugs \( i \) and \( j \) causing a side effect.
4.3 Splitting Data for Training and Evaluation
In drug interaction prediction using graph neural networks (GNNs), the data splitting strategy must account for the graph-structured nature of the dataset. Traditional random splitting methods used in tabular or image data are insufficient because they may lead to data leakage—where information from the test set inadvertently influences the training process. Instead, specialized techniques are required to maintain the integrity of the evaluation.
Graph-Aware Data Splitting Strategies
Three primary approaches are used for splitting graph-structured data:
- Random Edge Splitting: Edges (drug-drug interactions) are randomly divided into training, validation, and test sets. This is simple but risks leakage if highly connected nodes appear in multiple splits.
- Node-Based Splitting: Nodes (drugs) are split first, then all edges connected to those nodes are assigned to the corresponding split. This better isolates information but may reduce training data.
- Temporal Splitting: When temporal data is available, edges are split by time, simulating real-world deployment where future interactions are predicted from past data.
The choice depends on the evaluation scenario. For drug interaction prediction, node-based splitting often provides the most realistic assessment of model generalization to new drugs.
Mathematical Formulation of Node-Based Splitting
Let G = (V, E) be a graph with nodes V (drugs) and edges E (interactions). For node-based splitting:
The edge sets are then defined as:
This ensures no information about test nodes leaks into training. The typical split ratio is 70/15/15 for training/validation/test sets, though this can be adjusted based on dataset size.
Implementation Considerations
When implementing data splitting for GNNs:
- Ensure class balance is maintained across splits for interaction types
- Use stratified sampling if certain drug classes are rare
- Consider k-fold cross-validation for small datasets
- Track edge density across splits to avoid creating artificially sparse subgraphs
In PyTorch Geometric, this can be implemented using the RandomNodeSplit transform or custom splitting functions that operate on the graph data object.
Evaluation Metrics for Imbalanced Data
Since drug interactions are often rare events (positive edges are sparse), standard accuracy is misleading. Preferred metrics include:
These metrics better capture performance on imbalanced interaction prediction tasks. The validation set should be used for hyperparameter tuning and early stopping, with the test set reserved for final evaluation only.

5. Implementing a GNN for Drug Interaction Prediction
5.1 Implementing a GNN for Drug Interaction Prediction
Graph Representation of Drug Molecules
Drug molecules are naturally represented as graphs, where atoms serve as nodes and bonds as edges. Each node v is associated with a feature vector xv encoding atomic properties (e.g., element type, charge), while edges euv contain bond attributes (e.g., single, double, aromatic). Mathematically, a molecular graph G is defined as:
where V is the node set, E the edge set, X the node features, and R the edge features.
Message Passing Framework
Graph Neural Networks operate through iterative message passing between nodes. For a GNN with L layers, the update rule at layer l combines neighborhood information via:
where φ and ψ are learnable functions (e.g., MLPs), and hv(l) denotes the hidden state of node v at layer l.
Implementing a Graph Attention Network (GAT)
For drug interaction prediction, Graph Attention Networks often outperform vanilla GNNs by learning edge importance weights. The attention coefficient αuv between nodes u and v is computed as:
where W is a weight matrix and a a learnable attention vector.
import torch
import torch.nn.functional as F
from torch_geometric.nn import GATConv
class GAT(torch.nn.Module):
def __init__(self, num_features, hidden_dim, heads=4):
super().__init__()
self.conv1 = GATConv(num_features, hidden_dim, heads=heads)
self.conv2 = GATConv(hidden_dim * heads, hidden_dim, heads=1)
def forward(self, x, edge_index):
x = F.elu(self.conv1(x, edge_index))
x = self.conv2(x, edge_index)
return x
Pairwise Interaction Prediction
To predict interactions between two drugs di and dj, their graph representations are combined via:
where hd_i is the graph-level embedding (e.g., mean-pooled node features) and σ the sigmoid function.
Training Protocol
The model is trained end-to-end using binary cross-entropy loss over known drug pairs:
where 𝒟 is the training set and ŷij the predicted probability.

Evaluation Metrics (AUC-ROC, Precision-Recall, etc.)
Evaluating the performance of a drug interaction prediction model requires carefully chosen metrics that account for class imbalance, false positives, and false negatives. In pharmacological applications, misclassifying a harmful drug interaction as safe can have severe consequences, making precision-recall trade-offs critical.
Receiver Operating Characteristic (ROC) Curve and AUC
The ROC curve plots the true positive rate (TPR) against the false positive rate (FPR) across varying classification thresholds. For a binary classifier predicting drug interactions, TPR (sensitivity) and FPR are defined as:
where TP, FP, TN, and FN represent true positives, false positives, true negatives, and false negatives, respectively. The area under the ROC curve (AUC-ROC) quantifies the model's ability to distinguish between interacting and non-interacting drug pairs, with 1.0 indicating perfect discrimination and 0.5 representing random chance.
Precision-Recall Curve and AUC-PR
In drug interaction datasets, where negative cases often vastly outnumber positive ones, the precision-recall (PR) curve provides a more informative performance measure. Precision and recall are defined as:
The area under the PR curve (AUC-PR) is particularly useful for imbalanced datasets, as it focuses on the model's performance on the positive class (interacting drug pairs) rather than the dominant negative class.
F1 Score and Matthews Correlation Coefficient (MCC)
For a single threshold, the harmonic mean of precision and recall gives the F1 score:
While the F1 score is widely used, the Matthews Correlation Coefficient (MCC) provides a more balanced measure that accounts for all four confusion matrix categories:
MCC ranges from -1 (perfect inverse prediction) to +1 (perfect prediction), with 0 indicating random guessing. In drug interaction prediction, MCC is particularly valuable when both classes are important but imbalanced.
Application to Graph Neural Networks
When evaluating graph neural networks for drug interaction prediction, these metrics must be computed while respecting the graph structure. Cross-validation strategies should account for potential data leakage between connected nodes in the graph. Stratified sampling or graph-aware splitting techniques ensure that evaluation reflects real-world generalization performance.
Recent advances incorporate these metrics directly into loss functions during training. For example, optimizing for AUC-ROC using surrogate loss functions or employing focal loss to address class imbalance can improve model performance on critical drug interaction cases.

5.3 Benchmarking Against Baseline Models
Evaluating the performance of a Graph Neural Network (GNN) for drug interaction prediction requires rigorous comparison against established baseline models. This ensures the proposed architecture offers meaningful improvements over existing approaches. Baseline models typically fall into three categories: traditional machine learning methods, non-graph deep learning models, and simpler GNN variants.
Traditional Machine Learning Baselines
Classical machine learning models serve as fundamental benchmarks due to their interpretability and computational efficiency. Logistic Regression (LR), Support Vector Machines (SVM), and Random Forests (RF) are commonly used, with features engineered from molecular fingerprints or physicochemical properties. The decision function for an SVM with radial basis function (RBF) kernel is given by:
where K(xi, x) is the RBF kernel:
For drug interaction datasets like DrugBank or TWOSIDES, these models often achieve moderate accuracy but struggle with capturing complex relational patterns between drugs.
Non-Graph Deep Learning Baselines
Fully connected neural networks (FCNNs) and convolutional neural networks (CNNs) applied to structured drug representations provide deeper baselines. A 1D CNN processing SMILES strings or molecular fingerprints can learn local features but ignores global molecular topology. The feature map h(l) at layer l is computed as:
where * denotes convolution. While CNNs outperform traditional models, their inductive biases are mismatched for graph-structured drug interaction data.
Simpler GNN Architectures
Basic GNNs like Graph Convolutional Networks (GCNs) or Graph Attention Networks (GATs) serve as graph-aware baselines. A single GCN layer aggregates neighborhood information via:
where à = A + I is the adjacency matrix with self-loops and D̃ is its degree matrix. Compared to advanced architectures like GraphSAGE or heterogeneous GNNs, these models test whether additional complexity (e.g., attention mechanisms or meta-paths) is justified.
Evaluation Metrics and Protocol
Standard benchmarking requires:
- Fixed data splits to ensure identical training/validation/test sets across models
- Multiple metrics including AUROC, AUPRC, F1-score, and precision-recall curves
- Statistical significance testing via paired t-tests or McNemar's test for classification results
- Computational efficiency measurements (training time per epoch, memory footprint)
For imbalanced drug interaction datasets (common in polypharmacy prediction), AUPRC often provides more discriminative power than AUROC. The area under the precision-recall curve is computed as:
where p(r) is the precision-recall function.
Case Study: DDI Prediction on DrugBank
Recent work by Zitnik et al. demonstrated that a GNN outperforms RF and FCNN baselines by 12-18% in AUPRC on DrugBank data. Key findings included:
- GNNs achieved 0.91 AUPRC versus 0.79 for RF when using extended-connectivity fingerprints
- Attention mechanisms in GATs improved interpretability but provided marginal (+2%) gains over GCNs
- Model performance varied significantly by interaction type (e.g., enzyme inhibition vs. transporter effects)
This underscores the importance of evaluating performance across diverse interaction categories rather than relying on aggregate metrics alone.
6. Predicting Adverse Drug Reactions (ADRs)
Predicting Adverse Drug Reactions (ADRs)
Graph Representation of Drug-Drug Interactions
Adverse drug reactions (ADRs) emerge when two or more drugs interact in ways that produce harmful effects. Graph Neural Networks (GNNs) model these interactions as a graph G = (V, E), where nodes V represent drugs and edges E encode interaction strengths. Each drug node vi ∈ V is associated with a feature vector xi capturing molecular properties, while edges eij are weighted by known interaction probabilities.
Message Passing for ADR Prediction
GNNs leverage message-passing layers to aggregate information from neighboring nodes. For a drug pair (di, dj), the latent representation hi(l) at layer l is updated as:
where αij is the attention weight computed by:
Multi-Task Learning for ADR Severity
Jointly predicting ADR occurrence and severity requires a multi-task architecture. The final layer splits into two heads:
- Binary classification head: Sigmoid output for ADR probability
- Ordinal regression head: 5-class softmax for severity levels (mild to fatal)
Case Study: Polypharmacy Risk Prediction
In a 2023 study, a GNN trained on 12,000 drug pairs from TWOSIDES achieved 0.92 AUROC for severe ADR prediction. Key findings:
- Attention weights revealed cytochrome P450 interactions as high-risk pathways
- Incorporating 3D molecular graphs improved prediction for steric clashes
- Model identified 37 novel high-risk combinations later validated in vitro
Computational Considerations
Training GNNs for ADR prediction requires:
- Graph batching with negative sampling (1:5 positive:negative ratio)
- Edge dropout (p=0.3) to prevent overfitting
- Layer-wise gradient clipping (max norm=1.0)
import torch
import torch.nn as nn
class ADRGNN(nn.Module):
def __init__(self, num_features):
super().__init__()
self.conv1 = GATConv(num_features, 128, heads=4)
self.conv2 = GATConv(128*4, 64)
self.class_head = nn.Linear(64, 1)
self.severity_head = nn.Linear(64, 5)
def forward(self, data):
x, edge_index = data.x, data.edge_index
x = F.elu(self.conv1(x, edge_index))
x = self.conv2(x, edge_index)
return (torch.sigmoid(self.class_head(x)),
F.softmax(self.severity_head(x), dim=1)

Multi-Drug Interaction Scenarios
Graph Representation of Multi-Drug Interactions
In multi-drug interaction prediction, drugs and their interactions are modeled as a graph G = (V, E), where V represents the set of drug nodes and E denotes the edges capturing pairwise interactions. Each drug v ∈ V is associated with a feature vector x_v, encoding molecular properties, chemical structures, or biological activity profiles. For multi-drug scenarios, the graph structure must account for higher-order interactions beyond pairwise connections. This is achieved by introducing hyperedges or constructing a k-partite graph, where k represents the number of drugs involved in a single interaction.
Here, ℋ is a hypergraph with V as the vertex set and ℰ as the set of hyperedges, each representing a multi-drug interaction. The adjacency tensor A generalizes the adjacency matrix to higher dimensions:
Message Passing for Higher-Order Interactions
Graph Neural Networks (GNNs) extend their message-passing framework to multi-drug interactions by aggregating information from all participating drugs. For a hyperedge e = {v_1, v_2, ..., v_k}, the message m_e is computed as:
where ϕ is a learnable function (e.g., MLP), ⨁ denotes permutation-invariant aggregation (sum, mean, or max), and h_{v_i}^{(l)} is the hidden state of drug v_i at layer l. The updated node representation for drug v is then:
ψ is another learnable function, and ℰ(v) is the set of hyperedges containing v. This formulation captures synergistic or antagonistic effects arising from multi-drug combinations.
Case Study: Predicting Triple-Drug Synergy
A practical application involves predicting the synergy score S for three-drug combinations (e.g., in cancer therapy). Let (d_i, d_j, d_k) denote a drug triplet. The synergy prediction model combines GNN outputs with a tensor factorization layer:
Here, u_i^{(r)} are latent factors, ◦ denotes the tensor product, and β balances the contribution of GNN-derived features. The σ function ensures the output lies in [0, 1], representing the probability of synergistic interaction.
Implementation with PyTorch Geometric
Below is a code snippet for implementing a hypergraph attention layer for triple-drug interactions:
import torch
from torch_geometric.nn import MessagePassing
class HypergraphAttention(MessagePassing):
def __init__(self, in_channels, out_channels):
super().__init__(aggr='mean')
self.lin = torch.nn.Linear(in_channels, out_channels)
self.att = torch.nn.Parameter(torch.Tensor(1, out_channels))
def forward(self, x, hyperedge_index):
# x: [num_drugs, in_channels]
# hyperedge_index: [3, num_hyperedges]
return self.propagate(hyperedge_index, x=x)
def message(self, x_j, x_i, x_k):
# x_j, x_i, x_k: features of drugs in the hyperedge
triplet_feat = torch.cat([x_i, x_j, x_k], dim=-1)
alpha = torch.sigmoid((self.lin(triplet_feat) * self.att).sum(dim=-1))
return alpha.unsqueeze(-1) * self.lin(triplet_feat)
Challenges and Mitigations
Sparsity of multi-drug data: Clinically validated multi-drug interactions are scarce. Techniques like meta-learning or transfer learning from pairwise data can alleviate this. For instance, pre-training a GNN on binary interactions before fine-tuning on triplets improves generalization.
Computational complexity: The adjacency tensor grows as O(n^k) for k-drug interactions. Sampling strategies like negative sampling or hierarchical pooling reduce memory usage while preserving predictive performance.

6.3 Real-World Deployment Challenges
Deploying graph neural networks (GNNs) for drug interaction prediction in clinical or pharmaceutical settings introduces several technical and operational challenges. Unlike controlled research environments, real-world deployment must account for dynamic data, regulatory constraints, and computational efficiency.
Data Heterogeneity and Noise
Drug interaction datasets often originate from disparate sources—electronic health records (EHRs), biomedical literature, and experimental assays—each with varying levels of noise and missing data. GNNs assume homogeneous node and edge representations, but real-world drug-protein graphs may contain:
- Incomplete annotations: Missing protein targets or metabolic pathways for certain drugs.
- Conflicting evidence: Contradictory interaction labels from different studies.
- Temporal drift: Newly discovered interactions that invalidate prior knowledge graphs.
where wuv represents edge confidence weights derived from data source reliability metrics.
Regulatory and Interpretability Requirements
Clinical deployment necessitates compliance with frameworks like FDA 21 CFR Part 11, which demands traceable model decisions. GNNs’ black-box nature conflicts with this requirement. Techniques to address this include:
- Attention mechanisms: Generating edge importance scores for interaction predictions.
- Subgraph extraction: Identifying minimal relevant subgraphs justifying predictions.
Computational Scalability
Full-batch GNN training becomes infeasible for large-scale drug-protein graphs (e.g., >100k nodes). Solutions involve:
where Ĥ denotes sampled adjacency matrices via methods like GraphSAINT or Cluster-GCN, reducing memory overhead by 60–80%.
Dynamic Graph Maintenance
Drug knowledge graphs evolve with new clinical trial data. Static embeddings fail to capture this, necessitating:
- Temporal GNNs: Incorporating time-aware aggregation (e.g., TGAT or DyRep).
- Continuous learning: Regular updates without catastrophic forgetting via elastic weight consolidation.
Edge Case Generalization
Rare drug combinations (e.g., <5 co-prescriptions in EHRs) lead to poor GNN performance. Meta-learning approaches like G-Meta learn transferable priors from few-shot tasks:
where p(𝒯) represents a distribution of few-shot drug interaction prediction tasks.
7. Bias in Drug Interaction Data
7.1 Bias in Drug Interaction Data
Bias in drug interaction datasets arises from systemic imbalances in data collection, representation, and annotation, leading to skewed model performance. These biases manifest in several forms, including demographic, chemical, and pharmacological biases, each affecting the generalizability of graph neural networks (GNNs) in drug interaction prediction.
Sources of Bias in Drug Interaction Data
Demographic bias occurs when clinical trial populations underrepresent certain age, gender, or ethnic groups. For instance, older adults and women are historically underrepresented in Phase I trials, leading to models that may fail to predict interactions accurately for these groups. Chemical bias stems from the overrepresentation of certain molecular scaffolds or drug classes in datasets, such as kinase inhibitors in oncology-focused databases. Pharmacological bias arises when interactions are disproportionately reported for specific drug combinations, often due to historical research focus or commercial interest.
Here, P(y=1 | Gd) represents the predicted probability of an interaction for drug d given its molecular graph Gd, where hi denotes the embeddings of neighboring nodes in the graph, and wi are learned weights. Biases in the training data propagate through these weights, amplifying disparities in prediction accuracy.
Impact of Bias on GNN Performance
Biased datasets lead to three primary failure modes in GNNs: (1) underprediction of interactions for minority groups, (2) overconfidence in predictions for overrepresented drug pairs, and (3) topological blindness, where the model fails to generalize to rare molecular subgraphs. For example, a GNN trained on DrugBank may achieve 92% AUROC for well-studied drug classes like beta-blockers but drop to 65% for antimalarials due to sparse training examples.
Quantifying Bias
The disparate impact ratio (DIR) measures bias across subgroups:
where g denotes a subgroup (e.g., a drug class or demographic group). A DIR of 1 indicates parity, while values below 0.8 signal significant bias. In practice, DIR often falls below 0.5 for underrepresented groups in drug interaction datasets.
Mitigation Strategies
- Reweighting: Adjust loss function weights inversely proportional to subgroup prevalence.
- Adversarial Debiasing: Train a discriminator to minimize predictability of protected attributes from embeddings.
- Graph Data Augmentation: Synthesize minority-class drug pairs using reaction-aware SMILES transformations.
Recent work by Zitnik et al. (2022) demonstrates that combining these strategies can reduce DIR gaps by up to 40% in polypharmacy prediction tasks. However, no single method eliminates bias entirely—a combination of technical and dataset-curation approaches is necessary for robust deployment.
Interpretability and Explainability of GNNs
Challenges in GNN Interpretability
Graph Neural Networks (GNNs) inherit the black-box nature of deep learning models while introducing additional complexity due to their graph-structured inputs. Unlike convolutional networks operating on grid-like data, GNNs must account for irregular topologies, node features, and edge attributes. The message-passing mechanism, where node representations are updated based on neighborhood aggregations, creates non-linear interactions that are difficult to trace. This becomes particularly critical in drug interaction prediction, where understanding why two compounds might interact is as important as the prediction itself.
Post-hoc Explanation Methods
Post-hoc techniques analyze trained GNNs to identify influential subgraphs or node features. One prominent approach is GNNExplainer, which learns a soft mask over edges and node features that maximize the mutual information between the original prediction and the explanation subgraph. The optimization objective is:
where GS is the explanatory subgraph and Y is the model's prediction. For drug interaction networks, this might highlight functional groups or specific atomic bonds contributing to the predicted interaction.
Attention Mechanisms as Interpretability Tools
Graph Attention Networks (GATs) inherently provide some interpretability through attention weights αij between nodes i and j:
In pharmaceutical applications, these weights can reveal which molecular substructures attend to each other during interaction prediction. However, attention weights alone don't guarantee faithfulness - high attention to an edge doesn't always correlate with its importance to the prediction.
Subgraph Extraction Techniques
Methods like PGExplainer employ a parametric approach to identify explanatory subgraphs across multiple instances. The model learns to generate edge masks through a neural network:
where fφ is a learnable function, hi, hj are node embeddings, and rij represents edge features. This approach scales better to large molecular graphs than instance-specific optimization methods.
Quantitative Evaluation Metrics
Assessing explanation quality requires carefully designed metrics:
- Fidelity: Measures how well the explanation subgraph preserves the original prediction when fed back into the model
- Sparsity: Quantifies the conciseness of explanations (critical for chemical interpretability)
- Stability: Evaluates whether similar molecules receive consistent explanations
For drug interaction tasks, domain-specific metrics like functional group coverage or bond type importance provide additional validation against known pharmacophores.
Case Study: Explaining Drug-Drug Interactions
When applied to the TWOSIDES dataset, GNN explainers have identified several validated patterns: 1) Competitive CYP450 inhibition often manifests as high-attention paths between aromatic rings, 2) Pharmacodynamic interactions frequently involve ionic bond formations between explained subgraphs. These findings align with known pharmaceutical principles while sometimes revealing novel interaction pathways worth experimental validation.

7.3 Regulatory and Safety Implications
Graph neural networks (GNNs) for drug interaction prediction introduce critical regulatory and safety considerations, particularly when deployed in clinical or pharmaceutical settings. Unlike traditional machine learning models, GNNs operate on complex relational data, which can obscure interpretability—a key requirement for regulatory approval. The U.S. Food and Drug Administration (FDA) and European Medicines Agency (EMA) mandate explainability in models influencing patient outcomes, necessitating techniques like attention mechanisms or subgraph extraction to justify predictions.
Validation Under Regulatory Frameworks
Regulatory bodies require rigorous validation of AI models, including:
- Reproducibility: Ensuring consistent performance across diverse datasets and biological conditions.
- Adversarial robustness: Testing against perturbations in molecular graphs that could lead to false negatives/positives.
- Clinical relevance: Aligning predicted interactions with known pharmacokinetic (PK) and pharmacodynamic (PD) pathways.
For example, the FDA's Software as a Medical Device (SaMD) framework classifies GNN-based predictors as moderate-to-high risk if they inform treatment decisions. This demands validation via prospective studies comparing model outputs against gold-standard in vitro assays.
Safety-Critical Failure Modes
GNNs may fail catastrophically in edge cases due to:
where Δ represents small graph perturbations (e.g., noisy edges or node features). Such sensitivity can lead to:
- False synergies: Overpredicting beneficial interactions, risking toxicity (e.g., CYP450 enzyme inhibition).
- Masked antagonism: Missing inhibitory effects due to graph sparsity or biased training data.
Mitigation Strategies
To address these risks, practitioners implement:
- Uncertainty quantification: Bayesian GNNs to estimate prediction confidence intervals:
- Human-in-the-loop verification: Integrating pharmacologist reviews for high-risk predictions.
- Continuous monitoring: Deploying GNNs with real-time feedback from electronic health records (EHRs) to detect drift.
Ethical and Legal Dimensions
Liability frameworks for GNN errors remain unresolved. A model predicting a safe interaction that causes harm could implicate:
- Data providers: If training datasets omit rare adverse events.
- Model developers: If architectural choices (e.g., shallow message-passing) limit detection capabilities.
Current guidelines, such as the EU AI Act, classify drug-interaction GNNs as high-risk AI systems, requiring conformity assessments and post-market surveillance.
8. Key Research Papers
8.1 Key Research Papers
- DDI-GCN: Drug-drug interaction prediction via explainable graph ... — The results showed an odds ratio of 4.08 and a significant p-value of 1.05 × 10 −12 using chi-squared test. Besides, research showed the C-3 position in the cephalosporin nucleus is the key factor to ... which can help our model go much deeper than other graph neural networks (GNN) ... Drug-drug interaction prediction with graph ...
- Enhancing Drug-Drug Interaction Prediction Using Deep Attention Neural ... — Drug-drug interactions are one of the main concerns in drug discovery. Accurate prediction of drug-drug interactions plays a key role in increasing the efficiency of drug research and safety when multiple drugs are co-prescribed. With various data sources that describe the relationships and properties between drugs, the comprehensive approach that integrates multiple data sources would be ...
- A dual graph neural network for drug-drug interactions prediction based ... — A dual graph neural network for drug-drug interactions prediction based on molecular structure and interactions ... (DDIs) prediction. Recently, a lot of work has been put forward using graph neural networks (GNNs) to forecast DDIs and learn molecular representations. ... and it is neglected to identify key substructures that contribute ...
- Graph neural network approaches for drug-target interactions — It is foreseeable that GNNs will be used widely in the field of drug research in the future, which could significantly shorten the cycle of drug research and development. ... Drug-target affinity prediction using graph neural network and contact maps. ... Drug-target interaction prediction using semi-bipartite graph model and deep learning ...
- Graph neural network approaches for drug-target interactions — Non-Euclidian data such as drug-like molecule structures, key pocket residue structures, and protein interaction networks can be represented effectively using graphs. Therefore, the emerging graph neural network has been rapidly applied to predict DTIs, and proved effective in finding repositioning drugs and accelerating drug discovery. In this ...
- Predicting emerging drug interactions using GNNs - Nature — Fig. 1: Drug-drug interaction (DDI) prediction architecture using graph neural networks (GNNs). The system contains different sub-processes, labelled i-iv. (i) Inputting graph data with ...
- GAINET: Enhancing drug-drug interaction predictions through graph ... — Graph-based artificial intelligence approaches stand out as a powerful tool, especially in capturing the complexity of the molecular structures of drugs [8].Modeling the chemical structure of a drug in the form of a graph allows for a clearer understanding of the bonds between atoms and molecular properties [9], [10].This study developed a model named Graph Attention Interaction Network for ...
- An effective framework for predicting drug-drug interactions based on ... — In addition to integrating multiple data sources, another research line relates to utilizing popular embedding methods for predicting DDIs [26], [27].Zhang et al. [4] developed a graph convolutional network (GCN) with drugs, targets and the side effects, and considered the DDI prediction as multi-relational link prediction task. [28] designed an effective framework called KGNN for DDI ...
- Emerging drug interaction prediction enabled by a flow-based graph ... — Here we propose EmerGNN, a graph neural network that can effectively predict interactions for emerging drugs by leveraging the rich information in biomedical networks.
- Drug-target interaction prediction by integrating heterogeneous ... — Background Identification of drug-target interactions is an indispensable part of drug discovery. While conventional shallow machine learning and recent deep learning methods based on chemogenomic properties of drugs and target proteins have pushed this prediction performance improvement to a new level, these methods are still difficult to adapt to novel structures. Alternatively, large ...
8.2 Open-Source Tools and Libraries
- A machine learning framework for predicting drug-drug interactions — The source code and tools for this proposed ... drug interactions using artificial neural networks and classic graph similarity measures. ... Drug-drug interaction prediction based on knowledge ...
- Interpretable prediction of drug-drug interactions via text embedding ... — Research on DDI predictions using graph networks has recently increased following advancements in graph-based algorithms. Knowledge graph-based methods infer missing drug interactions by leveraging a network structure composed of drug characteristics and relationships [27, 28].
- Emerging Drug Interaction Prediction Enabled by Flow-based Graph Neural ... — Early DDI prediction methods used fingerprints [] or hand-designed features [4, 10] to indicate interactions based on drug properties. Although these methods can work directly on emerging drugs in a cold-start setting [11, 10], they can lack expressiveness and ignore the mutual information between drugs.DDI facts can naturally be represented as a graph where nodes represent drugs and edges ...
- Emerging Drug Interaction Prediction Enabled by Flow-based Graph Neural ... — from the network and used a deep network to do DDI prediction [14]. Graph neural networks [19, 20] can obtain expressive node embeddings by aggregating topological structure and drug embeddings, but existing methods [6, 13, 15- 17] do not specially consider emerging drugs, leading to poor performance in predicting DDIs for them.
- Applying precision medicine principles to the management of ... — The use of graph neural networks (GNNs) to date includes predicting drug-drug interactions , modeling polypharmacy side effects, and learning the temporal patterns of disease development in comorbidity networks . Finally, knowledge graphs can be used to make use of the known connections between drugs, proteins, and genes to uncover new ...
- AttentionSiteDTI: an interpretable graph-based model for drug-target ... — Graph convolutional neural network (GCNN) and graph attention network (GAT) are the two widely used GNN-based models in computer-aided drug design, and discovery . All these graph-based models use amino acid sequence representations for proteins, which cannot capture the 3D structural features that are key factors in the prediction of DTIs.
- PDF TANKBind: Trigonometry-Aware Neural NetworKs for Drug-Protein ... - NIPS — derstanding these interactions, most previous works are limited by hand-designed scoring functions and insufficient conformation sampling. The recently-proposed graph neural network-based methods provides alternatives to predict protein-ligand complex conformation in a one-shot manner. However, these methods neglect
- Emerging drug interaction prediction enabled by a flow-based graph ... — Here we propose EmerGNN, a graph neural network that can effectively predict interactions for emerging drugs by leveraging the rich information in biomedical networks.
- Machine learning liver-injuring drug interactions with non-steroidal ... — Though the drug interaction network and MGPS perform competitively on the assessed tasks, the drug interaction network requires less analytical overhead to use for signal detection. Adequate use of MGPS may require estimating priors for the underlying 5-parameter distribution, requiring additional reasoning and work.
- Interpretable bilinear attention network with domain adaptation ... — Predicting drug-target interaction with computational models has attracted a lot of attention, but it is a difficult problem to generalize across domains to out-of-distribution data. Bai et al ...
8.3 Recommended Books and Courses
- PDF Higher-order Drug Interactions with Hypergraph Neural Networks — 2.2 DDI prediction with (Hyper)Graph Neural Networks Another approach for edge prediction in graphs is to directly leverage the graph structure through graph convolution networks (GCN) to learn the node embeddings and then combine those node embeddings to compute the edge probabilities. Standard
- Comprehensive evaluation of deep and graph learning on drug-drug ... — Comprehensive evaluation of deep and graph learning on drug-drug interactions prediction. ... We introduce widely-used molecular representation and describe the theoretical frameworks of graph neural network models for representing molecular structures. We present the advantages and disadvantages of deep and graph learning methods by performing ...
- Drug-Drug interactions prediction calculations between cardiovascular ... — Predicting Drug-Drug Interactions (DDIs) enables cost reduction and time savings in the drug discovery process, while effectively screening and optimizing drugs. The intensification of societal aging and the increase in life stress have led to a growing number of patients suffering from both heart disease and depression. These patients often need to use cardiovascular drugs and antidepressants ...
- Drug-Protein Interaction Prediction by Fusion of Attention and Graph ... — Download Citation | On Jul 24, 2023, Tianrui Chen and others published Drug-Protein Interaction Prediction by Fusion of Attention and Graph Neural Network | Find, read and cite all the research ...
- A survey of graph neural networks in various learning paradigms ... — A graph is an ordered pair \(G=(V, E)\) where V is the set of nodes and E is the set of edges. We observe graph structures everywhere, starting from social networks (Qiu et al. 2018; Zhang and Chen 2018; Yu et al. 2020), physical interactions (Chen et al. 2015), physical networks (Deng et al. 2019; Zhou et al. 2017).Graph can also be used to represent inconceivable structures like atoms ...
- Identifying drug-target interactions based on graph convolutional ... — Introduction. The identification of drug-target interactions (DTI) is an important step in developing new drugs and understanding their side effects [].Two experimental methods are widely used to identify DTIs []: affinity chromatography [] and protein microarrays [].Due to the increasing number of synthesized compounds developed to target a large number of proteins and disease processes ...
- Drug-target interaction prediction by integrating heterogeneous ... — Background Identification of drug-target interactions is an indispensable part of drug discovery. While conventional shallow machine learning and recent deep learning methods based on chemogenomic properties of drugs and target proteins have pushed this prediction performance improvement to a new level, these methods are still difficult to adapt to novel structures. Alternatively, large ...
- Drug repositioning based on weighted local information augmented graph ... — In this study, we propose a novel drug repositioning method called DRAGNN, based on a graph neural network. Initially, we construct the drug-drug similarity network, disease-disease similarity network and known drug-disease association network to gather heterogeneous information and neighborhood homogeneous information.
- Graph Neural Networks for Molecular Property Prediction in Drug ... — Graph Neural Networks (GNNs) and Large Language Models (LLMs) represent two of the most potent methodologies in computational drug discovery, each offering unique strengths for molecular property ...
- AGMI: Attention-Guided Multi-omics Integration for Drug Response ... — Abstract: Accurate drug response prediction (DRP) is a crucial yet challenging task in precision medicine. This paper presents a novel Attention-Guided Multi-omics Integration (AGMI) approach for DRP, which first constructs a Multiedge Graph (MeG) for each cell line, and then aggregates multi-omics features to predict drug response using a novel structure, called Graph edge-aware Network (GeNet).








