Reinforcement Learning and Generative AI for Next-Generation Protein Degradation Drugs
Hämtar...
Ladda ner
Publicerad
Författare
Typ
Examensarbete för masterexamen
Master's Thesis
Master's Thesis
Modellbyggare
Tidskriftstitel
ISSN
Volymtitel
Utgivare
Sammanfattning
Targeted protein degradation using Proteolysis Targeting Chimeras (PROTACs) has
emerged as a new paradigm in drug discovery, enabling event-driven pharmacology
targeting disease-causing proteins previously considered undruggable. However, due
to the complex ternary structure of PROTACs comprising a POI warhead, E3 ligand,
and a linker, optimizing them for degradation presents a difficult multi-objective optimization task. This is because minor structural changes can significantly impact
the degradation potency (DC50) and maximum degradation level (Dmax). This thesis
introduces a pipeline for de novo generation of text-based SMILES representations
of PROTACs optimized for degradation of the Androgen receptor (AR) using the
Cereblon (CRBN) E3 ligase, and the LNCaP cell line. The thesis implements and
evaluates two deep learning architectures: a long short-term memory (LSTM)-based
recurrent neural network (RNN) trained using the REINVENT4 framework, and
an autoregressive Transformer-based model. For both models, the pipeline utilizes
multi-stage transfer learning (TL) across synthetic and curated PROTAC data, after
which reinforcement learning (RL) is applied, driven by rewards given by the ensemble of TACK surrogate degradation predictors. To rigorously test the generalization
capabilities of both models, they were trained using three increasingly restricting
datasets: (1) a dataset where data with an exact scaffold match to any PROTAC
in a held-out set were removed, (2) identical to (1), with additional removal of data
having a Tanimoto similarity > 0.4 to any PROTAC scaffold in the held-out set, and
(3) identical to condition (2), with the additional removal of all data associated with
the target POI (AR). In addition to the multi-stage training, the pipeline includes
a data-processing step and an evaluation step.
Out of 10,000 generated PROTACS, the fully trained LSTM on the largest dataset
achieved 97.9% validity and 99.8% novelty. Furthermore, 9222 out of the 10,000
sampled SMILES were valid, unique, and canonical, of which 96.5% had predicted
DC50 < 100 nM, 85.6% had predicted Dmax > 80%, and 91.7% were predicted to
be active degraders for the target POI according to surrogate degradation predictor
models. Additionally, the fully trained Transformer trained on the largest dataset
achieved 95.6% validity and 100% novelty. 9112 out of the 10,000 sampled SMILES
were valid, unique, and canonical, of which 99.4% had predicted DC50 < 100 nM,
98.6% had predicted Dmax > 80%, and 94.9% were predicted to be active degraders
for the target POI according to surrogate degradation predictor models. These
results indicate that generative AI pipelines utilizing reinforcement learning can
teach models to discover highly potent and novel PROTACs under multiple complex
objectives.
Beskrivning
Ämne/nyckelord
PROTAC, generative AI, drug discovery, targeted protein degradation, machine learning, reinforcement learning, transfer learning, data
