Reinforcement Learning and Generative AI for Next-Generation Protein Degradation Drugs
| dc.contributor.author | Persson, Alexander | |
| dc.contributor.author | Erngård, Felix | |
| dc.contributor.department | Chalmers tekniska högskola / Institutionen för data och informationsteknik | sv |
| dc.contributor.department | Chalmers University of Technology / Department of Computer Science and Engineering | en |
| dc.contributor.examiner | Mercado, Rocio | |
| dc.contributor.supervisor | Ribes, Stefano | |
| dc.date.accessioned | 2026-07-01T12:42:32Z | |
| dc.date.issued | 2026 | |
| dc.date.submitted | ||
| dc.description.abstract | Targeted protein degradation using Proteolysis Targeting Chimeras (PROTACs) has emerged as a new paradigm in drug discovery, enabling event-driven pharmacology targeting disease-causing proteins previously considered undruggable. However, due to the complex ternary structure of PROTACs comprising a POI warhead, E3 ligand, and a linker, optimizing them for degradation presents a difficult multi-objective optimization task. This is because minor structural changes can significantly impact the degradation potency (DC50) and maximum degradation level (Dmax). This thesis introduces a pipeline for de novo generation of text-based SMILES representations of PROTACs optimized for degradation of the Androgen receptor (AR) using the Cereblon (CRBN) E3 ligase, and the LNCaP cell line. The thesis implements and evaluates two deep learning architectures: a long short-term memory (LSTM)-based recurrent neural network (RNN) trained using the REINVENT4 framework, and an autoregressive Transformer-based model. For both models, the pipeline utilizes multi-stage transfer learning (TL) across synthetic and curated PROTAC data, after which reinforcement learning (RL) is applied, driven by rewards given by the ensemble of TACK surrogate degradation predictors. To rigorously test the generalization capabilities of both models, they were trained using three increasingly restricting datasets: (1) a dataset where data with an exact scaffold match to any PROTAC in a held-out set were removed, (2) identical to (1), with additional removal of data having a Tanimoto similarity > 0.4 to any PROTAC scaffold in the held-out set, and (3) identical to condition (2), with the additional removal of all data associated with the target POI (AR). In addition to the multi-stage training, the pipeline includes a data-processing step and an evaluation step. Out of 10,000 generated PROTACS, the fully trained LSTM on the largest dataset achieved 97.9% validity and 99.8% novelty. Furthermore, 9222 out of the 10,000 sampled SMILES were valid, unique, and canonical, of which 96.5% had predicted DC50 < 100 nM, 85.6% had predicted Dmax > 80%, and 91.7% were predicted to be active degraders for the target POI according to surrogate degradation predictor models. Additionally, the fully trained Transformer trained on the largest dataset achieved 95.6% validity and 100% novelty. 9112 out of the 10,000 sampled SMILES were valid, unique, and canonical, of which 99.4% had predicted DC50 < 100 nM, 98.6% had predicted Dmax > 80%, and 94.9% were predicted to be active degraders for the target POI according to surrogate degradation predictor models. These results indicate that generative AI pipelines utilizing reinforcement learning can teach models to discover highly potent and novel PROTACs under multiple complex objectives. | |
| dc.identifier.coursecode | DATX05 | |
| dc.identifier.uri | https://hdl.handle.net/20.500.12380/311764 | |
| dc.language.iso | eng | |
| dc.setspec.uppsok | Technology | |
| dc.subject | PROTAC, generative AI, drug discovery, targeted protein degradation, machine learning, reinforcement learning, transfer learning, data | |
| dc.title | Reinforcement Learning and Generative AI for Next-Generation Protein Degradation Drugs | |
| dc.type.degree | Examensarbete för masterexamen | sv |
| dc.type.degree | Master's Thesis | en |
| dc.type.uppsok | H | |
| local.programme | Data science and AI (MPDSC), MSc | |
| local.programme | Computer science -algorithms, languages and logic (MPALG), MSc |
