Evolution Strategies as an Alternative to Reinforcement Learning in De Novo Molecular Generation - Evaluating Distribution Based Gaussian Evolution Strategies in REINVENT vs. Policy Based Reinforcement Learning on the Practical Molecular Optimization Benchmark
| dc.contributor.author | Eklöv, Malva | |
| dc.contributor.department | Chalmers tekniska högskola / Institutionen för data och informationsteknik | sv |
| dc.contributor.department | Chalmers University of Technology / Department of Computer Science and Engineering | en |
| dc.contributor.examiner | Engkvist, Ola | |
| dc.contributor.supervisor | Österbacka, Nicklas | |
| dc.date.accessioned | 2026-08-05T06:46:41Z | |
| dc.date.issued | 2026 | |
| dc.date.submitted | ||
| dc.description.abstract | De novo molecular generation examines how computational methods can propose novel drug candidates by scoring generated molecules against desired properties and updating generation toward higher scoring regions. One approach trains a SMILES based language model and fine tunes it toward such regions. Fine tuning is commonly performed with policy gradient reinforcement learning (RL), as in AstraZenecas highly optimized REINVENT platform. An alternative is evolutionary strategies (ES), where a population of models is created and their parameters are updated to bias generation toward higher scores. This thesis investigates how distribution based ES compares with RL when implemented in REINVENT and evaluated on the Practical Molecular Optimization benchmark. OpenAI-ES shows near competitive performance on scoring and diversity metrics, whereas variance estimating Natural ES methods perform poorly. Variations of fixed variance ES are explored, and a novel sampling technique that biases generation toward high diversity shows promising performance across multiple metrics. | |
| dc.identifier.coursecode | DATX05 | |
| dc.identifier.uri | https://hdl.handle.net/20.500.12380/312072 | |
| dc.language.iso | eng | |
| dc.setspec.uppsok | Technology | |
| dc.subject | Evolution Strategies, Reinforcement Learning, De Novo Molecular Gen eration, Molecular AI, Drug Development, Fine-Tuning | |
| dc.title | Evolution Strategies as an Alternative to Reinforcement Learning in De Novo Molecular Generation - Evaluating Distribution Based Gaussian Evolution Strategies in REINVENT vs. Policy Based Reinforcement Learning on the Practical Molecular Optimization Benchmark | |
| dc.type.degree | Examensarbete för masterexamen | sv |
| dc.type.degree | Master's Thesis | en |
| dc.type.uppsok | H | |
| local.programme | Engineering mathematics and computational science (MPENM), MSc |
