Evaluating Retrieval-Augmented Generation for Automated Regulatory GapAnalysis in the Information Security domain
| dc.contributor.author | Mariyanayagam, Jenifa | |
| dc.contributor.department | Chalmers tekniska högskola / Institutionen för data och informationsteknik | sv |
| dc.contributor.department | Chalmers University of Technology / Department of Computer Science and Engineering | en |
| dc.contributor.examiner | Horkoff, Jennifer | |
| dc.contributor.supervisor | Fotrousi, Farnaz | |
| dc.date.accessioned | 2026-06-29T13:49:44Z | |
| dc.date.issued | 2026 | |
| dc.date.submitted | ||
| dc.description.abstract | Regulatory compliance in information security requires organisations to map their controls and mitigations against frameworks such as ISO/IEC 27001 and the NIS2 Directive. This process is traditionally manual, resource-intensive, and difficult to scale. Large Language Models (LLMs) offer potential for automating compliance tasks, but their tendency to hallucinate and produce unverifiable outputs limits their applicability in contexts where findings must be defensible and auditable. This thesis investigates whether Retrieval-Augmented Generation (RAG) improves the accuracy, recommendation quality, and clause-level traceability of LLM-assisted regulatory gap analysis compared to a non-grounded baseline. A controlled compu tational experiment was conducted following Wohlin et al.’s experimentation frame work. Two large language models (GPT-5.1 and Claude Opus 4.6) were evaluated under two configurations (baseline and RAG) on a dataset of 93 compliance risks mapped against ISO/IEC 27002 and ENISA NIS2 Technical Implementation Guid ance. Outputs were evaluated against an expert-validated Golden Standard Ref erence using clause prediction accuracy, coverage assessment accuracy, mitigation quality metrics, and LLM-as-Judge evaluation. Results show that RAG significantly improves gap detection for GPT-5.1, increasing ISO clause prediction accuracy from 43.0% to 79.6% (p = 0.000001) and enabling ENISA clause identification that is impossible from parametric knowledge alone (0% to 33.3%). Coverage assessment accuracy improved from 78.5% to 98.9% under re laxed evaluation. For Claude Opus 4.6, which already achieves 86.0% ISO accuracy from parametric knowledge, RAG does not significantly improve clause prediction but remains essential for ENISA identification. RAG acts as an equaliser across model architectures, reducing the inter-model performance gap from 43 percentage points to 5.3. LLM-as-Judge evaluation shows RAG produces higher-quality miti gations (winning 80.7% of pairwise comparisons), while expert validation confirms both pipelines achieve high ISO mitigation quality. The study contributes empirical evidence on the use of retrieval-augmented ground ing for information security GRC gap analysis, demonstrating that RAG improves accuracy and enables traceability for regulatory frameworks absent from LLM pre training data. | |
| dc.identifier.coursecode | DATX05 | |
| dc.identifier.uri | https://hdl.handle.net/20.500.12380/311628 | |
| dc.language.iso | eng | |
| dc.setspec.uppsok | Technology | |
| dc.subject | retrieval-augmented generation, regulatory compliance, gap analysis, large language models, information security, traceability, ISO 27001, NIS2, grounding, evaluation | |
| dc.title | Evaluating Retrieval-Augmented Generation for Automated Regulatory GapAnalysis in the Information Security domain | |
| dc.type.degree | Examensarbete för masterexamen | sv |
| dc.type.degree | Master's Thesis | en |
| dc.type.uppsok | H | |
| local.programme | Software engineering and technology (MPSOF), MSc |
