Evaluating Retrieval-Augmented Generation for Automated Regulatory GapAnalysis in the Information Security domain
Hämtar...
Ladda ner
Publicerad
Författare
Typ
Examensarbete för masterexamen
Master's Thesis
Master's Thesis
Modellbyggare
Tidskriftstitel
ISSN
Volymtitel
Utgivare
Sammanfattning
Regulatory compliance in information security requires organisations to map their
controls and mitigations against frameworks such as ISO/IEC 27001 and the NIS2
Directive. This process is traditionally manual, resource-intensive, and difficult to
scale. Large Language Models (LLMs) offer potential for automating compliance
tasks, but their tendency to hallucinate and produce unverifiable outputs limits
their applicability in contexts where findings must be defensible and auditable.
This thesis investigates whether Retrieval-Augmented Generation (RAG) improves
the accuracy, recommendation quality, and clause-level traceability of LLM-assisted
regulatory gap analysis compared to a non-grounded baseline. A controlled compu
tational experiment was conducted following Wohlin et al.’s experimentation frame
work. Two large language models (GPT-5.1 and Claude Opus 4.6) were evaluated
under two configurations (baseline and RAG) on a dataset of 93 compliance risks
mapped against ISO/IEC 27002 and ENISA NIS2 Technical Implementation Guid
ance. Outputs were evaluated against an expert-validated Golden Standard Ref
erence using clause prediction accuracy, coverage assessment accuracy, mitigation
quality metrics, and LLM-as-Judge evaluation.
Results show that RAG significantly improves gap detection for GPT-5.1, increasing
ISO clause prediction accuracy from 43.0% to 79.6% (p = 0.000001) and enabling
ENISA clause identification that is impossible from parametric knowledge alone (0%
to 33.3%). Coverage assessment accuracy improved from 78.5% to 98.9% under re
laxed evaluation. For Claude Opus 4.6, which already achieves 86.0% ISO accuracy
from parametric knowledge, RAG does not significantly improve clause prediction
but remains essential for ENISA identification. RAG acts as an equaliser across
model architectures, reducing the inter-model performance gap from 43 percentage
points to 5.3. LLM-as-Judge evaluation shows RAG produces higher-quality miti
gations (winning 80.7% of pairwise comparisons), while expert validation confirms
both pipelines achieve high ISO mitigation quality.
The study contributes empirical evidence on the use of retrieval-augmented ground
ing for information security GRC gap analysis, demonstrating that RAG improves
accuracy and enables traceability for regulatory frameworks absent from LLM pre
training data.
Beskrivning
Ämne/nyckelord
retrieval-augmented generation, regulatory compliance, gap analysis, large language models, information security, traceability, ISO 27001, NIS2, grounding, evaluation
