Towards Reliable Retrieval Systems: Design and Evaluation of an Agentic GraphRAG Pipeline

dc.contributor.authorSpreitz, Adam
dc.contributor.departmentChalmers tekniska högskola / Institutionen för fysiksv
dc.contributor.departmentChalmers University of Technology / Department of Physicsen
dc.contributor.examinerGranath, Mats
dc.contributor.supervisorEhrenström Roos, Christian
dc.date.accessioned2026-09-11T12:30:28Z
dc.date.issued2026
dc.date.submitted
dc.description.abstractLarge language models are increasingly used for knowledge-intensive question answering, but their reliability remains limited when answers must be grounded in domain-specific documents. This limitation is particularly important in high-assurance environments, where systems must be controllable, traceable, and deployable without relying on external infrastructure. This thesis investigates the design and evaluation of an agentic GraphRAG system for document-grounded question answering, using scientific literature as a controlled proxy for technical internal documentation. The implemented system combines document ingestion, section-aware chunking, knowledge graph construction, vector indexing, graph traversal, reranking, and languagemodel- based answer generation. Two retrieval architectures are compared under shared conditions: a VectorRAG baseline using iterative hybrid search and a GraphRAG system using community-first hierarchical traversal over an explicit knowledge graph. The systems are evaluated across multiple retrieval configurations, generation models, and query sets using automated RAG evaluation metrics, statistical tests, and pairwise LLM-as-judge comparisons. The results show a clear divergence between retrieval-oriented metrics and answerlevel evaluation. VectorRAG achieves stronger automated retrieval scores, particularly on context precision and recall, while GraphRAG is preferred in holistic answer comparisons and manual validation. This suggests that chunk-level retrieval metrics do not always capture the usefulness of structurally retrieved evidence for downstream answer generation. The findings indicate that graph-based retrieval can improve answer quality and interpretability in agentic RAG systems, but also highlight important limitations related to dataset construction, evaluator dependence, corpus scale, graph quality, and agentic control.
dc.identifier.coursecodeTIFX05
dc.identifier.urihttps://hdl.handle.net/20.500.12380/312439
dc.language.isoeng
dc.setspec.uppsokPhysicsChemistryMaths
dc.subjectGraphRAG, retrieval-augmented generation, knowledge graphs, agentic systems, evaluation, Neo4j, LangChain, large language models
dc.titleTowards Reliable Retrieval Systems: Design and Evaluation of an Agentic GraphRAG Pipeline
dc.type.degreeExamensarbete för masterexamensv
dc.type.degreeMaster's Thesisen
dc.type.uppsokH
local.programmeComplex adaptive systems (MPCAS), MSc

Ladda ner

Original bundle

Visar 1 - 1 av 1
Hämtar...
Bild (thumbnail)
Namn:
Adam_Spreitz.pdf
Size:
4.36 MB
Format:
Adobe Portable Document Format

License bundle

Visar 1 - 1 av 1
Hämtar...
Bild (thumbnail)
Namn:
license.txt
Size:
2.35 KB
Format:
Item-specific license agreed upon to submission
Description: