Semantic Similarity of LLM Explanations in Software Engineering Tasks

dc.contributor.authorAndersson, Simon
dc.contributor.authorWang, Qianyuan
dc.contributor.departmentChalmers tekniska högskola / Institutionen för data och informationstekniksv
dc.contributor.departmentChalmers University of Technology / Department of Computer Science and Engineeringen
dc.contributor.examinerFotrousi, Farnaz
dc.contributor.supervisorFeldt , Robert
dc.contributor.supervisorAkbarova, Sabina
dc.date.accessioned2026-07-06T12:06:22Z
dc.date.issued2026
dc.date.submitted
dc.description.abstractLarge language models (LLMs) are increasingly used to generate natural-language explanations for software engineering (SE) tasks. However, it is still unclear how similar these explanations are across models when the task remains the same. We address this gap in two SE contexts: Python code comprehension and explaining requirements classifications. We use a controlled experimental design to compare GPT-4.1, Claude Sonnet 4, and Gemini 2.5 Flash under two prompt types: high level and detailed. In total, the experiment produces 240 explanations. We analyze these explanations from three complementary perspectives: embedding-based semantic similarity, manual reasoning pattern analysis using a predefined codebook, and perceived similarity and quality through a human survey and an LLM-as-a-judge evaluation. Our results show that explanations from different models are often close in meaning, but their similarity varies across model pairs, task types, and prompt conditions. The reasoning pattern analysis shows that models also differ in how they structure their explanations, and these differences depend on the task and prompt type. The quality evaluation of the selected explanation pairs shows that no model is consistently better than the others across all dimensions, and that human and LLM-as-a-judge judgments do not always agree. Therefore, evaluating LLM explanations in SE should combine semantic similarity measures, reasoning-structure analysis, and human or LLM-based quality judgments to support trustworthy use of these models in practice.
dc.identifier.coursecodeDATX05
dc.identifier.urihttps://hdl.handle.net/20.500.12380/311873
dc.language.isoeng
dc.setspec.uppsokTechnology
dc.subjectLarge language models, semantic similarity, explanation quality, explanation structure, LLM-as-a-judge
dc.titleSemantic Similarity of LLM Explanations in Software Engineering Tasks
dc.type.degreeExamensarbete för masterexamensv
dc.type.degreeMaster's Thesisen
dc.type.uppsokH
local.programmeSoftware engineering and technology (MPSOF), MSc

Ladda ner

Original bundle

Visar 1 - 1 av 1
Hämtar...
Bild (thumbnail)
Namn:
CSE 26-51 SA QW.pdf
Size:
2.31 MB
Format:
Adobe Portable Document Format

License bundle

Visar 1 - 1 av 1
Hämtar...
Bild (thumbnail)
Namn:
license.txt
Size:
2.35 KB
Format:
Item-specific license agreed upon to submission
Description: