Semantic Similarity of LLM Explanations in Software Engineering Tasks
Hämtar...
Ladda ner
Publicerad
Författare
Typ
Examensarbete för masterexamen
Master's Thesis
Master's Thesis
Modellbyggare
Tidskriftstitel
ISSN
Volymtitel
Utgivare
Sammanfattning
Large language models (LLMs) are increasingly used to generate natural-language
explanations for software engineering (SE) tasks. However, it is still unclear how
similar these explanations are across models when the task remains the same. We
address this gap in two SE contexts: Python code comprehension and explaining
requirements classifications. We use a controlled experimental design to compare
GPT-4.1, Claude Sonnet 4, and Gemini 2.5 Flash under two prompt types: high
level and detailed. In total, the experiment produces 240 explanations. We analyze
these explanations from three complementary perspectives: embedding-based semantic similarity, manual reasoning pattern analysis using a predefined codebook,
and perceived similarity and quality through a human survey and an LLM-as-a-judge
evaluation. Our results show that explanations from different models are often close
in meaning, but their similarity varies across model pairs, task types, and prompt
conditions. The reasoning pattern analysis shows that models also differ in how they
structure their explanations, and these differences depend on the task and prompt
type. The quality evaluation of the selected explanation pairs shows that no model
is consistently better than the others across all dimensions, and that human and
LLM-as-a-judge judgments do not always agree. Therefore, evaluating LLM explanations in SE should combine semantic similarity measures, reasoning-structure
analysis, and human or LLM-based quality judgments to support trustworthy use
of these models in practice.
Beskrivning
Ämne/nyckelord
Large language models, semantic similarity, explanation quality, explanation structure, LLM-as-a-judge
