LLM-based Log Analysis - Evaluating LLM-Based Log Analysis in Industrial CI Environments: Technical Metrics, Practitioner Perspectives and Perceived Workflow Impact
Hämtar...
Ladda ner
Publicerad
Författare
Typ
Examensarbete för masterexamen
Master's Thesis
Master's Thesis
Modellbyggare
Tidskriftstitel
ISSN
Volymtitel
Utgivare
Sammanfattning
Continuous Integration (CI) pipelines generate large volumes of complex logs that
developers must analyze to diagnose failures and identify root causes. As soft
ware systems become increasingly distributed, manual log analysis has become time
consuming, cognitively demanding, and highly dependent on practitioners’ expertise.
Recent advances in Large Language Models (LLMs) have created new opportunities
for automated log analysis and debugging support. However, there remains a limited
understanding of how LLM-based log analysis tools should be evaluated and how
practitioners perceive their usefulness and influence on industrial CI environments.
This thesis addresses that gap through an industrial case study conducted in collaboration with Bosch. The study follows a mixed-methods research design that
combines qualitative and quantitative approaches, including a systematic literature
review, technical evaluation, descriptive statistical analysis, and correlation analyses
using Spearman’s rank correlation and Kendall’s tau.
The findings show that conventional text-overlap metrics, such as BLEU and ROUGE,
are not suitable for evaluating log analysis tool outputs because they require a reference and cannot handle the reasoning-intensive nature of CI log analysis. LLM-as-a
judge metrics, specifically G-Eval and GPTScore, were identified as more appropriate because they support reference-free and reasoning-aware evaluation. However,
the comparison between automated and practitioner evaluations revealed only weak
alignment and limited discriminative power in automated scoring, indicating that
automated evaluation cannot reliably replace human judgment. Practitioners were
found to evaluate outputs based on precise root-cause identification and whether the
outputs support debugging. The study further shows that practitioners perceive the
LLM-based tool as useful for supporting CI troubleshooting by accelerating failure
investigation, reducing manual log inspection, and providing interactive support,
particularly for complex and unfamiliar failures. It also highlights risks related to
hallucinations, over-reliance, and reduced critical reasoning.
Finally, practitioners perceived the integration of an MCP-based agent into the
development environment as useful for supporting workflow continuity by reducing context switching and enabling iterative conversational debugging. However,
they viewed it as complementary to, rather than a replacement for, the standalone
web-based tool. Together, the findings contribute to evaluation methodologies and
empirical insights into how LLM-based log analysis tools can be assessed and considered for adoption in industrial CI workflows.
Beskrivning
Ämne/nyckelord
Large Language Models, LLM-based Log Analysis, Human-Centered Metrics, LLM-as-a-Judge, Root Cause Analysis, MCP-based Agent, CI Workflows
