LLM-based Log Analysis - Evaluating LLM-Based Log Analysis in Industrial CI Environments: Technical Metrics, Practitioner Perspectives and Perceived Workflow Impact

Hämtar...
Bild (thumbnail)

Publicerad

Typ

Examensarbete för masterexamen
Master's Thesis

Modellbyggare

Tidskriftstitel

ISSN

Volymtitel

Utgivare

Sammanfattning

Continuous Integration (CI) pipelines generate large volumes of complex logs that developers must analyze to diagnose failures and identify root causes. As soft ware systems become increasingly distributed, manual log analysis has become time consuming, cognitively demanding, and highly dependent on practitioners’ expertise. Recent advances in Large Language Models (LLMs) have created new opportunities for automated log analysis and debugging support. However, there remains a limited understanding of how LLM-based log analysis tools should be evaluated and how practitioners perceive their usefulness and influence on industrial CI environments. This thesis addresses that gap through an industrial case study conducted in collaboration with Bosch. The study follows a mixed-methods research design that combines qualitative and quantitative approaches, including a systematic literature review, technical evaluation, descriptive statistical analysis, and correlation analyses using Spearman’s rank correlation and Kendall’s tau. The findings show that conventional text-overlap metrics, such as BLEU and ROUGE, are not suitable for evaluating log analysis tool outputs because they require a reference and cannot handle the reasoning-intensive nature of CI log analysis. LLM-as-a judge metrics, specifically G-Eval and GPTScore, were identified as more appropriate because they support reference-free and reasoning-aware evaluation. However, the comparison between automated and practitioner evaluations revealed only weak alignment and limited discriminative power in automated scoring, indicating that automated evaluation cannot reliably replace human judgment. Practitioners were found to evaluate outputs based on precise root-cause identification and whether the outputs support debugging. The study further shows that practitioners perceive the LLM-based tool as useful for supporting CI troubleshooting by accelerating failure investigation, reducing manual log inspection, and providing interactive support, particularly for complex and unfamiliar failures. It also highlights risks related to hallucinations, over-reliance, and reduced critical reasoning. Finally, practitioners perceived the integration of an MCP-based agent into the development environment as useful for supporting workflow continuity by reducing context switching and enabling iterative conversational debugging. However, they viewed it as complementary to, rather than a replacement for, the standalone web-based tool. Together, the findings contribute to evaluation methodologies and empirical insights into how LLM-based log analysis tools can be assessed and considered for adoption in industrial CI workflows.

Beskrivning

Ämne/nyckelord

Large Language Models, LLM-based Log Analysis, Human-Centered Metrics, LLM-as-a-Judge, Root Cause Analysis, MCP-based Agent, CI Workflows

Citation

Arkitekt (konstruktör)

Geografisk plats

Byggnad (typ)

Byggår

Modelltyp

Skala

Teknik / material

Index

Endorsement

Review

Supplemented By

Referenced By