From Alarms to Root-Cause: Autonomous Hardware Diagnostics through Multi-Agent Systems: A feasibility evaluation and proof-of-concept for LLM based diagnostics
Hämtar...
Ladda ner
Författare
Typ
Examensarbete för masterexamen
Master's Thesis
Master's Thesis
Modellbyggare
Tidskriftstitel
ISSN
Volymtitel
Utgivare
Sammanfattning
The rapid adoption of large language models (LLMs) and by extension LLM agents
is transforming the way work is done across various industrial sectors. Although
systemic performance increases are promised across many white-collar domains, certain
areas that could leverage the technology remain largely unexplored. Hardware
diagnostics, and by extension root-cause analysis are critical examples within the
telecommunications sector, where clients rely on high uptime, network availability,
and performance.
Modern telecommunication networks generate large volumes of daily log data that
can be leveraged for diagnostics. However, proactive manual human analysis of
this data is tedious, time-consuming, and unscalable across large fleets of units.
This thesis investigates the feasibility of autonomous LLM agents in this specialized
domain, evaluates these inherently non-deterministic workflows, and assesses how
custom investigation structure influences investigation behavior.
This thesis proposes Autonomous Diagnostics Agent (ADA), a multi-agent system
utilizing a hypothesis-driven investigation structure. ADA employs a dynamic planner
that builds a hypothesis tree of different plausible root-causes for a given hardware
unit, and autonomously explores and gathers evidence for each branch, while
verifying key insights in the logs. To this end, ADA coordinates a fleet of subagents,
specialized to tackle knowledge retrieval or log analysis.
ADA has been evaluated through qualitative domain expert feedback and through
a small automated test suite with known historical cases. The results show that
while ADA can provide utility mostly in-line with domain expert expectations for
well-documented issues, ADA also demonstrates multiple failure modes. The investigation
structure has an observable impact on the exploration breadth but results
in increased token usage, without any major correctness increase for the test cases
tried. Despite these challenges, domain-expert feedback suggests that ADA provides
a useful diagnostic baseline that helps save engineering triage time.
Ultimately, this work highlights that the overarching bottleneck for enterprise LLM
agents in complex engineering domains is evaluation. Since the number of possible
answer-solution pairs is vast, evaluation becomes reliant on proving stability in
outputs, and faithfulness to ground truth documentation and logs.
Beskrivning
Ämne/nyckelord
Agentic AI, AI, LLM, RCA, MAS, Autonomous, Diagnostics, Reporting
