Benchmarking Self-Supervised Log Embeddings Across Log Modalities and Evaluation Targets

Hämtar...
Bild (thumbnail)

Publicerad

Typ

Examensarbete för masterexamen
Master's Thesis

Modellbyggare

Tidskriftstitel

ISSN

Volymtitel

Utgivare

Sammanfattning

Log data are increasingly used to support monitoring, troubleshooting, anomaly detection, and operational decision making in software, industrial, and process systems. These records contain heterogeneous information, including structured fields, message text, event sequences, timestamps, and numerical values. Embedding methods can convert such records into fixed-dimensional vectors for downstream analysis, but the usefulness of an embedding depends on which parts of the original input unit remain accessible after representation construction. Evaluating log embeddings is challenging because log data do not form a single homogeneous modality. Input units can range from individual rows to fixed windows and larger operational traces, and each scale exposes different event-level, sequencelevel, numerical, and behavioural structure. At the same time, downstream scores can reflect only one aspect of an embedding space. A classifier score, retrieval result, clustering metric, numerical probe, weak outcome target, or runtime measurement may therefore lead to different conclusions about the same representation. The aim of this thesis is to benchmark and analyse log embeddings across heterogeneous log settings, with emphasis on what different representation methods preserve. The study compares strong transparent baselines with self-supervised learned embeddings based on reconstruction, contrastive, and hybrid objectives. Four datasets are used to cover structured row-level logs, operational row and window logs, blocklevel system event sequences, and process-oriented subprocess logs. The embeddings are evaluated using probes, retrieval, clustering, numerical recovery, weak outcome targets, and runtime. The results show that log embedding quality is conditional on the relation between the input unit, the available observable structure, and the evaluation target. Simple baselines remain strong when the target is close to explicit fields, tokens, counts, or numerical summaries. Learned embeddings provide clearer benefits when the target depends on context or relations across events. The main conclusion is methodological: log embeddings should be evaluated as profiles of preserved information, with method selection grounded in the intended analysis use.

Beskrivning

Ämne/nyckelord

log embeddings, representation learning, self-supervised learning, log analysis, benchmark evaluation, operational logs

Citation

Arkitekt (konstruktör)

Geografisk plats

Byggnad (typ)

Byggår

Modelltyp

Skala

Teknik / material

Index

Endorsement

Review

Supplemented By

Referenced By