Benchmarking Self-Supervised Log Embeddings Across Log Modalities and Evaluation Targets
Hämtar...
Publicerad
Författare
Typ
Examensarbete för masterexamen
Master's Thesis
Master's Thesis
Modellbyggare
Tidskriftstitel
ISSN
Volymtitel
Utgivare
Sammanfattning
Log data are increasingly used to support monitoring, troubleshooting, anomaly
detection, and operational decision making in software, industrial, and process systems.
These records contain heterogeneous information, including structured fields,
message text, event sequences, timestamps, and numerical values. Embedding methods
can convert such records into fixed-dimensional vectors for downstream analysis,
but the usefulness of an embedding depends on which parts of the original input
unit remain accessible after representation construction.
Evaluating log embeddings is challenging because log data do not form a single
homogeneous modality. Input units can range from individual rows to fixed windows
and larger operational traces, and each scale exposes different event-level, sequencelevel,
numerical, and behavioural structure. At the same time, downstream scores
can reflect only one aspect of an embedding space. A classifier score, retrieval result,
clustering metric, numerical probe, weak outcome target, or runtime measurement
may therefore lead to different conclusions about the same representation.
The aim of this thesis is to benchmark and analyse log embeddings across heterogeneous
log settings, with emphasis on what different representation methods preserve.
The study compares strong transparent baselines with self-supervised learned embeddings
based on reconstruction, contrastive, and hybrid objectives. Four datasets
are used to cover structured row-level logs, operational row and window logs, blocklevel
system event sequences, and process-oriented subprocess logs. The embeddings
are evaluated using probes, retrieval, clustering, numerical recovery, weak outcome
targets, and runtime.
The results show that log embedding quality is conditional on the relation between
the input unit, the available observable structure, and the evaluation target. Simple
baselines remain strong when the target is close to explicit fields, tokens, counts, or
numerical summaries. Learned embeddings provide clearer benefits when the target
depends on context or relations across events. The main conclusion is methodological:
log embeddings should be evaluated as profiles of preserved information, with
method selection grounded in the intended analysis use.
Beskrivning
Ämne/nyckelord
log embeddings, representation learning, self-supervised learning, log analysis, benchmark evaluation, operational logs
