All Malware Looks Similar Until It Doesn’t
Hämtar...
Publicerad
Författare
Typ
Examensarbete för masterexamen
Master's Thesis
Master's Thesis
Modellbyggare
Tidskriftstitel
ISSN
Volymtitel
Utgivare
Sammanfattning
This thesis investigates the use of machine learning to learn embeddings from static,
dynamic, and hybrid malware representations. The learned embeddings are evaluated
through similarity search and zero-shot clustering, with particular focus on
their ability to generalize beyond malware families observed during training. To
this end, Siamese metric learning approaches based on triplet loss are applied using
multi-layer perceptrons for static, dynamic, and hybrid representations, and
a transformer-based encoder for dynamic malware reports. The results highlight
the inherently difficult nature of learning semantically meaningful malware embeddings
that generalize to unseen malware families. In particular, the thesis shows
that strong performance on unseen samples from seen families in training does not
necessarily imply true generalization. Overall, dynamic and hybrid representations
provide a stronger basis for similarity learning than only static representations. The
findings further suggest that the main limitation is not necessarily the model architecture,
but mostly the noisy, imbalanced, and ambiguous nature of malware data
itself.
Beskrivning
Ämne/nyckelord
malware analysis, malware embeddings, metric learning, Siamese networks, similarity search, zero-shot generalization, hybrid malware representations, transformer encoders, EMBER, dynamic malware analysis
