All Malware Looks Similar Until It Doesn’t

Hämtar...
Bild (thumbnail)

Publicerad

Typ

Examensarbete för masterexamen
Master's Thesis

Modellbyggare

Tidskriftstitel

ISSN

Volymtitel

Utgivare

Sammanfattning

This thesis investigates the use of machine learning to learn embeddings from static, dynamic, and hybrid malware representations. The learned embeddings are evaluated through similarity search and zero-shot clustering, with particular focus on their ability to generalize beyond malware families observed during training. To this end, Siamese metric learning approaches based on triplet loss are applied using multi-layer perceptrons for static, dynamic, and hybrid representations, and a transformer-based encoder for dynamic malware reports. The results highlight the inherently difficult nature of learning semantically meaningful malware embeddings that generalize to unseen malware families. In particular, the thesis shows that strong performance on unseen samples from seen families in training does not necessarily imply true generalization. Overall, dynamic and hybrid representations provide a stronger basis for similarity learning than only static representations. The findings further suggest that the main limitation is not necessarily the model architecture, but mostly the noisy, imbalanced, and ambiguous nature of malware data itself.

Beskrivning

Ämne/nyckelord

malware analysis, malware embeddings, metric learning, Siamese networks, similarity search, zero-shot generalization, hybrid malware representations, transformer encoders, EMBER, dynamic malware analysis

Citation

Arkitekt (konstruktör)

Geografisk plats

Byggnad (typ)

Byggår

Modelltyp

Skala

Teknik / material

Index

Endorsement

Review

Supplemented By

Referenced By