Congesting Distributed AI Within the Host Network
Loading...
Download
Date
Type
Examensarbete för masterexamen
Master's Thesis
Master's Thesis
Model builders
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
As artificial intelligence models scale, they rely on distributed, multi-tenant hardware.
In these environments, the internal host network of e.g. CPUs, GPUs, memory
controllers and inter-socket links is critical to transferring data efficiently. The
internal host network is often viewed as a trusted environment, but congestion of
the shared paths can have substantial effect on performance.
Consequently, this thesis addresses the following research question: Can a co-located
adversary intentionally utilize host network congestion to execute an attack against
distributed AI workloads?
To evaluate this potential threat, we developed adversarial workloads mainly targeting
the memory and the inter-socket link. These attacks were executed concurrently with
Transformer and Graph Neural Network (GNN) models while capturing performance
and hardware statistics.
The results show that an adversary can weaponize the host network and cause
substantial end-to-end performance degradation for the AI models. The magnitude
of the performance degradation depends on how the models use the hardware. Models
that continuously rely on the CPU memory for data suffered the largest performance
drops, when an attacker saturated CPU memory. In contrast, the model that kept
it’s data in the GPUs’ VRAM and synchronized peer-to-peer was immune to this
kind of attack. It was however vulnerable to an attack against the inter-socket path.
Essentially, this thesis demonstrates a vulnerability in multi-tenant high-performance
computing. An adversary does not need to break virtual machine boundaries or have
elevated privileges to cause harm. Exploiting the host network can cause damage in
terms of performance degradation. Expensive GPUs being underutilized translates
to a substantial indirect financial cost.
Description
Keywords
Host network, distributed AI, hardware security, congestion attacks
