This article highlights the key concepts from the Security and Performance Implications of GPU Cache Eviction Priority Hints by Qizhong Wang, Xiangyue Huang, Yanan Guo, and Yuanchao X.
Cache intro
Two-level cache hierarchy: each Streaming Multiprocessor (SM) has a private L1 cache, and all SMs share the L2 cache
Least Recently Used (LRU) replacement policy: the cache line unused for the longest time is evicted when more space is needed
Problem: LRU’s predictions are not always reliable + data reuse is a complex task
Solution: cache eviction priority hints (1) evict_first: highest eviction priority → remove first
- Used when streaming data
(2) evict_last: lowest eviction priority → remove last
- Used when: needed for data that needs to remain persistent
(3) evict_normal: no priority is set → default hint
How it works? : using the hints to make decisions – evict_first priority hint, the sender and receiver can target the single slot allocated for evict_first cache lines, creating conflicts with fewer operations = efficient covert channel

Listing 1 — Security and Performance Implications of GPU Cache Eviction Priority Hints
//line 2: create a policy, specifying evict_last with 64 bit memory addresses, with the base memory address ([addr] — first 128 = memory range, second 128 = max range limit for the policy)
//line 5: loading instructions to the L2 cache controller with an unsigned 64-bit integer as the data type
//line 5: loads 64 bit val at [addr] into target register val → attaches the cache_policy
GPU Memory Utilization
- Each SM (Simulatenous Multiprocessor) has 1 L1 each
- On-chip network connects shared L2 cache to the SM
- L2 cache connected to memory controllers → manage the device memory DRAM
- GPUs can access each other memories without involving the host
- Multi-Process Service (MPS) = enables multiple processes to run simultaneously on a GPU → problem: processes share the same address space → Volta-MPS = workloads can run in separate address spaces in parallel
Figure 1 — Security and Performance Implications of GPU Cache Eviction Priority Hints
Covert Channels
A covert channel uses shared cache and memory to transfer data. A sender transmits information by causing conflicts in a cache set, while the receiver retrieves the data by measuring the conflicts.
Latency means the time delayed. An increased latency means the model has been degraded. The degradation attack methods are meant to slow down a model’s reaction time.
Finger Printing attack
(1) Finger printing attack: one user marks several cache lines per L2 set with the evict_last priority → “pin” those cache lines in the L2 cache set → replacement policy does not evict them → reduces the cache availability for other users → increases cache access latency due to the limited L2 bandwidth
– reveals that a cache access occurred or which set was accessed
– Since evict_last allows a user to pin data in the cache, which can lead to performance degradation in multi-tenant scenario
– Specifically: Marking more than 12/16 (or 3/16) of the L2 cache size:
GPUs only allow a max of 12 cache lines per set (or 3 – depending on the GPU driver version) evict_last cache lines per 16-way L2 cache set → more cache lines are marked with evict_last hint, they cannot remain pinned
– Why this is a stealthy attack? Uses prime+probe side-channel attack
Attacker prime and probes 16 cache lines (check if any lines evicted) → uses evict_last to pin n lines inside a set → attacker only needs to touch the remaining (16 - n) lines to see the victim’s memory
Ex. Attacker pins n = 8 cache lines, active monitoring footprint = 16 - 8 = 8 lines/set → victim runs GPU → attacker gets cache evictions tracked by time → from this data can make a spatial/temporal memory pattern (memorygram) → CNN (Convolutional Neural Network) can be trained on memorygrams to classify applications running
- Impact: access cut down from 16 to 8 access/set per iteration → attacker 1/2s cache access overhead → results in lower memory-bandwidth footprint + higher sampler rates = hardware performance counters are less likely to detect attacks
Ex. Another way is to set the set the priority_hint to a certain 3 evict_last cache lines in each set of an L2 cache
- Impact: time it takes to access each set increases which increases idle time
Takeaways:
(1) Loading evict_normal cache lines can evict existing evict_last cache lines
(2) A new evict_last cache line can evict an existing evict_last cache line
Both cause a high miss rate, rather than having them pinned in the cache…degrades the cache hit rate and increases latency
Cache eviction hints
(2) Cache eviction hints detection attack:
(a) The prep: Attack relies on how the hardware replacement policy chooses the next eviction candidate: victim accesses (using evict_last) causing hardware to pick/evict an LRU line → victim accesses (using evict_normal) causing hardware to not pick high-priority lines, and pick normal/unpinned lines
(b) The watch: what does the attacker watch? Attacker monitors specific cache lines to distinguish access types:
-
Eviction of overall LRU line: victim executed an evict_last access
-
Eviction of normal-group LRU line: victim executed an evict_normal access
Can identify which line was pushed out → sees the victim’s hints
(c) The hunt: able to attack more specifically if we know the hint policies
Ex. Matrix multiplication kernel tags an input matrix/tile with evict_last (pinned to L2 cache for reuse), other matrix/buffers have evict_normal (default priority)
– Attacker can now track the ratio, frequency, sequence of evict_last and evict_normal accesses (1 evict_last hit → 4 evict_normal hits = tells attacker the pattern)
– Attacker knows # of high-priority steps = 4 inner tiles fit in a dimension
– Attacker knows timing duration of each block (in turn attacker knows the size of the tiles)
– Attacker knows matrix size (tile dimensions * # of loop cycles) = exact dimensions of private matrices (batch sizes, layer widths)
(3) refresh the pinned cache lines ~ every 108 cycles before they are removed from the L2 cache
- brute-force scanning attack would need to hit the cache over 40x times to match the same impact of the (1) attack
(4) Fewer threads (256 threads) makes idle time 97.75%
Side note: Attacks won’t work on applications that run streaming workloads: streaming access patterns do not reuse data → incur high L2 rates → pinned lines do not affect the caching → attacker has to switch to high-frequency bandwidth saturation → 9% performance degradation
(5) Create conflicts within a covert channel based on the evict_first priority hint
– Goal: one evict_first cache line exists in a L2 set → each evict_first cache (sender and receiver’s cache line) competes for that single slot → creates conflicts
– Latency increases when multiple threads access L2 cache at the same time, also due to the limited bandwidth
Sources
Wang, Qizhong, et al. “Security and Performance Implications of GPU Cache Eviction Priority Hints.” Security and Performance Implications of GPU Cache Eviction Priority Hints, 2025, pp. 1058-1072. ACM Digital Library, https://dl.acm.org/doi/10.1145/3725843.3756116#core-fn1-1. Accessed 28 September 2026.
Rishabh Jain, Vivek M Bhasi, Adwait Jog, Anand Sivasubramaniam, Mahmut T Kandemir, and Chita R Das. 2024. Pushing the Performance Envelope of DNN-based Recommendation Systems Inference on GPUs. In 2024 57th IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 1217–1232.
Haoxuan Liu, Vasu Singh, Michał Filipiuk, and Siva Kumar Sastry Hari. 2024. ALBERTA: ALgorithm-Based Error Resilience in Transformer Architectures. IEEE Open Journal of the Computer Society (2024).
Link: https://ieeexplore.ieee.org/abstract/document/10530530
Zhenkai Zhang, Tyler Allen, Fan Yao, Xing Gao, and Rong Ge. 2023. T unne L s for B ootlegging: Fully Reverse-Engineering GPU TLBs for Challenging Isolation Guarantees of NVIDIA MIG. In Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security. 960–974.
Link: https://www.usenix.org/conference/osdi24/presentation/zhuang