What happens when a GPU writes memory
22 days ago
- STG.E is the GPU instruction for writing data to global memory, following a load and addition.
- The instruction reads three registers (address and data) and issues to the load/store unit (LSU), which sends addresses and masks to the coalescer.
- Stores can be issued every ~6.1 cycles per warp, but the SM is limited to 32 bytes stored per cycle.
- The coalescer converts 32 four-byte accesses into the minimum number of 32-byte sectors (usually four for contiguous writes).
- L1 cache uses write-through policy: stores pass directly to L2 regardless of whether the line is resident; L1 may allocate on miss.
- L2 cache uses a 16-way set-associative design with RRPV replacement (3 levels) and a write-back buffer; the 8-dirty rule proactively cleans lines to avoid burst evictions.
- Stores are acknowledged from L2 back to the SM, but data may remain dirty in L2 until evicted; fences (membar.cta, membar.gl, membar.sys) ensure visibility at different scopes.
- The driver inserts a system-scope membar after the kernel to make stores visible to the host for cudaMemcpyDeviceToHost.