Hasty Briefsbeta

Bilingual

What happens when a GPU writes memory

22 days ago
  • STG.E is the GPU instruction for writing data to global memory, following a load and addition.
  • The instruction reads three registers (address and data) and issues to the load/store unit (LSU), which sends addresses and masks to the coalescer.
  • Stores can be issued every ~6.1 cycles per warp, but the SM is limited to 32 bytes stored per cycle.
  • The coalescer converts 32 four-byte accesses into the minimum number of 32-byte sectors (usually four for contiguous writes).
  • L1 cache uses write-through policy: stores pass directly to L2 regardless of whether the line is resident; L1 may allocate on miss.
  • L2 cache uses a 16-way set-associative design with RRPV replacement (3 levels) and a write-back buffer; the 8-dirty rule proactively cleans lines to avoid burst evictions.
  • Stores are acknowledged from L2 back to the SM, but data may remain dirty in L2 until evicted; fences (membar.cta, membar.gl, membar.sys) ensure visibility at different scopes.
  • The driver inserts a system-scope membar after the kernel to make stores visible to the host for cudaMemcpyDeviceToHost.