the.ai

Hardware / Compute

verified

Memory Bandwidth Wall

Every accelerator generation adds far more arithmetic throughput than memory bandwidth. The gap has widened for decades, and the effect is that operations which used to be limited by how fast the chip could compute are now limited by how fast it can be fed. Buying a faster chip stops helping, which is a strange and specific kind of disappointment.

Viz primitive · budget-splitmemory-bound-kernels = 6

memory-bound-kernels holds 17% of the budget; rest holds the remaining 83%.

Kernels whose speed is set by memory bandwidth, against those still limited by arithmetic, in kernels. Drag the hardware generation forward to watch memory-bound take the workload — nothing in the code changed; the machine's demanded operations-per-byte rose underneath it.

6

Reviewed by opendroid · 2026-08-18