Peak training memory per device from model size, precision, optimizer state, and
sharding strategy. All figures derive live from the inputs; activations are taken
as measured.
Derived figures
Weights + gradients ⁄ device28.0 GB
Optimizer states ⁄ device84.0 GB
Total ⁄ device124 GB
Headroom−44.0 GB
Fig. C.1 — Peak per-device memory across data-parallel width under the selected sharding.
Dashed guide marks device capacity. CUDA context, fragmentation, and communication
buffers excluded — pad ~10%.