tools / vram

Appendix C — VRAM estimator

Peak training memory per device from model size, precision, optimizer state, and sharding strategy. All figures derive live from the inputs; activations are taken as measured.

Derived figures
Weights + gradients ⁄ device 28.0 GB
Optimizer states ⁄ device 84.0 GB
Total ⁄ device 124 GB
Headroom −44.0 GB
12481632641282561000300100301031data-parallel width — devicesGB80 GB 1416642561000100101
Fig. C.1 — Peak per-device memory across data-parallel width under the selected sharding. Dashed guide marks device capacity. CUDA context, fragmentation, and communication buffers excluded — pad ~10%.
Fig. C.1 — GB per device vs devices, log scale.