GitHub Projects @githubprojects.bsky.social · 16d

DeepSpeed provides training optimizations for large-scale models, reducing memory and compute costs through advanced offloading and pipeline techniques. - ZeRO offload and ZenFlow engine reduce memory stalls during LLM training

0 likes 1 replies

?

Replies

GitHub Projects · 16d

- SuperOffload enables large-scale training on superchips with efficient memory use - Muon Optimizer support integrates modern optimization methods into DeepSpeed - System DMA (SDMA) offloads collectives on AMD GPUs for better overlap with ZeRO-3 Explore it here: osp.fyi/deepspeed