具身智能观察

Setting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72

产业动态

来源:NVIDIA 技术博客发布时间待核实

Frontier model pre-training has converged on mixture of experts (MoE), which is fundamentally changing what limits large-scale AI training. As compute per token... Frontier model pre-training has converged on mixture of experts (MoE), which is fundamentally changing what limits large-scale AI training. As compute per token falls, communication increasingly determines how efficiently models scale across thousands of GPUs.

Setting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72 | 具身智能观察