NVIDIA 解释了 Nemotron 3.5 Lightning 等混合专家模型如何在拥有 30B 参数的情况下仅激活 3B 参数。

How can a 30B-parameter model activate only 3B parameters per token, and still use the capacity of the larger model? Nemotron 3.5 Lightning illustrates the...