Data Science Wire

Dense vs. MoE Models: Active Parameters, Throughput, and When to Choose Each

NVIDIA Technical Blog - AI9h4 min read

How can a 30B-parameter model activate only 3B parameters per token, and still use the capacity of the larger model? Nemotron 3.5 Lightning illustrates the...

Read the full story at NVIDIA Technical Blog - AI

More in AI