Kimi K3 MoE mixture of experts shorts
Watch on YouTube In practice, about fifty billion parameters are active for each computation, not 2.8 trillion. Mixture of experts. The MoE architecture that Mistral already popularized, which DeepSeek later adopted, and now Moonshot. Basically, the model is huge on paper but lightweight in execution. Exactly. And this is
In practice, about fifty billion parameters are active for each computation, not 2.8 trillion. Mixture of experts. The MoE architecture that Mistral already popularized, which DeepSeek later adopted, and now Moonshot. Basically, the model is huge on paper but lightweight in execution. Exactly. And this is where Moonshot does something worth understanding: they aren’t competing on raw computing power, because they can’t. Nvidia’s chip export restrictions to China cut off their access to the latest training GPUs. So they are betting on architectural innovation. And this isn’t the first time we’ve seen that.
Full episode: https://youtu.be/kWzw1qh8nHw
🤖 AI-generated content: the script, voices, and images for this episode were produced using artificial intelligence tools.
#Shorts