When you send a prompt, the model doesn't wake up all 141 billion parameters. Instead, a tiny routing network looks at your input and says: "For this specific word/token, these 2 experts are the most qualified to handle it." It activates only those 2 experts (out of 8), processes the token, and moves to the next one.
Think of it like a hospital. You don't wake up every single doctor for every patient. The triage nurse (the router) directs each case to the right specialist. The cardiologist handles heart stuff. The neurologist handles brain stuff. They don't all crowd around a broken ankle.
This is why MoE models are so efficient — and why the naming is confusing. 8x22B doesn't mean 176B active. It means 8 experts of ~22B each, with only ~39B active at any moment.
Why This Matters for Local Hosting
If I had actually needed to load all 141 billion parameters into memory at full precision (FP32), I'd need ~564 GB of VRAM (4 bytes per parameter × 141B parameters).
That's not a Mac Mini. That's not even a high-end gaming PC. That's a server rack.
But because of the MoE architecture and quantization (compressing the model's precision from 32-bit to 16-bit, 8-bit, or even 4-bit), the actual requirements drop dramatically: