
A 2.8-trillion-parameter model with a context window of 1 million tokens is built on a mixture-of-experts architecture: out of 896 experts, only 16 are activated, reducing inference costs. According to Moonshot’s own data, scaling efficiency has increased by approximately 2.5 times compared to the previous version. On broader benchmarks, K3 falls behind top configurations from Claude and OpenAI—its lead is limited to coding tasks.
Continue reading this article on source: coindesk.com