Modal, Fireworks, and Baseten gain cost advantage with Nvidia and AMD chips for Kimi K3

1 hour ago 1



Three American AI infrastructure companies have found a lucrative edge in an unlikely place: serving a Chinese-built model better and cheaper than its creators can. Modal, Fireworks AI, and Baseten are now offering hosted inference for Moonshot AI’s Kimi K3 at roughly one-tenth the cost of direct access, powered by Nvidia’s and AMD’s latest accelerator hardware that remains largely unavailable to Chinese firms under US export controls. What makes Kimi K3 worth the effort Kimi K3 is not a small model. Released by Moonshot AI on July 27, 2026, it packs 2.8 trillion parameters into a Mixture-of-Experts architecture, making it one of the largest open-weights models ever published. The “open weights” distinction matters: anyone can download and run the model, which is exactly what Modal, Fireworks, and Baseten have done. The MoE design means the full 2.8 trillion parameters don’t activate on every query. Instead, K3 routes each token through 16 of its 896 available experts, activating roughly 104 billion parameters per forward pass. K3 also supports a context window of 1 million tokens, which means it can ingest and reason over book-length documents in a single session. But running a mo...

Read Entire Article