Welcome to The AI Circle Brief #3 · Kimi K3

Moonshot's Kimi K3 open weights just went live and everyone’s fixated on its sizeand intelligence, however the more interesting story is how it was built.

Every time a normal LLM produces a new token, it looks back at every token that came before it. Which means the longer your prompt or document the more costly it becomes to use an AI model.

Kimi K3 does the opposite, it refuses to let its memory grow at all. It uses a novel technique called Kimi Delta Attention combined with something called "fine-grained forgetting" which buys up to 6.3x faster decoding at million-token lengths.

The second trick is arguably more impressive - Attention Residuals stops older model weight layers from getting ignored as models grow. Kimi K3 achieves roughly 25% better training efficiency for under 2% extra compute. Together both of these designs underpin Moonshot's claimed 2.5x scaling efficiency over Kimi K2.

But what does this mean for you?

Our full deep dive walks through the design, reasoning (with interactive demos) and what this means for anyone building AI that reads long documents, holds long conversations, or runs multiple agents with memory.

Enjoy!

AI Circle is a community of researchers and operators from the frontier labs and businesses building with AI. Find out more at ai-circle.org and join us at our next event!

Keep reading