Kimi K3 demos are everywhere, but the prompt is never shown

Kimi K3 demos are everywhere, but the prompt is never shown

I saw a video this morning of someone making a Halo CE 10v10 multiplayer game with one prompt in Kimi K3. No engine, no dev team. The WSJ is asking if this is a “DeepSeek 2.0” moment for markets.

S @Rubzem

Kimi K3 is actually crazy. Someone just remade HALO CE 10v10 multiplayer with one prompt. No engine. No dev team. No months of work. Kimi K3 is miles better than Anthropics Fable 5 from what I can see. We will be seing AI game making take over after the summer! https://t.co/zjOyITA9eq

View on X →

The video is genuinely impressive. The problem is what’s missing: the prompt. Nobody shares the exact text they gave the model. Without that, I have no way to tell if the demo is representative or a cherry-picked result after 20 attempts. The same thing happened with Sora and GPT-5 — the most viral demos rarely include the input that produced them.

I haven’t tried Kimi K3 yet. The benchmarks put it at #1 in frontend code, surpassing Claude Fable 5 with a 17-place jump from the previous version.

Arena.ai @arena

Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5. This is a 17-place jump from Kimi-k2.6 (#18 -> #1). In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Simulations, and Content Creation Tools, landing #2 only in Gaming behind Fable 5. The full model weights will be released by July 27. Congrats to the @Kimi_Moonshot team on this major milestone!

View on X →

That’s a real achievement, and the technical report mentions architecture improvements like Delta Attention for faster decoding on million-token contexts. But a benchmark is an average; my work is a stream of one-shot, ambiguous requests. I built most of my projects this year next to an agentic coding loop, and I wrote about the gap between what works in a clean eval and what actually helps at 2 a.m.. Models need to be predictable under my specific mess, not just score high.

Chetaslua @chetaslua

Kimi K3 vs GPT-5.6 Sol Difference in taste is so stark , like if i swap kimi k 3 name with fable 5 people will trust it but this is not only about visuals , its also about function , you can see both achieve same result but the way to achieve is different kimi is more creative

View on X →

When open weights drop July 27, I’ll probably download the model and run it locally. No API cost, and I can test it with my own prompts — where I control the input and see the full output. Until then, the demos are entertaining, but they’re not evidence I can build on. Show me the prompt, and then we can talk.