AI Models

What Is Kimi K3? Moonshot's Frontier-Class Reasoner, Explained

Moonshot AI's Kimi K3 scores within 3 points of Claude Fable 5 at a third of the cost — but it thinks before it talks. The verified numbers, the speed trade-off, and when to pick it.

Leo Parker·July 25, 20268 min read

TL;DR: Kimi K3 is Moonshot AI's frontier reasoning model, launched July 16, 2026. Per Artificial Analysis it scores 57 on the Intelligence Index — #7 of 190 models, within 3 points of Claude Fable 5's leading 60 — at roughly 1/3.3 of Fable's blended API cost. The catch: it's slow. 32.6 tokens/sec output and ~161 seconds to first token make it a think-first model, not a chat sprinter. On TulexAI it costs 7 PT per message on every plan, from $11/mo.

What is Kimi K3?

Kimi K3 is the flagship model from Moonshot AI, the Beijing lab behind the Kimi family of open-weight models. It shipped on July 16, 2026, and the pitch is simple: frontier-class deep reasoning with a 1-million-token context window, priced well below the Western frontier. Simon Willison's launch-day writeup (simonwillison.net) captured why people paid attention — this isn't a budget model punching up on one cherry-picked benchmark; it's a top-ten model on a neutral scoreboard, from a lab most casual users hadn't tried yet.

K3 went live on TulexAI the same day it launched. If you've been running the same prompt through Claude, GPT and Grok to see who wins, there's now a fourth serious contender in the dropdown.

The numbers, from a neutral scoreboard

All benchmark figures below come from Artificial Analysis, the independent evaluation shop — not from Moonshot's own launch post.

MetricKimi K3Claude Fable 5GPT-5.6 Sol
AA Intelligence Index57 (#7 of 190)60 (leader)59
Context window1M tokens1M tokens
API price (in / out, per 1M)$3 / $15$10 / $50$5 / $30
Blended cost (0.25 in / 0.75 out)$12.00 / MTok$40.00 / MTok$23.75 / MTok
Output speed (AA)32.6 tok/s (#147 of 190)
On TulexAI7 PT / message, every plan25 PT, Pro ($23) and up12 PT, every plan

Read that middle column pairing honestly: K3 lands within 3 Index points of Fable 5 at 1/3.3 of the blended cost ($12 vs $40 per million tokens on the standard chat mix). It doesn't beat Fable — nothing currently does on absolute intelligence, and Fable also leads long-horizon agent work with 56% on AA-Briefcase. But "95% of the leader for 30% of the price" is exactly the kind of trade-off that should change what you reach for by default.

API pricing has one more lever worth knowing: cached input drops to $0.30 per million tokens — 90% off (verified against Moonshot's pricing docs). For workloads that repeatedly hit the same long document, the effective cost falls even further below the frontier.

The trade-off nobody headlines: K3 is slow on purpose

We sell honest billing, so here's the honest catch. On Artificial Analysis' measurements, K3 outputs at 32.6 tokens per second — ranking #147 of 190 models — and takes roughly 161 seconds to first token on their test. That's not a bug or an overloaded server; K3 is a deep-reasoning model that thinks extensively before it answers. It is brilliant, and it is not snappy.

What that means in practice: you send a hard question, you wait — sometimes a couple of minutes — and then you get an answer that reads like it came from a model three price tiers up. For a 40-page contract review, that's a fantastic trade. For "rewrite this sentence three ways," it's agony. Pick your tool for the job:

  • Use Kimi K3 for: long documents (the 1M context swallows entire codebases, contract stacks, or book manuscripts), multi-step reasoning problems, research synthesis, and any task where you'd normally pay Fable-tier prices for careful analysis. At 7 PT it's the cheapest frontier-class analysis on TulexAI.
  • Don't use Kimi K3 for: rapid back-and-forth chat, quick drafts, or anything where a two-minute wait breaks your flow. For snappy conversation, Kimi K2.6 at 2 PT or GPT-5.6 Luna at 2.5 PT will feel ten times better and cost a third as much.

K3 vs Fable 5 vs GPT-5.6 Sol: who should pick which?

  • Pick Kimi K3 when: the question is hard, the document is long, and you can wait. It's the value play for deep reasoning — 7 PT on every plan, including Basic at $11/mo.
  • Pick Claude Fable 5 when: you need the absolute ceiling. It leads the Intelligence Index at 60 and wins long-horizon agentic work (56% on AA-Briefcase). It runs 25 PT per message and requires the Pro plan ($23/mo) or higher.
  • Pick GPT-5.6 Sol when: the work is coding. Sol tops the AA Coding Agent Index outright (80, vs Terra's 77 and Luna's 75) and is unusually token-efficient at ~15K output tokens per Index task. 12 PT, every plan.

The rest of the Kimi family just arrived too

As of today, TulexAI carries two more Moonshot models alongside K3:

ModelAPI price (in / out, per 1M)ContextWhat it's forOn TulexAI
Kimi K2.6$0.95 / $4.00262KFast everyday chat; text, image and video input; thinking and non-thinking modes2 PT / message
Kimi K2.7 Code$1.90 / $8.00 (cached $0.38)262KAgentic coding (kimi-k2.7-code-highspeed)4 PT / message
Kimi K3$3.00 / $15.00 (cached $0.30)1MDeep reasoning, long documents7 PT / message

All three are available on every plan. That gives you a clean ladder inside one vendor family: K2.6 for speed, K2.7 Code for agentic coding runs, K3 when you need the big brain.

What this costs you in practice

The subscription math is the quiet headline. A month of heavy K3 use through the API means managing keys, top-ups and another dashboard. On TulexAI, K3 sits in the same dropdown as Fable 5, GPT-5.6, Grok and 50 models in all — one login, one bill, and a live token meter showing exactly what each message costs before you send it. Ask K3 the hard question at 7 PT, then sanity-check its answer with Sol at 12 PT — that side-by-side workflow is the whole point, and it's why an aggregator beats a stack of single-vendor subscriptions. See the full breakdown on the pricing page, or compare the approach to a single ChatGPT subscription on our ChatGPT alternative page.

Try Kimi K3 free on TulexAI — no card, and your first prompts are on us.

Frequently Asked Questions

Is Kimi K3 as good as Claude Fable 5?

Close, but not quite. On the Artificial Analysis Intelligence Index, Fable 5 leads at 60 to K3's 57 (#7 of 190), and Fable also wins long-horizon agent evals. But K3's blended API cost is $12 per million tokens versus Fable's $40 — about 1/3.3 the price. On TulexAI that gap is 7 PT versus 25 PT per message, and Fable 5 requires the Pro plan ($23/mo) or higher while K3 is on every plan.

Why is Kimi K3 so slow?

By design. K3 is a deep-reasoning model that thinks extensively before answering — Artificial Analysis measured 32.6 tokens/sec output (#147 of 190) and roughly 161 seconds to first token on their test. It trades latency for answer quality. For fast conversational work, use Kimi K2.6 (2 PT) or GPT-5.6 Luna (2.5 PT) instead.

How much does Kimi K3 cost?

Via Moonshot's API: $3 per million input tokens, $15 per million output tokens, with cached input at $0.30 (90% off), verified against Moonshot's pricing docs in July 2026. On TulexAI it's 7 Platform Tokens per message on any plan, from $11/mo — no separate Moonshot account needed.

What is Kimi K3's context window?

1 million tokens — matching Claude Fable 5 and. Combined with the deep-reasoning behavior, that makes it one of the strongest picks anywhere for analyzing very long documents in a single pass.

Is Kimi K3 good for coding?

It's capable, but it isn't the specialist. GPT-5.6 Sol leads the Artificial Analysis Coding Agent Index (score 80), and Moonshot's own agentic coding pick is Kimi K2.7 Code, live on TulexAI at 4 PT per message. Reach for K3 when the problem is reasoning-heavy — architecture decisions, tricky debugging analysis — rather than high-volume code generation.

Can I use Kimi K3 without a Moonshot account?

Yes. TulexAI serves Kimi K3, K2.6 and K2.7 Code through the Moonshot API on every plan, from $11/mo, alongside Claude, GPT-5.6, Grok and 50 models in all. Sign up free and your first prompts are on us.

Kimi K3Moonshot AIClaude Fable 5GPT-5.6model comparison

Ready to consolidate your AI tools?

40+ AI models — GPT-5.6, Claude Opus 5, Gemini, Flux, Sora & more. One subscription from $11/mo.

Try 1 Free Prompt

Continue reading