What Is Kimi K3? Moonshot's Frontier-Class Reasoner, Explained
Moonshot AI's Kimi K3 scores within 3 points of Claude Fable 5 at a third of the cost — but it thinks before it talks. The verified numbers, the speed trade-off, and when to pick it.
TL;DR: Kimi K3 is Moonshot AI's frontier reasoning model, launched July 16, 2026. Per Artificial Analysis it scores 57 on the Intelligence Index — #7 of 190 models, within 3 points of Claude Fable 5's leading 60 — at roughly 1/3.3 of Fable's blended API cost. The catch: it's slow. 32.6 tokens/sec output and ~161 seconds to first token make it a think-first model, not a chat sprinter. On TulexAI it costs 7 PT per message on every plan, from $11/mo.
What is Kimi K3?
Kimi K3 is the flagship model from Moonshot AI, the Beijing lab behind the Kimi family of open-weight models. It shipped on July 16, 2026, and the pitch is simple: frontier-class deep reasoning with a 1-million-token context window, priced well below the Western frontier. Simon Willison's launch-day writeup (simonwillison.net) captured why people paid attention — this isn't a budget model punching up on one cherry-picked benchmark; it's a top-ten model on a neutral scoreboard, from a lab most casual users hadn't tried yet.
K3 went live on TulexAI the same day it launched. If you've been running the same prompt through Claude, GPT and Grok to see who wins, there's now a fourth serious contender in the dropdown.
The numbers, from a neutral scoreboard
All benchmark figures below come from Artificial Analysis, the independent evaluation shop — not from Moonshot's own launch post.
| Metric | Kimi K3 | Claude Fable 5 | GPT-5.6 Sol |
|---|---|---|---|
| AA Intelligence Index | 57 (#7 of 190) | 60 (leader) | 59 |
| Context window | 1M tokens | 1M tokens | — |
| API price (in / out, per 1M) | $3 / $15 | $10 / $50 | $5 / $30 |
| Blended cost (0.25 in / 0.75 out) | $12.00 / MTok | $40.00 / MTok | $23.75 / MTok |
| Output speed (AA) | 32.6 tok/s (#147 of 190) | — | — |
| On TulexAI | 7 PT / message, every plan | 25 PT, Pro ($23) and up | 12 PT, every plan |
Read that middle column pairing honestly: K3 lands within 3 Index points of Fable 5 at 1/3.3 of the blended cost ($12 vs $40 per million tokens on the standard chat mix). It doesn't beat Fable — nothing currently does on absolute intelligence, and Fable also leads long-horizon agent work with 56% on AA-Briefcase. But "95% of the leader for 30% of the price" is exactly the kind of trade-off that should change what you reach for by default.
API pricing has one more lever worth knowing: cached input drops to $0.30 per million tokens — 90% off (verified against Moonshot's pricing docs). For workloads that repeatedly hit the same long document, the effective cost falls even further below the frontier.
The trade-off nobody headlines: K3 is slow on purpose
We sell honest billing, so here's the honest catch. On Artificial Analysis' measurements, K3 outputs at 32.6 tokens per second — ranking #147 of 190 models — and takes roughly 161 seconds to first token on their test. That's not a bug or an overloaded server; K3 is a deep-reasoning model that thinks extensively before it answers. It is brilliant, and it is not snappy.
What that means in practice: you send a hard question, you wait — sometimes a couple of minutes — and then you get an answer that reads like it came from a model three price tiers up. For a 40-page contract review, that's a fantastic trade. For "rewrite this sentence three ways," it's agony. Pick your tool for the job:
- Use Kimi K3 for: long documents (the 1M context swallows entire codebases, contract stacks, or book manuscripts), multi-step reasoning problems, research synthesis, and any task where you'd normally pay Fable-tier prices for careful analysis. At 7 PT it's the cheapest frontier-class analysis on TulexAI.
- Don't use Kimi K3 for: rapid back-and-forth chat, quick drafts, or anything where a two-minute wait breaks your flow. For snappy conversation, Kimi K2.6 at 2 PT or GPT-5.6 Luna at 2.5 PT will feel ten times better and cost a third as much.
K3 vs Fable 5 vs GPT-5.6 Sol: who should pick which?
- Pick Kimi K3 when: the question is hard, the document is long, and you can wait. It's the value play for deep reasoning — 7 PT on every plan, including Basic at $11/mo.
- Pick Claude Fable 5 when: you need the absolute ceiling. It leads the Intelligence Index at 60 and wins long-horizon agentic work (56% on AA-Briefcase). It runs 25 PT per message and requires the Pro plan ($23/mo) or higher.
- Pick GPT-5.6 Sol when: the work is coding. Sol tops the AA Coding Agent Index outright (80, vs Terra's 77 and Luna's 75) and is unusually token-efficient at ~15K output tokens per Index task. 12 PT, every plan.
The rest of the Kimi family just arrived too
As of today, TulexAI carries two more Moonshot models alongside K3:
| Model | API price (in / out, per 1M) | Context | What it's for | On TulexAI |
|---|---|---|---|---|
| Kimi K2.6 | $0.95 / $4.00 | 262K | Fast everyday chat; text, image and video input; thinking and non-thinking modes | 2 PT / message |
| Kimi K2.7 Code | $1.90 / $8.00 (cached $0.38) | 262K | Agentic coding (kimi-k2.7-code-highspeed) | 4 PT / message |
| Kimi K3 | $3.00 / $15.00 (cached $0.30) | 1M | Deep reasoning, long documents | 7 PT / message |
All three are available on every plan. That gives you a clean ladder inside one vendor family: K2.6 for speed, K2.7 Code for agentic coding runs, K3 when you need the big brain.
What this costs you in practice
The subscription math is the quiet headline. A month of heavy K3 use through the API means managing keys, top-ups and another dashboard. On TulexAI, K3 sits in the same dropdown as Fable 5, GPT-5.6, Grok and 50 models in all — one login, one bill, and a live token meter showing exactly what each message costs before you send it. Ask K3 the hard question at 7 PT, then sanity-check its answer with Sol at 12 PT — that side-by-side workflow is the whole point, and it's why an aggregator beats a stack of single-vendor subscriptions. See the full breakdown on the pricing page, or compare the approach to a single ChatGPT subscription on our ChatGPT alternative page.
Try Kimi K3 free on TulexAI — no card, and your first prompts are on us.
Frequently Asked Questions
Is Kimi K3 as good as Claude Fable 5?
Close, but not quite. On the Artificial Analysis Intelligence Index, Fable 5 leads at 60 to K3's 57 (#7 of 190), and Fable also wins long-horizon agent evals. But K3's blended API cost is $12 per million tokens versus Fable's $40 — about 1/3.3 the price. On TulexAI that gap is 7 PT versus 25 PT per message, and Fable 5 requires the Pro plan ($23/mo) or higher while K3 is on every plan.
Why is Kimi K3 so slow?
By design. K3 is a deep-reasoning model that thinks extensively before answering — Artificial Analysis measured 32.6 tokens/sec output (#147 of 190) and roughly 161 seconds to first token on their test. It trades latency for answer quality. For fast conversational work, use Kimi K2.6 (2 PT) or GPT-5.6 Luna (2.5 PT) instead.
How much does Kimi K3 cost?
Via Moonshot's API: $3 per million input tokens, $15 per million output tokens, with cached input at $0.30 (90% off), verified against Moonshot's pricing docs in July 2026. On TulexAI it's 7 Platform Tokens per message on any plan, from $11/mo — no separate Moonshot account needed.
What is Kimi K3's context window?
1 million tokens — matching Claude Fable 5 and. Combined with the deep-reasoning behavior, that makes it one of the strongest picks anywhere for analyzing very long documents in a single pass.
Is Kimi K3 good for coding?
It's capable, but it isn't the specialist. GPT-5.6 Sol leads the Artificial Analysis Coding Agent Index (score 80), and Moonshot's own agentic coding pick is Kimi K2.7 Code, live on TulexAI at 4 PT per message. Reach for K3 when the problem is reasoning-heavy — architecture decisions, tricky debugging analysis — rather than high-volume code generation.
Can I use Kimi K3 without a Moonshot account?
Yes. TulexAI serves Kimi K3, K2.6 and K2.7 Code through the Moonshot API on every plan, from $11/mo, alongside Claude, GPT-5.6, Grok and 50 models in all. Sign up free and your first prompts are on us.
Ready to consolidate your AI tools?
40+ AI models — GPT-5.6, Claude Opus 5, Gemini, Flux, Sora & more. One subscription from $11/mo.
Try 1 Free PromptContinue reading
Kimi K2.6 and K2.7 Code Are Live on TulexAI — the Moonshot Family Is Complete
Two new Moonshot AI models land on TulexAI today: Kimi K2.6 at 2 PT and the agentic coder K2.7 Code at 4 PT, joining the K3 frontier reasoner. Every plan, honest prices, and when to pick each one over Claude or GPT.
Claude Fable 5 vs Kimi K3 vs GPT-5.6 Sol: The Real 2026 Frontier Showdown
Anthropic's leader, OpenAI's coding king, and Moonshot's $3/$15 disruptor — the verified Artificial Analysis numbers, the honest caveats, and which one to use for which job. All three live on TulexAI.