Kimi AI Pricing (2026): Plans, API Cost, Usage Explained
Kimi's lineup spans a five-fold price range — K3 at $3/M input against K2.5 at $0.60 — so model choice is the whole cost decision. Full API table, batch rates, subscription tiers and the V1 sunset.
Advertisement

Kimi's pricing has an unusual shape. Most model families get cheaper as you move down the lineup by a modest amount. Moonshot's spans a five-fold range: the flagship K3 costs $3.00 per million input tokens, and K2.5 costs $0.60.
That gap is wide enough that model selection is not a preference — it is the entire cost decision. Running a workload on K3 that K2.5 could handle costs five times more on input and five times more on output, for the same request.
The second thing worth knowing is that Moonshot V1 is being sunset on 31 August. If you have anything pointing at it, that is a migration with a deadline rather than a someday task.
The model lineup and the sunset below are confirmed from Moonshot's own platform documentation. The per-token rates are drawn from multiple independent pricing trackers, because the vendor's individual model rate pages were not machine-readable when we checked — treat those figures as well-corroborated but verify before committing budget.
API Pricing at a Glance
| Model | Input | Cache hit | Output | Context |
|---|---|---|---|---|
| Kimi K3 | $3.00 | $0.30 | $15.00 | 1M |
| Kimi K2.7 Code | ~$0.95 | — | ~$4.00 | 262K |
| Kimi K2.6 | $0.95 | — | $4.00 | 262K |
| Kimi K2.5 | $0.60 | — | $3.00 | 262K |
| Moonshot V1 | Sunsetting 31 August — migrate off | |||
All figures per million tokens, excluding tax, which is calculated at checkout by jurisdiction.
Batch pricing
Batch processing discounts the rate by roughly 40% on the K2 family:
| Model | Standard in / out | Batch in / out |
|---|---|---|
| Kimi K2.6 | $0.95 / $4.00 | $0.57 / $2.40 |
| Kimi K2.5 | $0.60 / $3.00 | $0.36 / $1.80 |
For non-interactive work — overnight enrichment, bulk classification, evaluation runs — that is a straightforward 40% saving on a model that is already the cheapest in the lineup. K2.5 on batch at $0.36 input is the floor of Moonshot's pricing.
What You Actually Pay Per Model
The spread across the lineup is the thing to internalise:
| Comparison | Input | Output |
|---|---|---|
| K3 vs K2.5 | 5.0× more expensive | 5.0× more expensive |
| K3 vs K2.6 | 3.2× more expensive | 3.75× more expensive |
| K2.6 vs K2.5 | 1.6× more expensive | 1.3× more expensive |
So the ladder has one big step and one small one. Moving from K2.5 to K2.6 costs modestly more. Moving to K3 is a different order of expense entirely — and you are buying two specific things for it.
What K3 Actually Buys You
Two capabilities justify the premium, and neither is "it is newer".
A 1M-token context window, against 262K on the K2 family. That is roughly four times the room. If your workload genuinely needs to hold a very large codebase, a long document set or an extended agent trajectory in a single call, the K2 family cannot do it at any price and the comparison ends there.
Cache-hit input at $0.30 per million — a 90% discount on K3's standard input rate. This is the detail that changes the maths for anyone running repeated calls against a stable prefix. A K3 request whose input is largely cached costs $0.30 per million rather than $3.00, which lands it below K2.6's uncached input rate of $0.95.
The non-obvious consequence: if your prompts share a long stable prefix and cache reliably, K3 can be cheaper on input than the mid-tier models. Structure your calls for cacheability before you conclude the flagship is unaffordable.
Output is where K3 stays expensive regardless — $15.00 per million against $4.00 and $3.00, with no cache discount on generated tokens. So the workloads where K3 makes economic sense are those with large cached inputs and short outputs. Long-generation tasks on K3 get costly quickly.
Subscription Plans
Alongside the API, Kimi sells consumer and prosumer subscriptions on a tiered ladder using musical tempo names, running from a free tier up to around $199 a month.
| Tier | Price | Position |
|---|---|---|
| Adagio | Free | Entry — try the product |
| Moderato | ~$19/mo | First paid tier |
| Allegretto | ~$39/mo | Mid tier |
| Upper tiers | up to ~$199/mo | Higher allowances |
Two honest caveats. The exact composition of the upper tiers was not confirmable from the vendor's own documentation when we checked, so the ~$199 ceiling is the reported top of the range rather than a verified tier price. And subscription allowances are typically expressed as usage quotas rather than raw tokens, so they do not map cleanly onto the API table above.
The practical rule: subscriptions suit interactive use through the product; the API suits anything programmatic. If you are building on top of Kimi, price the API and ignore the subscription ladder entirely.
Which Model Should You Use?
| If you... | Use |
|---|---|
| Are doing high-volume classification or extraction | K2.5, on batch — the cheapest option available |
| Want general-purpose quality at low cost | K2.5, or K2.6 if quality falls short |
| Are running coding workloads | K2.7 Code |
| Need more than 262K context in one call | K3 — nothing else in the lineup can |
| Have long stable prefixes that cache well | Price K3 with cache hits before assuming it is too expensive |
| Generate long outputs | Avoid K3 — output is $15/M with no cache discount |
| Are running non-interactive jobs | Batch, for roughly 40% off |
| Are still on Moonshot V1 | Migrate before 31 August |
How to Cut Your Kimi Bill
- Start at the bottom of the ladder, not the top. Run your evaluations on K2.5 first and only move up when it demonstrably fails. A five-fold price difference makes this the highest-leverage decision available.
- Batch anything that is not interactive. Roughly 40% off for accepting latency you probably did not need.
- Engineer for cache hits. On K3 this is a 90% input discount. Put your stable system prompt, schema and reference documents at the front of the request and keep them byte-identical between calls.
- Watch output length. Output is the expensive side on every model here and the gap widens at K3. Constrain response length explicitly rather than hoping.
- Do not pay for context you are not using. K3's 1M window is only worth its premium if you actually fill it. A 20K-token request does not benefit from a 1M ceiling.
The Verdict
K2.5 is the default — the cheapest model in the lineup at $0.60 input and $3.00 output, dropping to $0.36 and $1.80 on batch. Most workloads should start here and stay unless evaluations say otherwise.
K2.6 is the modest upgrade when K2.5 falls short, at roughly 1.6× the input cost, and K2.7 Code is the coding-specific option at similar rates.
K3 is a capability purchase. Buy it for the 1M context window when you genuinely need it, or for cached-prefix workloads where the $0.30 cache rate makes it cheaper than the mid-tier. Do not buy it because it is the newest model — at 5× K2.5 on both input and output, that reasoning gets expensive fast.
And whatever you are running: if it still points at Moonshot V1, the 31 August sunset is the deadline that matters more than any of this.
For how this compares across the market, see our ChatGPT pricing breakdown, Grok pricing and Claude pricing explained.
Keep Reading
More pricing breakdowns: ChatGPT pricing 2026, Grok pricing 2026, Claude pricing explained, ElevenLabs pricing, Hedra pricing, and OpenArt pricing. Or browse all guides and prompts on PromptsRush.
Frequently Asked Questions
10 questions answered


