PromptsRush
Prompts

Browse

All PromptsThe full curated libraryPrompts GalleryVisual, Pinterest-style browsingImage PromptsMidjourney, DALL·E & SDXLVideo PromptsRunway, Kling & SoraText & TemplatesChatGPT & Claude system prompts

Discover

CategoriesExplore prompts by topicAI ModelsBest prompts per modelPrompt PacksCommunity, passcode-protectedSubmit a PromptShare with the community

For Creators

Turn prompts into followers

Share passcode-protected prompt packs and grow your audience with Auto DM.

Start sharing
Marketplace

Explore

Shared PromptsPasscode-protected prompt packsAI SkillsNewInstallable Agent SkillsDesign SystemsNewLive themes & design tokens

Contribute

Submit a PromptPublish a prompt packSubmit a SkillShip an Agent SkillSubmit a DesignShare a design system

New · Skills

Teach your AI new tricks

Install ready-made skills for Claude, ChatGPT, Gemini, n8n & more.

Browse skills
Learn

Learning Tracks

Prompt EngineeringWrite prompts that deliverAI SkillsBuild & ship Agent SkillsAI AutomationWorkflows, agents & MCPDesign SystemsOn-brand UI with AI

More

Learning HubAll tracks · 40+ lessonsBlogGuides, news & deep diveseBooksPremium prompt packs & guides

100% Free

Learn AI, the practical way

From fundamentals to advanced across four hands-on tracks — no fluff.

Explore the hub
Blog
LoginSign Up
PromptsRush

The ultimate directory for finding, sharing, and managing production-ready AI prompts, system instructions, and advanced templates.

TwitterGitHubYouTubeInstagramEmail

Platform

  • Home
  • Browse Prompts
  • Marketplace
  • Skills
  • Categories
  • Submit a Skill

Top Categories

  • Image PromptPopular
  • Video Prompts
  • Text Templates

Company

  • Privacy Policy
  • Terms of Service
  • Contact Us

Subscribe on YouTube

New AI prompt & skills tutorials every week.

Subscribe

© 2026 PromptsRush. Crafted with & Passion.

All systems operational
HomeBlogAI Tools
AI Tools

Kimi AI Pricing (2026): Plans, API Cost, Usage Explained

Kimi's lineup spans a five-fold price range — K3 at $3/M input against K2.5 at $0.60 — so model choice is the whole cost decision. Full API table, batch rates, subscription tiers and the V1 sunset.

P
PromptsRushAugust 25, 2026
•7 min read15 views

Advertisement

Kimi AI Pricing (2026): Plans, API Cost, Usage Explained

Kimi's pricing has an unusual shape. Most model families get cheaper as you move down the lineup by a modest amount. Moonshot's spans a five-fold range: the flagship K3 costs $3.00 per million input tokens, and K2.5 costs $0.60.

That gap is wide enough that model selection is not a preference — it is the entire cost decision. Running a workload on K3 that K2.5 could handle costs five times more on input and five times more on output, for the same request.

The second thing worth knowing is that Moonshot V1 is being sunset on 31 August. If you have anything pointing at it, that is a migration with a deadline rather than a someday task.

The model lineup and the sunset below are confirmed from Moonshot's own platform documentation. The per-token rates are drawn from multiple independent pricing trackers, because the vendor's individual model rate pages were not machine-readable when we checked — treat those figures as well-corroborated but verify before committing budget.

API Pricing at a Glance

A small context container beside a much larger one, showing the difference in capacity
ModelInputCache hitOutputContext
Kimi K3$3.00$0.30$15.001M
Kimi K2.7 Code~$0.95—~$4.00262K
Kimi K2.6$0.95—$4.00262K
Kimi K2.5$0.60—$3.00262K
Moonshot V1Sunsetting 31 August — migrate off

All figures per million tokens, excluding tax, which is calculated at checkout by jurisdiction.

Batch pricing

Batch processing discounts the rate by roughly 40% on the K2 family:

ModelStandard in / outBatch in / out
Kimi K2.6$0.95 / $4.00$0.57 / $2.40
Kimi K2.5$0.60 / $3.00$0.36 / $1.80

For non-interactive work — overnight enrichment, bulk classification, evaluation runs — that is a straightforward 40% saving on a model that is already the cheapest in the lineup. K2.5 on batch at $0.36 input is the floor of Moonshot's pricing.

What You Actually Pay Per Model

The spread across the lineup is the thing to internalise:

ComparisonInputOutput
K3 vs K2.55.0× more expensive5.0× more expensive
K3 vs K2.63.2× more expensive3.75× more expensive
K2.6 vs K2.51.6× more expensive1.3× more expensive

So the ladder has one big step and one small one. Moving from K2.5 to K2.6 costs modestly more. Moving to K3 is a different order of expense entirely — and you are buying two specific things for it.

What K3 Actually Buys You

Two capabilities justify the premium, and neither is "it is newer".

A 1M-token context window, against 262K on the K2 family. That is roughly four times the room. If your workload genuinely needs to hold a very large codebase, a long document set or an extended agent trajectory in a single call, the K2 family cannot do it at any price and the comparison ends there.

Cache-hit input at $0.30 per million — a 90% discount on K3's standard input rate. This is the detail that changes the maths for anyone running repeated calls against a stable prefix. A K3 request whose input is largely cached costs $0.30 per million rather than $3.00, which lands it below K2.6's uncached input rate of $0.95.

The non-obvious consequence: if your prompts share a long stable prefix and cache reliably, K3 can be cheaper on input than the mid-tier models. Structure your calls for cacheability before you conclude the flagship is unaffordable.

Output is where K3 stays expensive regardless — $15.00 per million against $4.00 and $3.00, with no cache discount on generated tokens. So the workloads where K3 makes economic sense are those with large cached inputs and short outputs. Long-generation tasks on K3 get costly quickly.

Genspark

Try Genspark — the AI super-agent

Genspark researches, plans and acts across the web for you — multi-step agentic workflows in one prompt.

Affiliate link · We may earn a commission

Try Genspark Free

Subscription Plans

Alongside the API, Kimi sells consumer and prosumer subscriptions on a tiered ladder using musical tempo names, running from a free tier up to around $199 a month.

TierPricePosition
AdagioFreeEntry — try the product
Moderato~$19/moFirst paid tier
Allegretto~$39/moMid tier
Upper tiersup to ~$199/moHigher allowances

Two honest caveats. The exact composition of the upper tiers was not confirmable from the vendor's own documentation when we checked, so the ~$199 ceiling is the reported top of the range rather than a verified tier price. And subscription allowances are typically expressed as usage quotas rather than raw tokens, so they do not map cleanly onto the API table above.

The practical rule: subscriptions suit interactive use through the product; the API suits anything programmatic. If you are building on top of Kimi, price the API and ignore the subscription ladder entirely.

Which Model Should You Use?

If you...Use
Are doing high-volume classification or extractionK2.5, on batch — the cheapest option available
Want general-purpose quality at low costK2.5, or K2.6 if quality falls short
Are running coding workloadsK2.7 Code
Need more than 262K context in one callK3 — nothing else in the lineup can
Have long stable prefixes that cache wellPrice K3 with cache hits before assuming it is too expensive
Generate long outputsAvoid K3 — output is $15/M with no cache discount
Are running non-interactive jobsBatch, for roughly 40% off
Are still on Moonshot V1Migrate before 31 August

How to Cut Your Kimi Bill

  1. Start at the bottom of the ladder, not the top. Run your evaluations on K2.5 first and only move up when it demonstrably fails. A five-fold price difference makes this the highest-leverage decision available.
  2. Batch anything that is not interactive. Roughly 40% off for accepting latency you probably did not need.
  3. Engineer for cache hits. On K3 this is a 90% input discount. Put your stable system prompt, schema and reference documents at the front of the request and keep them byte-identical between calls.
  4. Watch output length. Output is the expensive side on every model here and the gap widens at K3. Constrain response length explicitly rather than hoping.
  5. Do not pay for context you are not using. K3's 1M window is only worth its premium if you actually fill it. A 20K-token request does not benefit from a 1M ceiling.

The Verdict

K2.5 is the default — the cheapest model in the lineup at $0.60 input and $3.00 output, dropping to $0.36 and $1.80 on batch. Most workloads should start here and stay unless evaluations say otherwise.

K2.6 is the modest upgrade when K2.5 falls short, at roughly 1.6× the input cost, and K2.7 Code is the coding-specific option at similar rates.

K3 is a capability purchase. Buy it for the 1M context window when you genuinely need it, or for cached-prefix workloads where the $0.30 cache rate makes it cheaper than the mid-tier. Do not buy it because it is the newest model — at 5× K2.5 on both input and output, that reasoning gets expensive fast.

And whatever you are running: if it still points at Moonshot V1, the 31 August sunset is the deadline that matters more than any of this.

For how this compares across the market, see our ChatGPT pricing breakdown, Grok pricing and Claude pricing explained.

Recommended · Genspark

Try Genspark — the AI super-agent

Genspark researches, plans and acts across the web for you — multi-step agentic workflows in one prompt.

Try Genspark Free

Affiliate link · We may earn a commission

Keep Reading

More pricing breakdowns: ChatGPT pricing 2026, Grok pricing 2026, Claude pricing explained, ElevenLabs pricing, Hedra pricing, and OpenArt pricing. Or browse all guides and prompts on PromptsRush.

❓

Frequently Asked Questions

10 questions answered

Per million tokens: Kimi K3 is $3.00 input, $0.30 cache-hit and $15.00 output; K2.6 and K2.7 Code are around $0.95 input and $4.00 output; and K2.5 is $0.60 input and $3.00 output. All rates exclude tax, which is applied at checkout by jurisdiction.
Kimi K2.5 at $0.60 input and $3.00 output, dropping to $0.36 and $1.80 on batch — which is the floor of Moonshot's pricing. Most workloads should start there and only move up the ladder when evaluations show it genuinely falls short.
Only for two specific reasons. It has a 1M-token context window against 262K on the K2 family, which nothing else in the lineup can match at any price. And its cache-hit input at $0.30 per million is a 90% discount that can make it cheaper on input than the mid-tier models. Buying it simply for being newest costs 5× K2.5 on both input and output.
1M tokens on K3 and 262,144 tokens on K2.6 and K2.7 Code. If a single call genuinely needs to hold more than 262K, K3 is the only option in the lineup — but a 20K-token request gains nothing from a 1M ceiling, so do not pay the premium for headroom you will not use.
Roughly 40%. K2.6 drops from $0.95 / $4.00 to $0.57 / $2.40, and K2.5 from $0.60 / $3.00 to $0.36 / $1.80 per million tokens. For non-interactive work such as bulk classification, enrichment or evaluation runs, that is a straightforward saving for accepting latency you probably did not need.
A tiered ladder using musical tempo names, from a free Adagio tier through Moderato at around $19 a month and Allegretto at around $39, rising to a reported ceiling near $199. Subscription allowances are expressed as usage quotas rather than raw tokens, so they do not map directly onto API rates.
Subscriptions suit interactive use through the product interface; the API suits anything programmatic. If you are building a product on top of Kimi, price the API per token and ignore the subscription ladder — the two are billed on entirely different mechanics and mixing them up produces bad estimates.
Yes — Moonshot's own platform documentation lists V1 as sunsetting on 31 August. If anything in your stack still points at it, treat that as a migration with a hard deadline rather than a task for later.
In order of impact: start on K2.5 and only move up when evaluations force it, since the lineup spans a five-fold price range; batch anything non-interactive for roughly 40% off; engineer stable prompt prefixes for cache hits, which is a 90% input discount on K3; and constrain output length explicitly, because output is the expensive side on every model here.
K2.7 Code, which is purpose-built for coding workloads and priced in line with K2.6 at around $0.95 input and $4.00 output with a 262K context window. Reach for K3 only if a single call genuinely needs more context than 262K can hold.
Back to Blog

Table of Contents

In this article

  • 1API Pricing at a Glance
  • Batch pricing
  • 2What You Actually Pay Per Model
  • 3What K3 Actually Buys You
  • 4Subscription Plans
  • 5Which Model Should You Use?
  • 6How to Cut Your Kimi Bill
  • 7The Verdict
  • 8Keep Reading

Recent Posts

Best Funnel Builder Software of 2026

Aug 27 · 11 min

Best AI Landing Page Builder Tools

Aug 27 · 11 min

Claude Max vs ChatGPT Pro: Detailed Benefits Comparison

Aug 25 · 11 min

ChatGPT Statistics (2026): Usage, Trend, Market & Growth

Aug 25 · 7 min

25+ Best Prompts for Seedance 2.5

Aug 25 · 28 min

Category

AI Tools

Advertisement

You May Also Like

AI Tools

Best Funnel Builder Software of 2026

Aug 2711 min
AI Tools

Best AI Landing Page Builder Tools

Aug 2711 min
AI Tools

Claude Max vs ChatGPT Pro: Detailed Benefits Comparison

Aug 2511 min