PromptsRush
Prompts

Browse

All PromptsThe full curated libraryPrompts GalleryVisual, Pinterest-style browsingImage PromptsMidjourney, DALL·E & SDXLVideo PromptsRunway, Kling & SoraText & TemplatesChatGPT & Claude system prompts

Discover

CategoriesExplore prompts by topicAI ModelsBest prompts per modelPrompt PacksCommunity, passcode-protectedSubmit a PromptShare with the community

For Creators

Turn prompts into followers

Share passcode-protected prompt packs and grow your audience with Auto DM.

Start sharing
Marketplace

Explore

Shared PromptsPasscode-protected prompt packsAI SkillsNewInstallable Agent SkillsDesign SystemsNewLive themes & design tokens

Contribute

Submit a PromptPublish a prompt packSubmit a SkillShip an Agent SkillSubmit a DesignShare a design system

New · Skills

Teach your AI new tricks

Install ready-made skills for Claude, ChatGPT, Gemini, n8n & more.

Browse skills
Learn

Learning Tracks

Prompt EngineeringWrite prompts that deliverAI SkillsBuild & ship Agent SkillsAI AutomationWorkflows, agents & MCPDesign SystemsOn-brand UI with AI

More

Learning HubAll tracks · 40+ lessonsBlogGuides, news & deep diveseBooksPremium prompt packs & guides

100% Free

Learn AI, the practical way

From fundamentals to advanced across four hands-on tracks — no fluff.

Explore the hub
Blog
LoginSign Up
PromptsRush

The ultimate directory for finding, sharing, and managing production-ready AI prompts, system instructions, and advanced templates.

TwitterGitHubYouTubeInstagramEmail

Platform

  • Home
  • Browse Prompts
  • Marketplace
  • Skills
  • Categories
  • Submit a Skill

Top Categories

  • Image PromptPopular
  • Video Prompts
  • Text Templates

Company

  • Privacy Policy
  • Terms of Service
  • Contact Us

Subscribe on YouTube

New AI prompt & skills tutorials every week.

Subscribe

© 2026 PromptsRush. Crafted with & Passion.

All systems operational
HomeBlogNews
News

Google Gemini 3.8 Flash vs Opus 5 vs GPT 5.6

Google says Gemini 3.8 Flash beats Claude Opus 5 and GPT-5.6 at a sixth of the price. The HLE result is a 0.5-point tie — but the price at which it ties is the real story.

P
PromptsRushSeptember 2, 2026
•8 min read562 views

Advertisement

Google Gemini 3.8 Flash vs Opus 5 vs GPT 5.6

Google's Gemini 3.8 Flash landed on 2 September 2026 with benchmark charts placing it above Claude Opus 5 and GPT-5.6 — while costing roughly a sixth as much per token.

That claim deserves scrutiny rather than a headline, because the benchmarks come from Google, on Google's page, chosen by Google. So this comparison does two things: reports the numbers exactly as published, then examines what they actually support. The short version is that one of the three headline results is a statistical tie, and the real difference between these models is not the scores at all.

The Three Models

Gemini 3.8 FlashClaude Opus 5GPT-5.6 Sol
Model IDgemini-3.8-flashclaude-opus-5gpt-5.6-sol
ClassFast workhorseFrontierFlagship
Context~1.05M1M~1.05M
Max output65,536128,000128,000
Input modalitiesText, image, video, audio, PDFText, image, PDFText, image
Reasoning controllow / medium / highlow → max (5 levels)none → max (6 levels)
Knowledge cutoffNot publishedNot published16 Feb 2026

Note the modality asymmetry: Gemini 3.8 Flash is the only one of the three that ingests video and audio natively. If your input is a meeting recording or a screen capture, this comparison is already over.

Price: The Real Gap

A small cost and an enormous cost producing an identical result

Per million tokens, USD:

ModelInputOutputvs Flash (input)
Gemini 3.8 Flash (intro)$0.75$3.75—
Gemini 3.8 Flash (from Jan 2027)$1.50$7.502×
GPT-5.6 Luna$0.20$1.200.27×
GPT-5.6 Terra$2.00$12.002.7×
GPT-5.6 Sol$4.00$20.005.3×
Claude Opus 5$5.00$25.006.7×

Two things the headline framing hides. First, Gemini Flash is not the cheapest model here — GPT-5.6 Luna undercuts it nearly fourfold. Second, Flash's advantage halves in January when introductory pricing ends.

GPT-5.6 also charges more above a context threshold (Sol rises to $8/$30 on long context), and both OpenAI and Anthropic discount heavily for cached input — Sol's cached input is $0.40, and Opus 5 cache reads run at a tenth of base input price. On a workload with a large stable prefix, the headline gap narrows considerably.

The Benchmarks, As Published

These are Google's figures from its own DeepMind pages, reproduced exactly.

HLE-Verified — expert multidisciplinary reasoning

Three columns of near-identical height with a laser line across their tops
ModelScore
Gemini 3.8 Flash54.9%
GPT-5.6 Sol54.5%
Claude Opus 554.4%
Gemini 3.7 Flash53.6%
GPT-5.6 Terra51.1%
Claude Sonnet 531.0%
The top three are separated by 0.5 percentage points. On a benchmark of this size that is a tie, not a ranking — run it again on a different day and the order could reshuffle. Reporting "Gemini beats Opus 5 and GPT-5.6 on HLE" is technically true and practically meaningless.

What is meaningful: a model at $0.75/$3.75 landed in a dead heat with models at $4/$20 and $5/$25. The story is not that Flash won. It is that the price of that score fell by roughly 85%.

Vals Finance Agent v2

ModelScore
Gemini 3.8 Flash61.4%
Gemini 3.7 Flash59.0%
Claude Opus 558.6%
GPT-5.6 Terra54.4%
Claude Sonnet 553.9%
GPT-5.6 Sol53.8%

A 2.8-point lead over Opus 5 — narrow but outside tie territory. The oddity worth flagging: GPT-5.6 Terra outscores GPT-5.6 Sol here, despite Sol being the flagship at double the price. When a vendor's cheaper model beats its flagship on your benchmark, it is a sign the benchmark is measuring something specific rather than general capability.

Harvey Legal Agent

ModelScore
Gemini 3.8 Flash10.0%
Gemini 3.7 Flash8.8%
Claude Opus 56.7%
Claude Sonnet 55.0%
GPT-5.6 Sol2.5%
GPT-5.6 Terra0.8%

Read the axis before the ranking. The winning score is 10%. Every model here fails at least 90% of the tasks. Gemini leading by 3.3 points over Opus 5 is a 1.5× relative difference on a benchmark where the honest summary is "none of these models can do this job yet."

If you are evaluating AI for legal work, the actionable finding is not which model to pick — it is that this task class remains unsolved and needs a human in the loop regardless of vendor.

How Much to Trust These Numbers

Three caveats that apply to every vendor benchmark chart, this one included:

  • Google selected the benchmarks. A vendor publishes the evaluations it does well on. The absence of GPQA, AIME, MMMU, LiveCodeBench and standard SWE-bench from this comparison is information in itself.
  • Configuration is rarely disclosed. Reasoning effort, scaffolding and retry policy move agentic scores substantially. Opus 5 runs five effort levels and GPT-5.6 six — a comparison against them at low effort is a different result than at max, and the settings are not published.
  • Google's own charts are not a clean sweep, which is a point in their favour. On CWE-Bench patching, Google reports its Flash Cyber variant at 47.2% against a leading frontier model's 47.8% — a published loss. That is a more credible presentation than uniform victory.

The reasonable posture: treat these as evidence that Flash is competitive with the frontier tier on some agentic tasks, not that it surpasses it. Then run your own eval, because the only benchmark that predicts your results is your workload.

What Each Model Is Actually For

Gemini 3.8 Flash

The price-performance pick, and the only one that ingests video and audio. Strongest fit: high-volume agentic pipelines, document and media processing, finance-style analysis, anything where per-token cost is the binding constraint. Weakest fit: work needing more than 64K of output, or the deepest reasoning settings — its three thinking levels top out below what Opus 5 and GPT-5.6 offer.

Claude Opus 5

The frontier option when correctness outranks cost. Five effort levels up to max, 128K output, adaptive thinking on by default, and the deepest agentic tooling of the three — task budgets, compaction, context editing, mid-conversation system messages. Strongest fit: long-horizon coding and agent work where a wrong answer is expensive — the workloads our Opus 5 prompt collection is built around. Weakest fit: high-volume routine calls, where you pay 6.7× for capability you are not exercising.

GPT-5.6

The one that is really three models, and the variant choice matters more than the family. Luna at $0.20/$1.20 is the cheapest serious model in this comparison and belongs in any high-volume cost calculation. Terra at $2/$12 is the balanced middle — and outscored Sol on Vals Finance. Sol is the flagship for demanding professional work. Six reasoning levels including none gives the finest-grained cost control of the three.

Which Should You Choose?

If you…Choose
Run high-volume agentic pipelinesGemini 3.8 Flash
Process video or audio inputGemini 3.8 Flash — the only option here
Need the absolute cheapest per tokenGPT-5.6 Luna ($0.20/$1.20)
Do long-horizon coding or agent workClaude Opus 5
Need output longer than 64K tokensOpus 5 or GPT-5.6 (128K)
Need the deepest reasoning availableOpus 5 (max) or GPT-5.6 (max)
Want fine-grained cost control per callGPT-5.6 — six reasoning levels including none
Have a large stable cached prefixRe-run the maths — caching narrows the gap sharply
Are budgeting past December 2026Price Flash at $1.50/$7.50, not $0.75/$3.75
Are automating legal analysisNone yet — the best score here is 10%

The Verdict

The headline is wrong and the underlying story is bigger than the headline. Gemini 3.8 Flash did not decisively beat Claude Opus 5 and GPT-5.6 — on HLE the three are separated by half a point, which is noise. On Google's two agentic benchmarks Flash leads more clearly, but they are Google's chosen benchmarks and one of them has a 10% ceiling.

What is genuinely notable is the price at which Flash reaches that dead heat. A workhorse model matching frontier scores at a sixth of the per-token cost changes the default: the burden of proof has shifted onto the expensive models. You should now be able to say why a task needs Opus 5 or Sol rather than assuming it does.

There are good answers to that question — 128K output, the deepest reasoning settings, Opus 5's agentic tooling, and the reliability that matters when errors are costly. But "it is the best model" is no longer one of them by default, and that is a real change from six months ago, when Opus 5 comfortably outclassed Gemini 3.6 Flash.

Two practical notes to close on. Model your 2027 costs at Flash's post-January rate. And if you have a large cached prefix, redo the arithmetic with cache pricing before concluding anything — it moves the answer more than any benchmark on this page.

Recommended · Genspark

Try Genspark — the AI super-agent

Genspark researches, plans and acts across the web for you — multi-step agentic workflows in one prompt.

Try Genspark Free

Affiliate link · We may earn a commission

Keep Reading

If you are weighing the subscription route rather than the API, Claude pricing explained covers Pro, Max, Team and Enterprise. Or browse all guides and prompts on PromptsRush.

❓

Frequently Asked Questions

10 questions answered

On Google's published benchmarks it edges ahead — HLE-Verified 54.9% vs 54.4%, Vals Finance 61.4% vs 58.6%, Harvey Legal 10.0% vs 6.7%. But the HLE gap is half a percentage point, which is a statistical tie, and these are vendor-selected benchmarks. Opus 5 still leads on max output (128K vs 64K), reasoning depth and agentic tooling.
Per million tokens: Gemini 3.8 Flash $0.75/$3.75 (introductory, doubling to $1.50/$7.50 in January 2027), GPT-5.6 Luna $0.20/$1.20, GPT-5.6 Terra $2/$12, GPT-5.6 Sol $4/$20, Claude Opus 5 $5/$25. Opus 5 costs 6.7× Flash on input — though GPT-5.6 Luna is cheaper than Flash.
Gemini 3.8 Flash at 54.9%, followed by GPT-5.6 Sol at 54.5% and Claude Opus 5 at 54.4%. Treat that as a three-way tie rather than a ranking — a 0.5-point spread is within noise. The meaningful part is that a model costing a sixth as much reached the same score.
Treat it as directional. Google chose the benchmarks, and notable ones are absent — GPQA, AIME, MMMU, LiveCodeBench and standard SWE-bench do not appear. Reasoning effort and scaffolding are not disclosed, which matters when Opus 5 has five effort levels and GPT-5.6 six. To Google's credit the charts are not a clean sweep: it publishes a narrow CWE-Bench loss at 47.2% against 47.8%.
Only Gemini 3.8 Flash. It accepts text, image, video, audio and PDF natively. Claude Opus 5 handles text, image and PDF; GPT-5.6 handles text and image. If your input is a recording or screen capture, Gemini is the only one of the three that takes it directly.
None of them, on this evidence. The top score on Harvey's Legal Agent Benchmark is Gemini 3.8 Flash at 10.0% — meaning it fails 90% of tasks. Opus 5 scores 6.7% and GPT-5.6 Sol 2.5%. The actionable conclusion is that this task class needs a human in the loop regardless of which vendor you choose.
On Vals Finance Agent v2, Terra scores 54.4% against Sol's 53.8% despite Sol being the flagship at double the price. When a vendor's cheaper model outperforms its flagship on a benchmark, it usually indicates the benchmark measures something narrow rather than general capability — and it is a good reason to test both on your own workload.
They are effectively tied at around one million tokens — Gemini 3.8 Flash and GPT-5.6 at roughly 1.05M, Claude Opus 5 at 1M. Output is where they diverge: Opus 5 and GPT-5.6 support up to 128K output tokens, while Gemini 3.8 Flash caps at 65,536.
Considerably, yes. GPT-5.6 Sol's cached input drops to $0.40 and Claude Opus 5 cache reads cost a tenth of base input price. On a workload with a large stable prefix the headline gap narrows sharply, so redo the arithmetic with cache rates before concluding Flash is cheapest for your case.
Test it for high-volume agentic work, document processing and anything cost-constrained — the value case is strong. Stay on Opus 5 for long-horizon coding, work needing more than 64K output, the deepest reasoning settings, or where an error is expensive. The change is that you should now be able to justify the frontier model rather than defaulting to it.
Back to Blog

Table of Contents

In this article

  • 1The Three Models
  • 2Price: The Real Gap
  • 3The Benchmarks, As Published
  • HLE-Verified — expert multidisciplinary reasoning
  • Vals Finance Agent v2
  • Harvey Legal Agent
  • 4How Much to Trust These Numbers
  • 5What Each Model Is Actually For
  • Gemini 3.8 Flash
  • Claude Opus 5
  • GPT-5.6
  • 6Which Should You Choose?
  • 7The Verdict
  • 8Keep Reading

Recent Posts

22 Prompts to Improve Landing Page Conversion Rates

Sep 7 · 15 min

20 Prompts to Improve an Ugly AI-Generated Website

Sep 7 · 15 min

20 Lead Generation Tools with Top-Notch AI Integrations

Sep 7 · 21 min

25+ Vibe Coding Prompts and AI Tools for Building Beautiful Sites

Sep 5 · 16 min

How to Create a Landing Page using AI with Prompts

Sep 5 · 12 min

Category

News

Advertisement

You May Also Like

News

Google Gemini 3.8 Flash: What's New?

Sep 27 min
News

ChatGPT Statistics (2026): Usage, Trend, Market & Growth

Aug 257 min
C
News

Claude AI Stats 2026: User Growth, Market, Trends & Finance

Aug 138 min