Gemini 3.8 Flash is three times faster than GPT-5.6 Luna. Its price doubles on 1 January.
Gemini 3.8 Flash scores 41 to Luna's 38 at 332 tokens per second. Google lists $0.75 / $3.75 until 31 December and $1.50 / $7.50 from 1 January 2027.
GTM Architect & Growth Operator · Now · 16 September 2026
TL;DR · Key insights
- On Artificial Analysis's Intelligence Index v4.3, Gemini 3.8 Flash scores 41.19 and GPT-5.6 Luna scores 37.50. Close to four points to Gemini.
- The gap that stands out is speed: Gemini 3.8 Flash generates 332 output tokens per second, Luna 115. For anything a user waits on, that is felt.
- Luna is the cheaper model today, at $0.20 / $1.20 against Gemini's $0.75 / $3.75. Google's own pricing page says Gemini 3.8 Flash moves to $1.50 / $7.50 on 1 January 2027.
- From January the price gap roughly doubles, to 7.5x on input and 6.25x on output. Build on Gemini 3.8 Flash for its speed, and price the product on January's numbers, not September's.
Most comparisons in the cheap tier come down to price per million tokens. This one comes down to a number that rarely decides anything at this level: how fast the answer arrives.
And to a date on Google’s pricing page.
| Info | Gemini 3.8 Flash | GPT-5.6 Luna |
|---|---|---|
| Input / output per 1M, now | $0.75 / $3.75 | $0.20 / $1.20 |
| Input / output from 1 Jan 2027 | $1.50 / $7.50 | No change stated |
| Intelligence Index v4.3 | 41.19 (high) | 37.50 (max) |
| Output speed, tokens/s | 332.0 | 115.0 |
| Blended price per 1M | $0.58 | $0.17 |
| Cost to run the full index | $1,623 | $320 |
| Context window | 1M | 1,050,000 |
Prices from Google's Gemini API pricing page and OpenAI's GPT-5.6 Luna model page. Index to two decimals from Artificial Analysis's v4.3 announcement; speed, blended price and run cost from its model pages. Read on 16 September 2026.
Speed is the story
Gemini 3.8 Flash generates 332 output tokens per second on Artificial Analysis’s measurement. Luna generates 115. That is close to three times as fast, from a model that also scores higher.
At this tier speed usually gets mentioned and then ignored, because both models are fast enough. At 332 tokens per second it stops being a footnote. A 1,000-token answer takes about three seconds to generate on Gemini and about nine on Luna, before any thinking time. In a support chat, or an agent loop where a person watches each step, that difference is the product.
For batch work nobody watches, tokens per second is nearly irrelevant and price per token decides. For anything interactive, the order flips. Know which one you are building before you compare prices.
The scores are close
On Intelligence Index v4.3, Gemini 3.8 Flash at high reasoning scores 41.19 and Luna at max effort scores 37.50. Under four points.
That is enough to call Gemini modestly stronger. On the same scale Luna sits just under a point below Claude Sonnet 5, so Gemini leads Anthropic’s mid tier by nearly three points and Luna by nearly four.
Luna is cheaper, and the gap grows in January
Today Luna lists at $0.20 / $1.20 and Gemini 3.8 Flash at $0.75 / $3.75, so Gemini costs 3.75 times as much on input and about 3.1 times on output. Artificial Analysis’s numbers agree. Its blended prices are $0.58 against $0.17, and running the full index cost $1,623 against $320.
Then there is the line on Google’s pricing page. Gemini 3.8 Flash’s prices are listed through 31 December 2026, with $1.50 input and $7.50 output from 1 January 2027. Context caching and batch prices double on the same date.
From January the ratio becomes 7.5 times on input and 6.25 times on output, assuming Luna’s price does not move.
Luna has a pricing trap of its own. Above 272K input tokens OpenAI bills the whole request at 2x input and 1.5x output, which I explained in the 272K pricing cliff.
Which one I would use
| The work | Model | Why |
|---|---|---|
| Interactive assistant or chat | Gemini 3.8 Flash | Close to three times faster, and slightly stronger. Latency is the product. |
| Agent loops a person watches | Gemini 3.8 Flash | Speed compounds across every visible step. |
| Batch classification or extraction | GPT-5.6 Luna | Nobody waits, so price decides, and Luna is several times cheaper. |
| A product priced for 2027 | Model both at January rates | Gemini's listed price doubles on 1 January. The gap to Luna widens with it. |
Speed for interactive work, price for batch work, and January's numbers for anything that lasts.
For Gemini 3.8 Flash against Anthropic’s mid tier, see Gemini 3.8 Flash against Claude Sonnet 5. For another cheap-tier challenger to Luna, DeepSeek V4.1 Flash against Luna.
What would change my mind
Google revising the January price. It is listed, but pricing pages change.
A Luna speed improvement. The whole case for paying more for Gemini at this tier rests on a 3x speed gap.
Cost per task on an interactive workload. Across Artificial Analysis’s index Gemini 3.8 Flash wrote 170M tokens and Luna 150M, so the extra writing eats only a small part of Gemini’s speed lead there. A workload that makes it think far longer than Luna would change that.
Questions people asked
Is Gemini 3.8 Flash better than GPT-5.6 Luna?
Somewhat, and much faster: 41.19 against 37.50 on the Intelligence Index v4.3, and 332 against 115 output tokens per second. Luna is cheaper.
Which is cheaper, Gemini 3.8 Flash or GPT-5.6 Luna?
Luna, at $0.20 / $1.20 against Gemini’s $0.75 / $3.75 today and $1.50 / $7.50 from 1 January 2027.
Is Gemini 3.8 Flash’s price going up?
Yes. From 1 January 2027 Google’s pricing page lists $1.50 input and $7.50 output, and the prices for context caching and batch double on the same date.
Should I use Gemini 3.8 Flash or GPT-5.6 Luna?
Gemini 3.8 Flash where latency decides, Luna where cost per token decides. Price any Gemini product on its January 2027 rates.
If you are choosing a fast model for an interactive product, the contact page is the fastest route to me.