DeepSeek V4.1 Flash beats GPT-5.6 Luna on score and speed. Whether it beats it on price depends on the clock.
V4.1 Flash scores 40 to Luna's 38 at almost twice the speed. Off-peak it undercuts Luna; at peak it costs more on input. When the peak hours fall.
GTM Architect & Growth Operator · Now · 16 September 2026
TL;DR · Key insights
- On Artificial Analysis's Intelligence Index v4.3, DeepSeek V4.1 Flash scores 39.55 and GPT-5.6 Luna scores 37.50. V4.1 Flash is also close to twice as fast, at 214.4 against 115 output tokens per second.
- DeepSeek prices by time of day. Off-peak V4.1 Flash costs $0.15 / $0.60 per million tokens, below Luna's $0.20 / $1.20. At peak it costs $0.30 / $1.20, above Luna on input and level on output.
- Peak runs 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays. In Warsaw or Berlin in September that is 03:00 to 06:00 and 08:00 to 12:00, which covers the European working morning.
- V4.1 Flash is open weights and has a 384K max output against Luna's 128K. Luna's edge is a flat price, no peak window to schedule around.
Most API pricing has one number per model. DeepSeek’s has two, and which one you pay depends on what time it is in UTC.
That turns a straightforward cheap-tier comparison into a scheduling question.
| Info | DeepSeek V4.1 Flash | GPT-5.6 Luna |
|---|---|---|
| Input / output per 1M, off-peak | $0.15 / $0.60 | $0.20 / $1.20 |
| Input / output per 1M, peak | $0.30 / $1.20 | $0.20 / $1.20 |
| Cached input per 1M | $0.003 off-peak, $0.006 peak | $0.02 |
| Intelligence Index v4.3 | 39.55 (reasoning, max) | 37.50 (max) |
| Output speed, tokens/s | 214.4 | 115.0 |
| Context / max output | 1M / 384K | 1,050,000 / 128,000 |
| Weights | Open, downloadable | None found |
Prices, peak hours, context and max output from DeepSeek's pricing page and OpenAI's GPT-5.6 Luna model page. Index, speed and weights status from Artificial Analysis, read on 16 September 2026 against Intelligence Index v4.3.
On score and speed, V4.1 Flash leads
Artificial Analysis scores V4.1 Flash at 39.55 with reasoning at max effort, against Luna at 37.50. That is a lead of about two points.
The speed gap is bigger. V4.1 Flash generates 214.4 output tokens per second to Luna’s 115, close to twice as fast, and at this tier speed is often the reason to pick a model in the first place.
V4.1 Flash also allows far longer answers, with a maximum output of 384K tokens against Luna’s 128K, and it is open weights, so you can download it and run it on your own hardware. I could not find published weights for Luna.
On price, it depends on the clock
DeepSeek’s pricing page lists two rates for V4.1 Flash:
- Off-peak: $0.15 input, $0.60 output, $0.003 for a cache hit.
- Peak: $0.30 input, $1.20 output, $0.006 for a cache hit.
Luna costs $0.20 / $1.20 with $0.02 cached, at any hour.
So off-peak, V4.1 Flash is cheaper than Luna on everything: 25% less on input and half on output. At peak it costs 50% more than Luna on input and the same on output, though cached input stays far cheaper on DeepSeek at either hour.
The same model is the cheaper choice at 13:00 in Warsaw and the more expensive one on input at 10:00. For a steady workload that runs all day, you pay a blend of the two.
When peak actually falls
DeepSeek lists peak as 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday through Friday. All other hours, including weekends, are off-peak.
Converted for a European team:
| Peak window | UTC | Warsaw / Berlin, until late October | Warsaw / Berlin, after |
|---|---|---|---|
| First | 01:00 - 04:00 | 03:00 - 06:00 | 02:00 - 05:00 |
| Second | 06:00 - 10:00 | 08:00 - 12:00 | 07:00 - 11:00 |
DeepSeek's peak hours from its pricing page, converted for Central European Summer Time (UTC+2) and Central European Time (UTC+1).
The second window is the one that matters. In September it covers 08:00 to 12:00 in Central Europe, so an assistant that serves a European team during office hours pays peak rates for the whole working morning.
A US team sees the reverse. Both windows fall in the late evening and early morning on the East Coast, outside most working hours.
Scheduling is the whole strategy
The pricing rewards moving work you control:
- Batch jobs such as nightly classification or extraction can run off-peak and pay $0.15 / $0.60.
- Heavy cached context is cheap at any hour, at $0.003 or $0.006 per million.
- Interactive work during the European morning pays peak, where Luna’s flat $0.20 input is cheaper.
If you cannot schedule the work, the comparison narrows to score and speed against a higher input price during the hours you probably use most.
Luna has a pricing trap of its own. Above 272K input tokens OpenAI bills the whole request at 2x input and 1.5x output, which I went through in the 272K pricing cliff.
Which one I would use
| The work | Model | Why |
|---|---|---|
| Batch jobs you can schedule | DeepSeek V4.1 Flash | Off-peak it is cheaper than Luna on every rate, and it scores higher. |
| European daytime interactive use | GPT-5.6 Luna | Peak pricing covers the working morning; Luna's input is cheaper then. |
| Very long outputs | DeepSeek V4.1 Flash | 384K max output against Luna's 128K. |
| Need to self-host | DeepSeek V4.1 Flash | Open weights. I found none for Luna. |
| Predictable monthly spend | GPT-5.6 Luna | One price at every hour, nothing to schedule around. |
Score and speed favour V4.1 Flash. Price favours whichever model matches when your work runs.
For other challengers to Luna at this tier, see Gemini 3.8 Flash against Luna and Haiku 4.5 against Luna.
What would change my mind
DeepSeek moving the peak windows. They are a pricing policy, and the European morning overlap is the biggest single factor here.
A Luna price cut. Its input is already within five cents of DeepSeek’s off-peak rate, so output, at twice DeepSeek’s off-peak price, is where a cut would change this comparison.
Cost per task on a real workload. DeepSeek V4.1 Flash cost Artificial Analysis $476.89 to run its index against $319.93 for Luna, but that run was priced at DeepSeek’s peak input rate, 50% above Luna’s. I have not seen a like-for-like run at off-peak prices.
Questions people asked
Is DeepSeek V4.1 Flash better than GPT-5.6 Luna?
Modestly, and faster: 39.55 against 37.50 on the Intelligence Index v4.3, and 214.4 against 115 output tokens per second. It is also open weights with a 384K max output.
Which is cheaper, DeepSeek V4.1 Flash or GPT-5.6 Luna?
It depends on the time. Off-peak, DeepSeek at $0.15 / $0.60 is cheaper on every rate, while at peak its $0.30 / $1.20 is more than Luna on input and the same on output. Its cached input costs less at any hour.
When are DeepSeek’s peak hours?
01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday. In Central European Summer Time, 03:00 to 06:00 and 08:00 to 12:00; after late October, 02:00 to 05:00 and 07:00 to 11:00.
Should I use DeepSeek V4.1 Flash or GPT-5.6 Luna?
V4.1 Flash for batch work you can schedule off-peak, and for very long outputs or self-hosting. Luna for interactive work during the European morning and for predictable spend.
If you are deciding where to run a high-volume pipeline and when, the contact page is the fastest route to me.