What’s actually changing
Tech-press launch coverage (9to5google.com, “Google launches Gemini 3.6 Flash and 3.5 Flash-Lite,” dated 2026-07-21) reports that Gemini 3.6 Flash uses 17% fewer output tokens than Gemini 3.5 Flash for comparable tasks. That figure is sourced to reputable tech-press coverage, not Google’s own blog post directly, and is stated honestly as such.
Separately, and confirmed directly against Google’s own Gemini API pricing page, Gemini 3.6 Flash launched at $0.75 per million input tokens and $3.75 per million output tokens, with a 1M-token context window. That page also states the rate rises to $1.50 input / $7.50 output on 2027-01-01 — a scheduled increase, not a current price, and worth budgeting around separately if a project’s usage extends past that date.
Structural Comparison Matrix
| Operational Aspect | Gemini 3.5 Flash | Gemini 3.6 Flash (now, through 2026-12-31) | Gemini 3.6 Flash (from 2027-01-01) |
|---|---|---|---|
| Output tokens for comparable tasks | Baseline | ~17% fewer (reported) | ~17% fewer (reported) |
| Input price (per million tokens) | Not part of the sourced figures for this post | $0.75 | $1.50 |
| Output price (per million tokens) | Not part of the sourced figures for this post | $3.75 | $7.50 |
| Context window | Not part of the sourced figures for this post | 1M tokens | 1M tokens |
Why two levers matter more than one
A comparison that only checks the per-token rate misses half the picture here. If Gemini 3.6 Flash really does produce fewer output tokens for the same task, the cost saving compounds: a lower rate applied to a smaller token count, not just a lower rate applied to the same count as before. Two providers can advertise similar-looking per-token prices and still produce meaningfully different bills for the same workload, once actual token counts per task are accounted for.
That is exactly the trap a sticker-price-only comparison falls into, and exactly why Same Prompt, Different Bill treats token count and price per token as two separate numbers to check, not one.
Measure your own task, not the reported average
A 17% average reduction across “comparable tasks” may not match your specific prompt shape. Run your actual prompts through the free AI Token Counter to see your own estimated Gemini token counts before assuming the reported figure applies directly to your workload, then use the LLM Pricing Calculator to turn that count into an actual dollar figure against the current 2026 rate above.
Confirmed version
Pricing (both the current 2026 rate and the scheduled 2027-01-01 change) is confirmed directly against Google’s own Gemini API pricing page, re-checked 2026-08-13. The 17% output-token reduction claim is sourced to 9to5google.com’s launch coverage, dated 2026-07-21 — a reputable tech-press source, but not Google’s own blog post directly, and labeled as such throughout. Browse more coverage in the AI Productivity archive, or start from The 2026 LLM Token & Pricing Reset hub.







