---
title: Gemini 3.6 Flash Cuts Output Tokens 17%, Price Too
description: Gemini 3.6 Flash uses 17% fewer output tokens than 3.5 Flash and costs $0.75/$3.75 per million tokens through 2026, rising to $1.50/$7.50 in 2027.
date: 2026-08-12T00:00:00.000Z
category: ai-productivity
tags: tokens, gemini, google, llm-pricing
---

## Quick Answer

Gemini 3.6 Flash reportedly uses 17% fewer output tokens than Gemini 3.5 Flash for comparable tasks, and launched at $0.75 per million input tokens and $3.75 per million output tokens, with a 1M-token context window. That current rate holds through 2026-12-31 — Google's own pricing page confirms it rises to $1.50 and $7.50 on 2027-01-01. That's two cost levers moving at once: a lower per-token rate, and fewer tokens spent per task.

<Callout type="warning" title="Correction">
  An earlier version of this post stated Gemini 3.6 Flash's price as $1.50 input
  / $7.50 output per million tokens — that is the rate scheduled to take effect
  2027-01-01, not the price actually in effect today. The figures throughout
  this post now reflect the current 2026 rate, with the 2027 change stated
  separately and dated.
</Callout>

## What's actually changing

Tech-press launch coverage (9to5google.com, "Google launches Gemini 3.6 Flash and 3.5 Flash-Lite," dated 2026-07-21) reports that Gemini 3.6 Flash uses 17% fewer output tokens than Gemini 3.5 Flash for comparable tasks. That figure is sourced to reputable tech-press coverage, not Google's own blog post directly, and is stated honestly as such.

Separately, and confirmed directly against [Google's own Gemini API pricing page](https://ai.google.dev/gemini-api/docs/pricing), Gemini 3.6 Flash launched at $0.75 per million input tokens and $3.75 per million output tokens, with a 1M-token context window. That page also states the rate rises to $1.50 input / $7.50 output on 2027-01-01 — a scheduled increase, not a current price, and worth budgeting around separately if a project's usage extends past that date.

## Structural Comparison Matrix

| Operational Aspect                     | Gemini 3.5 Flash                              | Gemini 3.6 Flash (now, through 2026-12-31) | Gemini 3.6 Flash (from 2027-01-01) |
| :------------------------------------- | :-------------------------------------------- | :----------------------------------------- | :--------------------------------- |
| **Output tokens for comparable tasks** | Baseline                                      | ~17% fewer (reported)                      | ~17% fewer (reported)              |
| **Input price (per million tokens)**   | Not part of the sourced figures for this post | $0.75                                      | $1.50                              |
| **Output price (per million tokens)**  | Not part of the sourced figures for this post | $3.75                                      | $7.50                              |
| **Context window**                     | Not part of the sourced figures for this post | 1M tokens                                  | 1M tokens                          |

## Why two levers matter more than one

A comparison that only checks the per-token rate misses half the picture here. If Gemini 3.6 Flash really does produce fewer output tokens for the same task, the cost saving compounds: a lower rate applied to a smaller token count, not just a lower rate applied to the same count as before. Two providers can advertise similar-looking per-token prices and still produce meaningfully different bills for the same workload, once actual token counts per task are accounted for.

That is exactly the trap a sticker-price-only comparison falls into, and exactly why [Same Prompt, Different Bill](/ai-productivity/same-prompt-different-bill-gpt-claude-gemini/) treats token count and price per token as two separate numbers to check, not one.

<Callout type="tip" title="Measure your own task, not the reported average">
  A 17% average reduction across "comparable tasks" may not match your specific
  prompt shape. Run your actual prompts through the free{" "}
  <a href="/tools/ai-token-counter/">AI Token Counter</a> to see your own
  estimated Gemini token counts before assuming the reported figure applies
  directly to your workload, then use the{" "}
  <a href="/tools/llm-pricing-calculator/">LLM Pricing Calculator</a> to turn
  that count into an actual dollar figure against the current 2026 rate above.
</Callout>

## Confirmed version

Pricing (both the current 2026 rate and the scheduled 2027-01-01 change) is confirmed directly against [Google's own Gemini API pricing page](https://ai.google.dev/gemini-api/docs/pricing), re-checked 2026-08-13. The 17% output-token reduction claim is sourced to 9to5google.com's launch coverage, dated 2026-07-21 — a reputable tech-press source, but not Google's own blog post directly, and labeled as such throughout. Browse more coverage in the [AI Productivity](/ai-productivity) archive, or start from [The 2026 LLM Token & Pricing Reset](/ai-productivity/2026-llm-token-pricing-reset/) hub.
