- Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are now generally available (GA) for production use.
- Gemini 3.6 Flash offers stronger performance on complex agentic and multimodal tasks at a lower price point ($1.50/1M input, $7.50/1M output tokens).
- Gemini 3.5 Flash-Lite is the fastest and lowest-cost model in the 3.5 family ($0.30/1M input, $2.50/1M output tokens).
- Both models support a 1M token context window, 64k max output tokens, thinking, and Computer Use.
- Key improvements in 3.6 Flash include token reduction, better code generation, instruction following, and multimodal reasoning.
- Key improvements in 3.5 Flash-Lite include reduced latency, enhanced reasoning, subagent orchestration, and document understanding.
- API changes: sampling parameters (temperature, top_p, top_k) are deprecated; prefilled model turns are no longer allowed.