The speed of AI development is breathtaking. New models are being released almost daily. Google has recently significantly expanded its Gemini series. The latest additions include the “3.6 Flash” and “3.5 Flash-Lite,” joining the existing high-intelligence “3.1 Pro.” With the addition of “Extended” versions for each—which include extra “thinking time”—there are now six distinct options. But is it always the right move to use the fastest or the most “thoughtful” model?
The 1,000,000 Token Limit and Extended Costs
All six Gemini models share the ability to process 1,000,000 tokens at once. This capacity, equivalent to roughly 750,000 words, allows the AI to understand dozens of thick books or an entire hour-long video in one go.
Then there is the “Extended” concept. Before providing an answer, the model performs internal logical reasoning. For a human, this is akin to being asked a question and saying, “Hmm…” while organizing thoughts deeply.
This significantly increases accuracy for math problems or puzzles. However, keep in mind that the thinking process consumes “output tokens.” This means that the longer the model “thinks,” the more it costs to run.
When Every Millisecond Counts: 3.5 Flash-Lite
As the name suggests, “3.5 Flash-Lite” is an extremely lightweight and fast model. Inputting 1 million tokens costs only $0.30, and the output cost is roughly $2.50, making it the most affordable in the lineup.
It is specialized for simple, repetitive tasks rather than complex reasoning. It is ideal for tasks like rapidly classifying millions of customer inquiries or extracting specific keywords from massive amounts of text. It is most commonly used in services where real-time response is critical.
There is also an Extended version of this model, used when slight fact-checking is needed during light tasks. However, it is used sparingly, as it diminishes the “ultra-high-speed” advantage that defines the Lite model.
I personally use this primarily for document translation within my n8n automation workflows.
The Reliable Workhorse: 3.6 Flash
“3.6 Flash” is Google’s latest flagship model, offering the best balance of speed, performance, and cost. Input costs are $1.50 per million tokens, and output is roughly $7.50, making it very reasonable.
Its language comprehension is significantly smarter than previous generations, and its output is more stable. It handles everything from daily task assistance to coding and image analysis well. It is the most suitable engine for core workflow automation programs.
The 3.6 Flash Extended version adds a layer of logic to this. It is deployed when coding requirements are highly specific or when complex mathematical calculations are needed. It offers high accuracy at a reasonable price, making it highly effective for professional use.
The Expert Problem Solver: 3.1 Pro
“3.1 Pro” is built for high-difficulty tasks. It is the most expensive, at $2.00 input and $12.00 output per million tokens (for processing under 200,000 tokens).
While the cost is higher, it is meticulously accurate. It is optimized for complex knowledge work, such as reading and analyzing hundreds of documents at once or synthesizing various pieces of information to reach a new conclusion.
The pinnacle of the lineup is “3.1 Pro Extended.” It is like giving a brilliant professor ample time to think. It is used for designing entire system architectures or cross-verifying academic papers. While slower, it represents the highest intelligence currently available in the Gemini ecosystem.
Gemini 6-Model Comparison Table
I have summarized the characteristics of the six Gemini models in the table below. The key is the combination of three “weight classes” and the “thinking time” (Extended).
| Model Name | Key Features | Input Cost (1M Tokens) | Output Cost (1M Tokens) |
|---|---|---|---|
| 3.5 Flash-Lite | Ultra-fast, simple repetition | ~$0.30 | ~$2.50 |
| (Extended) | Light tasks + basic logic | Same (+ thinking cost) | Same (+ thinking cost) |
| 3.6 Flash | Great balance, pro standard | ~$1.50 | ~$7.50 |
| (Extended) | Pro standard + precise calculation | Same (+ thinking cost) | Same (+ thinking cost) |
| 3.1 Pro | High-difficulty complex reasoning | ~$2.00 | ~$12.00 |
| (Extended) | Top intelligence, academic/design | Same (+ thinking cost) | Same (+ thinking cost) |
As shown in the table, cost and performance are proportional. The price difference between the cheapest and most expensive models is more than sixfold. This is why you must objectively evaluate the difficulty of your tasks.
Choosing the Perfect AI for You
Each of the six Gemini models has a distinct role based on processing speed and cost per token. Using the expensive 3.1 Pro for simple text summarization is a massive waste.
Why?
When using Gemini Pro, your usage limits deplete faster as you move toward higher tiers and utilize extended reasoning.
Of course, if you don’t hit those usage limits and can afford to wait a few seconds for an answer, the 3.1 Pro version is the best choice.
Ultimately, the best AI is the one that fits your purpose and budget perfectly. Before you start, ask yourself: do you need the $0.30 ultra-high-speed processing, or do you need the $2.00 sophisticated insights? Answering that question is the first step toward using AI wisely.
📸 Behind the scenes
Check out more behind-the-scenes shots on my Naver blog (Korean):
신규 모델 제미나이 3.6 플래시부터 3.1 프로까지, 목적별 최적의 AI 추천
👉 https://blog.naver.com/PostView.naver?blogId=k5kun&logNo=224358339091