Google · OpenAI — Interactive pricing comparison
Gemini 2.5 Flash edges ahead on benchmarks (coding 45 vs 42, reasoning 50 vs 48) and adds a reasoning mode GPT-4.1 Mini lacks. But Mini charges $1.60 output against Flash's $2.50, so generation-heavy workloads pay less on the OpenAI side.
OpenAI
GPT-4.1 Mini is 23% cheaper at this usage
Gemini 2.5 Flash has a larger context window
Both ship a 1M token context window at sub-dollar input pricing. That alone makes them the two most cost-effective models for long-document work in production today. The question is which one you pick when both fit.
Flash wins on input cost: $0.30 versus Mini's $0.40. Mini wins on output: $1.60 versus Flash's $2.50. This is not a trivial difference. A summarization pipeline that reads 100K tokens and writes 2K tokens favors Flash. A code generation pipeline that reads a short prompt and writes 10K tokens favors Mini. Your token ratio determines your winner, not the spec sheet.
Benchmarks give Flash a slight lead across the board — coding 45 to 42, reasoning 50 to 48, math 58 to 55. The gaps are narrow enough that you will not feel them on routine tasks. They surface on harder problems where Flash also has the option of engaging its reasoning mode, a capability Mini does not offer. When Flash thinks before answering, its effective reasoning lifts above its base score. Mini has no such gear to shift into.
Both carry vision and function-calling. Both support prompt caching. Flash adds PDF input support. Neither supports computer-use.
Flash for input-heavy or reasoning-occasional workloads. Mini for output-heavy generation where the $0.90 per million output tokens saved compounds at scale. The model that costs less depends entirely on which end of the pipe you are heavier on.
Last updated June 2026