OpenAI · Anthropic — Interactive pricing comparison
Sonnet 4 leads GPT-4o on coding by 8 points (58 vs 50), reasoning by 6 (64 vs 58), and carries a 1M context window against 4o's 128K. GPT-4o costs less on output ($10 vs $15 per 1M), but the capability gap is wide.
OpenAI
Anthropic
GPT-4o is 29% cheaper at this usage
Claude Sonnet 4 has a larger context window
This comparison has a shelf-life problem. GPT-4o launched in mid-2024. Claude Sonnet 4 launched in mid-2025. A full generation separates them, and it shows. Sonnet 4 leads coding 58 to 50, reasoning 64 to 58. Eight points on coding is not a rounding error — it is the difference between a model that handles complex refactors and one that struggles with them.
Context is where the gap becomes structural. Sonnet 4 processes 1M tokens. GPT-4o caps at 128K. For any work involving large codebases, long documents, or extended multi-turn conversations, 4o simply cannot compete. You hit its ceiling before the interesting work begins.
Sonnet 4 also carries computer-use and extended thinking — capabilities 4o does not have and will not receive. Computer-use enables agentic workflows against real interfaces. Extended thinking lets Sonnet spend additional reasoning tokens on hard problems, trading cost for quality in a controlled way.
Price is 4o's last foothold: $2.50/$10 versus Sonnet's $3/$15. The input gap is negligible. The output gap is real ($10 vs $15) but buys you 8 fewer coding points and 900K less context. For high-volume, low-complexity generation where neither model is stretched, that $5 per million output tokens saved is defensible. For everything else, you are paying less to get less.
Sonnet 4 unless your pipeline is locked into OpenAI infrastructure and the migration cost exceeds the capability gain. GPT-4o is now the model you keep running because switching has a cost, not because it wins on merit.
Last updated June 2026