Anthropic · OpenAI — Interactive pricing comparison
Sonnet 4 wins on raw coding (58 vs 55) and adds computer-use, but GPT-4.1 costs a third less on output, so the choice comes down to whether you need agentic browser control or just cheaper tokens.
Anthropic
OpenAI
GPT-4.1 is 43% cheaper at this usage
GPT-4.1 has a larger context window
Computer-use is the dividing line here. Claude Sonnet 4 can drive a browser, click through interfaces, and complete multi-step UI tasks. GPT-4.1 cannot. If your workload involves any kind of agentic automation against real software, the comparison ends before it starts.
For everyone else, the numbers get closer. Sonnet 4 leads coding 58 to 55 and reasoning 64 to 62. Those gaps are real but narrow. You will notice them on hard refactors and tangled debugging sessions, less so on routine generation. Both ship a 1M token context window, so neither has an advantage on large codebases or long documents.
Price is where GPT-4.1 fights back. It runs $2/$8 against Sonnet's $3/$15. Output tokens cost nearly half. On a high-volume generation pipeline, that difference compounds fast. A team pushing millions of tokens daily will feel $15 output far more than a two-point coding score.
There is also the matter of reasoning depth. Sonnet 4 supports an extended thinking mode that spends extra tokens before answering, which lifts quality on the hardest problems but also inflates the bill in ways that are easy to miss. GPT-4.1 has no such mode, so its costs stay predictable.
Pick Sonnet 4 when you need computer-use, the strongest coding scores, or that extended reasoning. Pick GPT-4.1 when throughput and a tight budget matter more than a few benchmark points. Both are capable enough that most teams will be productive on either.
Last updated June 2026