Switching from GPT-5.5 to GPT-5.6 Made Me Less Productive
I pay for three Codex subscriptions at $200 each, and for the past week they have mostly bought me waiting. Since I…
Anthropic re-wrote the Sonnet 5 story post-launch. The first BrowseComp cost-performance chart showed Sonnet 5 lagging Opus 4.8. The replacement chart paints a much nicer picture! A wider curve, more room on the cost axis, and a model that looks like it could reach Opus territory if you spend enough money.
The official explanation now lives in Anthropic’s Sonnet 5 launch post. The changelog notes that the original BrowseComp chart did not apply Anthropic's standard agentic-search methodology and underestimated Sonnet 5 with a simpler methodology. The new chart is using the Sonnet 5 system-card setup, with 10M token budget, compaction, and programmatic tool calling. Anthropic says it also changed the text around the text.
Perhaps the new methodology is superior. Maybe the first chart did understate Sonnet 5 for real. Then the first chart shouldn’t have been included in the launch post.
This is important because Sonnet 5 was already in trouble on economics. It wasn’t just “here is a new model.” It was “here is the new default for Claude Code, here is the cost-performance story, and here is why developers should take it seriously.” The evidence in support of that claim is the graph.
Then the evidence was tampered with after people had already seen it.
The clean fix would have been easy: publish the old graph, the new graph, both data tables, the exact methodology diff and a time-stamped correction note. Shows the benchmark artifact for inspection. After the launch went badly the marketing story got fixed or Sonnet got a better harness – let the readers decide.
Instead, the public is left to compare screenshots.
That’s reputation problem. The new chart may be more accurate, Anthropic might think. What the users are seeing is something else. The weaker-looking chart is gone and the replacement makes Sonnet 5 look a lot more defensible.
The visual change is clear. The old screenshot visually caps out at around $10. Current chart goes to $50. Old story it was, "Sonnet trailing Opus". The new story is “Sonnet has a useful curve if you spend enough. If that’s the right answer, show your work.
Model labs already require a lot of trust from users. They pick the benchmark harness, tool setup, effort level, token budget, grader, and tokenizer assumptions, pricing assumptions, and chart design. These decisions can affect the result. Once a chart is launched, every hidden choice suddenly matters.
This is how goodwill disappears for developers. Not in one big scandal, but in tiny moments when the official page starts feeling less like evidence and more like a sales deck you can reshape after the fact.
Anthropic can say that the old graph underestimated Sonnet 5. The takeaway for the public is worse: the first graph made Sonnet 5 look weak, and the replacement makes Anthropic look unreliable.
Give Vroni a GitHub issue, bug report, spec, or rough idea. It reads the repo, plans the change, writes code, runs checks, and works toward a review-ready pull request.
Take a look at vroni.com