Specific Labs dropped Real-SWE, an enterprise-code SWE benchmark, and the leaderboard is a great study in why you should never read "Claude Code" or "Codex CLI" as a model name. Same harness, two different brains: GPT-6 Astra on Codex CLI: 33.8% resolution GPT-5.6 Sol on Codex CLI: 16.2% resolution Same vendor's CLI, same harness, same benchmark. More than a 2x gap. If you'd just read "Codex CLI scored 16%," you'd write off the tool. If you read "Codex CLI scored 34%," you'd maybe believe it. Neither reading is right, because the harness is just the routing layer, not the thing that does the thinking. It's the same on the Claude side. Fable 5.1 on Claude Code is 38.8%. GLM 5.3 on Claude Code is 28.8%. One harness name, ten points apart. Real-SWE is careful about this, actually: they frame every result as a model-and-harness combination, not a model in isolation. That's the honest way to present it, and most vendors won't do it because a low number under their own tooling looks bad. This matters for teams actually shopping for coding agents, because "we use Claude Code" says nothing about the skill of the agent you get. You've picked a route, not a brain. The model swap is the biggest lever, and it's completely invisible in the marketing. The eval habit I'd take from this: when someone hands you a benchmark score, pin both halves. Model. Harness. Conventions, context carry-over, tool loop, judge. If a tool vendor won't tell you which model a number belongs to, that's a red flag, not a detail. The scaffold can move a score more than reasoning effort does.