AI & ML
Two "Codex CLI" models on the same benchmark: the harness hides the model
Cole Halton Dev.to (EN Zone)
3 views
Specific Labs dropped Real-SWE, an enterprise-code SWE benchmark, and the leaderboard is a great study in why you should never read "Claude Code" or "Codex CLI" as a model name.
Same harness, two different brains:
GPT-6 Astra on Codex CLI: 33.8% resolution
GPT-5.6 Sol on Codex CLI: 16.2% resolution
Same vendor's CLI, same harness, same benchmark. More than a 2x gap. If you'd just read "Codex CLI scored 16%," you'd write off the tool. If you read "Codex CLI scored 34%," you'd maybe believe it. Neither reading is right, because the harness is just the routing layer, not the thing that does the thinking.
It's the same on the Claude side. Fable 5.1 on Claude Code is 38.8%. GLM 5.3 on Claude Code is 28.8%. One harness name, ten points apart. Real-SWE is careful about this, actually: they frame every result as a model-and-harness combination, not a model in isolation. That's the honest way to present it, and most vendors won't do it because a low number under their own tooling looks bad.
This matters for teams actually shopping for coding agents, because "we use Claude Code" says nothing about the skill of the agent you get. You've picked a route, not a brain. The model swap is the biggest lever, and it's completely invisible in the marketing.
The eval habit I'd take from this: when someone hands you a benchmark score, pin both halves. Model. Harness. Conventions, context carry-over, tool loop, judge. If a tool vendor won't tell you which model a number belongs to, that's a red flag, not a detail. The scaffold can move a score more than reasoning effort does.
Read original: https://dev.to/cole_halton_42f71d71b809b/two-codex-cli-models-on-the-same-benchmark-the-harness-hides-the-model-1jn
← Previous
Giving a coding agent more time barely helps
Next →
Cutting PR review time is really changing where review happens
Related
AX-RAY & K-MYTHOS: Inside Korea's Consortium-Built Security-Specialized AI Foundation Model
AI & ML
0
Dev.to (EN Zone)
Next.js & AI Systems Architecture: Scaling Real-Time Agents (2026)
AI & ML
0
Dev.to (EN Zone)
Just Train More: Measuring the Exchange Rate
AI & ML
0
Dev.to (EN Zone)
AI Won't Fix a Broken Process. Map It First, Then Automate
AI & ML
0
Dev.to (EN Zone)
Comments0
No comments yet — be the first