Loading benchmark data…
EyesInAI·Loading live benchmark data
We gave four coding agents the same 24 bug-fix and feature tasks and graded each result with hidden tests the agent never saw, run on a clean copy of its work. We also counted the result that matters most when you hand an agent real work: it said DONE, and the hidden tests failed.
| Agent | Passed | Easy | Real-world | Hard | Said DONE, wrong | Timed out | Median time |
|---|---|---|---|---|---|---|---|
| CodexGPT-6 Astra (high effort) | 22/24 | 6/6 | 9/10 | 7/8 | 2 | 0 | 1.8 min |
| Claude SonnetSonnet 5.5 in Claude Code (high effort) | 20/24 | 6/6 | 8/10 | 6/8 | 4 | 0 | 26 s |
| Z.ai GLMGLM-5.3 | 17/24 | 6/6 | 6/10 | 5/8 | 5 | 2 | 3.1 min |
| Google AntigravityGemini 3.1 Pro (high) | 13/24 | 6/6 | 5/10 | 2/8 | 11 | 0 | 2.6 min |
A pass means the hidden tests passed and the project's whole existing suite stayed green, inside the 25-minute limit. Median time is agent time only, excluding grading.
passed · failed but said DONE · hit the 25-minute limit
| Task | Codex | Claude Sonnet | Z.ai GLM | Google Antigravity |
|---|---|---|---|---|
| Easy pilot. Four public upstream bugs (more-itertools, humanize, boltons, click) and two synthetic tasks: a root-cause bug and a small feature with a CLI. | ||||
| more-itertools | ||||
| humanize | ||||
| boltons | ||||
| click | ||||
| synthetic root-cause bug | ||||
| synthetic feature + CLI | ||||
| Rebuilt real-world. New standalone code we wrote to reproduce the shape of bugs found in past code reviews. Published as numbered items only. | ||||
| Real-world 1 | ||||
| Real-world 2 | ||||
| Real-world 3 | ||||
| Real-world 4 | ||||
| Real-world 5 | ||||
| Real-world 6 | ||||
| Real-world 7 | ||||
| Real-world 8 | ||||
| Real-world 9 | ||||
| Real-world 10 | ||||
| Hard public bugs. Real bugs from public open-source projects: networkx, pyparsing, sqlglot, lark and flask. | ||||
| networkx: connectivity | ||||
| networkx: ISMAGS | ||||
| networkx: weak views | ||||
| sqlglot: BD literal | ||||
| pyparsing: match previous | ||||
| pyparsing: bounded repeat | ||||
| lark: interactive copy | ||||
| flask: IPv6 | ||||
What the failing hidden tests checked. Each of these was reported as finished.
| Z.ai GLM | networkx: weak views | A view pickled on its own comes back with a dead link to its graph (4 hidden tests). |
| Claude Sonnet | networkx: weak views | The same 4 tests as GLM: a view pickled on its own loses its graph. |
| Google Antigravity | networkx: weak views | Views were no longer cached, and pickled graphs lost their view cache (6 tests). |
| Claude Sonnet | sqlglot: BD literal | The integer form 0BD parses as an alias, and Spark 2's 1E3BD loses its suffix. |
| Codex | sqlglot: BD literal | The same 0BD case; Spark 2's negative scale crashes; Hive emits CAST where TRY_CAST is expected. |
| Google Antigravity | sqlglot: BD literal | Crashes on every BD literal: 123 hidden failures. |
| Google Antigravity | networkx: connectivity | Directed results still wrong for some flow functions, and non-connected graphs no longer raise (5 tests). The issue said “for every flow function”. |
| Google Antigravity | pyparsing: match previous | The hidden tests passed, but it broke an existing test, so the suite was red. |
| Google Antigravity | pyparsing: bounded repeat | Bounded repeats truncate results names to their first match (72 tests). |
| Google Antigravity | flask: IPv6 | SERVER_NAME="localhost:0" starts on port 5000, not 0 (1 test). |