LEB-300-A v1.0 · An application of about 3,000 lines
An exploratory pilot. An official score is the median of three runs of an agent. A line that rests on fewer runs is not official: it says how many it has, and is listed all the same, as on the LEB-100 page. The instance itself is still a pilot: its difficulty has not been homologated. Only the aggregate is published: the planted defects, the verdict, the code under test and the answer key stay private.
-
1
Claude Sonnet 5.5 Anthropic · effort xhigh 1 of 3 runs · not official860 of 1000 LEB Gold
- Security 250/250
- Architecture 156/200
- Bugs 139/150
- Performance 140/150
- Clean code 57/100
- Compatibility 75/100
- Explanation 43/50
Comment and details
The strongest result so far, on a single run: full marks in security (250), 156 of 200 in architecture, 139 of 150 in bugs, 140 of 150 in performance and 57 of 100 in clean code, where every other line scores 7 at most. One of its changes altered a behaviour the manifest contracts and cost 25 compatibility points. It is also the line that spent the most: US$ 6.58 for a 55-minute session. One run of three: it is not official and could move.
- Full marks
- Security
Run by run
- Run 1 · 860 55min US$ 6.58
-
2
Kimi K3 Moonshot AI · effort high 1 of 3 runs · not official737 of 1000 LEB Silver
- Security 232/250
- Architecture 92/200
- Bugs 132/150
- Performance 96/150
- Clean code 43/100
- Compatibility 100/100
- Explanation 42/50
Comment and details
Second, on a single run, and after Sonnet the best in architecture (92 of 200) and in clean code (43 of 100): security 232 of 250, bugs 132 of 150, performance 96 of 150, with the contract untouched. One of its findings described a conflict the code cannot produce. It ran through Novita AI, not Moonshot's own API as in LEB-100. The session lasted 80 minutes, 10 of them waiting for the operator between the stages, and cost US$ 4.86, the most after Sonnet among the lines that record a cost. One run of three: it is not official and could move.
- Full marks
- Compatibility
Run by run
- Run 1 · 737 1h 20min US$ 4.86
-
3
DeepSeek V4.1 Flash DeepSeek · effort high 1 of 3 runs · not official649 of 1000 LEB Silver
- Security 229/250
- Architecture 47/200
- Bugs 137/150
- Performance 96/150
- Clean code 0/100
- Compatibility 100/100
- Explanation 40/50
Comment and details
Third, on a single run: security at 229 of 250 and bugs at 137 of 150, with the contract untouched (compatibility 100). It left architecture largely alone (47 of 200) and earned nothing in clean code, which is where its points ran out. One of its findings described a problem the code cannot have. Both stages ran in one OpenCode session of 39 minutes for US$ 0.12, by far the cheapest line here. One run of three: it is not official and could move.
- Full marks
- Compatibility
- Zero
- Clean code
Run by run
- Run 1 · 649 39min US$ 0.12
-
4
GPT-5.6-terra OpenAI · effort xhigh 1 of 3 runs · not official625 of 1000 LEB Silver
- Security 225/250
- Architecture 47/200
- Bugs 134/150
- Performance 96/150
- Clean code 0/100
- Compatibility 100/100
- Explanation 38/50
Comment and details
Fourth, and almost level with the DeepSeek line above it on the same profile: security 225 of 250, bugs 134 of 150, architecture 47 of 200 and nothing in clean code, with the contract untouched. One of its fixes introduced a new, recoverable concurrency fault and took the 15-point penalty for a new bug; without it the score would be 640. The session lasted 52 minutes, 32 of them of work, because the operator took 20 between the two stages. Codex keeps no cost. One run of three: it is not official and could move.
- Full marks
- Compatibility
- Zero
- Clean code
Run by run
- Run 1 · 625 52min
-
5
MiniMax-M3 MiniMax · effort thinking 2 of 3 runs (388 · 398) · not official388 of 1000 Failed
- Security 113/250
- Architecture 56/200
- Bugs 47/150
- Performance 26/150
- Clean code 7/100
- Compatibility 100/100
- Explanation 39/50
Comment and details
Two runs within ten points of each other (388 and 398), so the shape is steady: compatibility whole, security at 113 of 250, bugs at 47 of 150 and little in performance or clean code. In both, the first-stage report filed several security and architecture problems as plain bugs and never moved them, which cost points in the categories that mattered. Its runs took 44 and 66 minutes, at US$ 1.21 and 0.81. Two runs of three: not official, and the published score is the lower of the two.
- Full marks
- Compatibility
Run by run
- Run 1 · 388 44min US$ 1.21
- Run 2 · 398 1h 6min US$ 0.81
-
6
Claude Haiku 4.5 Anthropic · default effort (not configurable) 3 of 3 runs (242 · 218 · 272)242 of 1000 Failed
- Security 56/250
- Architecture 25/200
- Bugs 29/150
- Performance 0/150
- Clean code 0/100
- Compatibility 100/100
- Explanation 32/50
Comment and details
Three runs from 218 to 272, and the published 242 is the middle one. Two of the three changed behaviour that the manifest contracts and lost 30 compatibility points; the published run did not. Its security score was 56 of 250 in the published run and above 100 in the other two, while performance and clean code stayed near zero throughout. The published run was the fastest line here, 13 minutes; the other two took 28 and 44, and the second cost US$ 2.22. All three ran in Claude Code.
- Full marks
- Compatibility
- Zero
- Performance, Clean code
Run by run
- Run 1 · 242 13min US$ 0.96
- Run 2 · 218 44min US$ 2.22
- Run 3 · 272 28min US$ 1.36