Skip to content

The first level, LEB-100, is the reference instance: every agent evaluated so far has run on it, and every run is published.

AI Benchmark · LEB

LEB-300 exploratory pilot

The second level of LEB: an application of about 3,000 lines, in multiple files, whose flaws cross files. The instance is active, so its answer key stays private and only the aggregate of each agent is published. How a run works, how it is scored and how this level differs from LEB-100 is on the Benchmark page.

The results so far

One line per agent: the total, the grade, the score in each category, how many runs it rests on, and the cost and the time. Every agent solved the same package, byte for byte, so the numbers compare like for like.

LEB-300-A v1.0 · An application of about 3,000 lines

An exploratory pilot. An official score is the median of three runs of an agent. A line that rests on fewer runs is not official: it says how many it has, and is listed all the same, as on the LEB-100 page. The instance itself is still a pilot: its difficulty has not been homologated. Only the aggregate is published: the planted defects, the verdict, the code under test and the answer key stay private.

mode A · 60 turns edition 2026 key private matrix c42c82878d2c…
  1. 1
    Claude Sonnet 5.5 Anthropic · effort xhigh 1 of 3 runs · not official
    860 of 1000 LEB Gold
    • Security 250/250
    • Architecture 156/200
    • Bugs 139/150
    • Performance 140/150
    • Clean code 57/100
    • Compatibility 75/100
    • Explanation 43/50

    Cost US$ 6.58 a run Session 54.8 min

    Comment and details

    The strongest result so far, on a single run: full marks in security (250), 156 of 200 in architecture, 139 of 150 in bugs, 140 of 150 in performance and 57 of 100 in clean code, where every other line scores 7 at most. One of its changes altered a behaviour the manifest contracts and cost 25 compatibility points. It is also the line that spent the most: US$ 6.58 for a 55-minute session. One run of three: it is not official and could move.

    A written reading of the aggregate; it is not part of the score, and it names no flaw.

    Full marks
    Security
    Run by run
    • Run 1 · 860 55min US$ 6.58
  2. 2
    Kimi K3 Moonshot AI · effort high 1 of 3 runs · not official
    737 of 1000 LEB Silver
    • Security 232/250
    • Architecture 92/200
    • Bugs 132/150
    • Performance 96/150
    • Clean code 43/100
    • Compatibility 100/100
    • Explanation 42/50

    Cost US$ 4.86 a run Session 79.8 min

    Comment and details

    Second, on a single run, and after Sonnet the best in architecture (92 of 200) and in clean code (43 of 100): security 232 of 250, bugs 132 of 150, performance 96 of 150, with the contract untouched. One of its findings described a conflict the code cannot produce. It ran through Novita AI, not Moonshot's own API as in LEB-100. The session lasted 80 minutes, 10 of them waiting for the operator between the stages, and cost US$ 4.86, the most after Sonnet among the lines that record a cost. One run of three: it is not official and could move.

    A written reading of the aggregate; it is not part of the score, and it names no flaw.

    Full marks
    Compatibility
    Run by run
    • Run 1 · 737 1h 20min US$ 4.86
  3. 3
    DeepSeek V4.1 Flash DeepSeek · effort high 1 of 3 runs · not official
    649 of 1000 LEB Silver
    • Security 229/250
    • Architecture 47/200
    • Bugs 137/150
    • Performance 96/150
    • Clean code 0/100
    • Compatibility 100/100
    • Explanation 40/50

    Cost US$ 0.12 a run Session 38.6 min

    Comment and details

    Third, on a single run: security at 229 of 250 and bugs at 137 of 150, with the contract untouched (compatibility 100). It left architecture largely alone (47 of 200) and earned nothing in clean code, which is where its points ran out. One of its findings described a problem the code cannot have. Both stages ran in one OpenCode session of 39 minutes for US$ 0.12, by far the cheapest line here. One run of three: it is not official and could move.

    A written reading of the aggregate; it is not part of the score, and it names no flaw.

    Full marks
    Compatibility
    Zero
    Clean code
    Run by run
    • Run 1 · 649 39min US$ 0.12
  4. 4
    GPT-5.6-terra OpenAI · effort xhigh 1 of 3 runs · not official
    625 of 1000 LEB Silver
    • Security 225/250
    • Architecture 47/200
    • Bugs 134/150
    • Performance 96/150
    • Clean code 0/100
    • Compatibility 100/100
    • Explanation 38/50

    Session 52.2 min

    Comment and details

    Fourth, and almost level with the DeepSeek line above it on the same profile: security 225 of 250, bugs 134 of 150, architecture 47 of 200 and nothing in clean code, with the contract untouched. One of its fixes introduced a new, recoverable concurrency fault and took the 15-point penalty for a new bug; without it the score would be 640. The session lasted 52 minutes, 32 of them of work, because the operator took 20 between the two stages. Codex keeps no cost. One run of three: it is not official and could move.

    A written reading of the aggregate; it is not part of the score, and it names no flaw.

    Full marks
    Compatibility
    Zero
    Clean code
    Run by run
    • Run 1 · 625 52min
  5. 5
    MiniMax-M3 MiniMax · effort thinking 2 of 3 runs (388 · 398) · not official
    388 of 1000 Failed
    • Security 113/250
    • Architecture 56/200
    • Bugs 47/150
    • Performance 26/150
    • Clean code 7/100
    • Compatibility 100/100
    • Explanation 39/50

    Cost US$ 1.21 a run Session 43.7 min

    Comment and details

    Two runs within ten points of each other (388 and 398), so the shape is steady: compatibility whole, security at 113 of 250, bugs at 47 of 150 and little in performance or clean code. In both, the first-stage report filed several security and architecture problems as plain bugs and never moved them, which cost points in the categories that mattered. Its runs took 44 and 66 minutes, at US$ 1.21 and 0.81. Two runs of three: not official, and the published score is the lower of the two.

    A written reading of the aggregate; it is not part of the score, and it names no flaw.

    Full marks
    Compatibility
    Run by run
    • Run 1 · 388 44min US$ 1.21
    • Run 2 · 398 1h 6min US$ 0.81
  6. 6
    Claude Haiku 4.5 Anthropic · default effort (not configurable) 3 of 3 runs (242 · 218 · 272)
    242 of 1000 Failed
    • Security 56/250
    • Architecture 25/200
    • Bugs 29/150
    • Performance 0/150
    • Clean code 0/100
    • Compatibility 100/100
    • Explanation 32/50

    Cost US$ 0.96 a run Session 13.4 min

    Comment and details

    Three runs from 218 to 272, and the published 242 is the middle one. Two of the three changed behaviour that the manifest contracts and lost 30 compatibility points; the published run did not. Its security score was 56 of 250 in the published run and above 100 in the other two, while performance and clean code stayed near zero throughout. The published run was the fastest line here, 13 minutes; the other two took 28 and 44, and the second cost US$ 2.22. All three ran in Claude Code.

    A written reading of the aggregate; it is not part of the score, and it names no flaw.

    Full marks
    Compatibility
    Zero
    Performance, Clean code
    Run by run
    • Run 1 · 242 13min US$ 0.96
    • Run 2 · 218 44min US$ 2.22
    • Run 3 · 272 28min US$ 1.36

What the record does not have

  • No checkpoint was taken between the two stages of the task in any of the runs, so it is not on record that the model stayed the same through the first stage. The transcripts show a single model.
  • The task text the agent read names the previous scoring-matrix hash. The matrix this run is scored against differs from it only by a header field that marks the instance as active. The scoring is the same.
  • The agents did not all run in the same client, Claude Code or OpenCode, each with its own tools. The client of every run is in its record.

What is published

This page shows one line per agent: the total, the grade, the score in each category, how many runs it rests on, and the cost and the time. It never shows the list of planted defects, the code under test or the answer key. The same instance has to keep measuring new agents, so those stay private.

The instance is in an exploratory pilot. The results above are the ones published so far, and other agents follow. Its first runs also served to check the procedure.

The protocol, the scoring and the tools are public in the ai-benchmark repository. How a run works and how the two levels differ is on the Benchmark page, and the results of the first level are on the LEB-100 page.