Can an AI maintain legacy code without breaking it?
LEB — the LLM Engineering Benchmark — hands an AI agent a legacy system in production, with flaws planted in it and consumers that depend on how it behaves today. It measures the work that dominates real engineering: finding the flaws, fixing them, keeping every contract intact, and explaining the decisions like a senior engineer would.
Why another benchmark
Most benchmarks measure code written from scratch, or one isolated issue solved. Neither is what most engineering is: evolving a system that other people already depend on. LEB scores security, architecture, bugs, performance, clean code, compatibility and the quality of the explanation — and it takes points away from the agent that rewrites everything, swaps technologies without need, or breaks a public contract.
Rewriting from scratch is not engineering. It is running away.
How a run works
-
01
A legacy system, with planted flaws
The agent receives the code, a manifest of its public surface — the contract — and a neutral task: report the problems, fix what should be fixed, keep compatibility, justify every decision. It is never told which flaws exist, how many, or where.
-
02
The agent works alone
In mode A it gets tools and a budget of turns; in mode S, one prompt and one answer. It hands back the changed code, a technical report, and an index of its findings, each with a 0–100 confidence.
-
03
Machines check the code
Characterization tests run on the legacy code and on the delivery: public behaviour that changed is a regression. Probes then attack each fixable flaw — the injection payload, the empty dataset, the query counter — and report whether it is still there.
-
04
A judge checks the report
Each finding is matched against the Official Failure Matrix, a hidden answer key that also holds decoys: plausible flaws that do not exist, and cost points when reported. A second judge scores the explanation blind. A deterministic scorer turns it all into 0–1000.
LEB-100 and LEB-300
LEB is a ladder: each level is a larger legacy system than the one below it, and an instance is one concrete system at a level. An agent is ranked only against others on the same instance, never against a level in general. Two instances exist today.
| Compared on | LEB-100-A | LEB-300-A |
|---|---|---|
| Size | About 300 lines, in one or two files. | About 3,000 lines, in multiple files. |
| What it tests | The fundamentals: finding a flaw, fixing it, and keeping the contract intact. | Navigation: its flaws cross files, so the agent has to move through the code instead of reading it top to bottom. |
| Difficulty | Easy, moderate and hard flaws, none at expert level. | Expert-level flaws can appear from this level on. |
| The run | Mode A, 30 turns. | Mode A, 60 turns, in two stages. |
| The answer key | Public since 13 July 2026. A model may have seen it in training, and the leaderboard marks the ones that may have. | Private. The instance is active, so only the SHA-256 of its key is published. |
| What is published | Every delivery, mechanical report, verdict and scorecard, and the results as spreadsheets. | The aggregate only: the total, the grade, the score in each category, the runs, the cost and the time. |
| Where it stands | The reference instance: every agent evaluated so far has run on it. | An exploratory pilot: its difficulty has not been homologated. |
Both are scored out of 1000, in the same seven categories and with the same grades. The comparison that means something is within one instance, so a LEB-300 score is never set beside a LEB-100 one.
LEB-100-A is where the method was proven. Its answer key is public, so a model that trained on it may know the flaws in advance, and the leaderboard marks those. LEB-300-A is the level above, and it keeps its key private so that the same instance can keep measuring new agents.
The instances
Each one has a page of its own, with its results and what to read before quoting them.
-
LEB-100-A
A legacy PHP application of about 300 lines: the support-ticket panel of an internet provider. Every run is published, with its delivery, its verdict and its scorecard.
Open LEB-100-A -
LEB-300-A
The level above, an application of about 3,000 lines. The instance is active, so only the aggregate of each agent is published.
Open LEB-300-A
1000 points, and how they are lost
Every instance is worth exactly 1000: the raw points of each category are normalised to its weight, so scores compare across instances of a level.
Compatibility starts at 100 and only goes down. Migrating mysqli to PDO without need costs 20; changing a public signature costs 30, per function.
Global penalties come off the total: a new bug −15, each broken characterization test −20, a needless rewrite −25, each decoy reported −5.
- Security SEC 250
- Architecture ARCH 200
- Bugs BUG 150
- Performance PERF 150
- Clean code CLN 100
- Compatibility COMP 100
- Explanation EXPL 50
Grades
- LEB Platinum 900–1000 · ready for critical legacy
- LEB Gold 750–899 · solid engineering
- LEB Silver 600–749 · useful with supervision
- LEB Bronze 400–599 · needs a full review
- Failed < 400 · a risk to the system
Runs happen where the answer key is out of reach
Agents run on a dedicated Linux VM isolated from GitHub. Its names resolve to loopback, its address ranges (the prefixes it announces and the edge addresses it publishes) are blackhole routes, and the VM has no IPv6 connectivity. The address blocks came on 30 September 2026. The runs of 29 September had the name block only: no agent could reach GitHub by name, but a deliberate connection straight to one of its addresses was not blocked, and the session logs show none was attempted. Every published run of 30 September had every layer. What a model saw in training is a separate question, answered by the training cutoff each run records.
Where the method lives
The specification, the scoring and the tools are public in the ai-benchmark repository. The runs of the first level are there too, with their scorecards, and as two spreadsheets on the LEB-100 page.