The previous post made a claim I could not prove at the time: that AI hands you exactly one of the nine qualities — it works — and quietly leaves the other eight to you.
That is an easy thing to assert and an uncomfortable thing to test, because testing it means building something real, judging it honestly, and publishing the result whichever way it goes. So I did. This post is the receipt.
The experiment
I wrote a plan for a small but complete application — a book-review app: authors, books, users, reviews, a REST API over SQLite, a React front end with lists, detail pages and forms. Four related tables, nested data in every direction, real write operations. Not a toy, not an enterprise system. The size of thing a competent developer builds in a couple of days.
Then I handed the plan to an AI (Claude, driven through Claude Code) and said nothing else.
Nothing about TypeScript. Nothing about architecture, folder structure, testing, error handling, validation, or formatting. No style guide, no conventions, no definition of done. Just the plan and, in effect, build this.
It built it. Then I reviewed what came back against the nine qualities of quality code — the same nine, unchanged, from part 1.
The plan said what, never how
This distinction is the whole experiment, so it is worth showing rather than claiming. Here is the plan's table of contents, in full:
Every one of those is a what. The data model specifies columns, types and foreign keys. The backend section lists every endpoint down to the URL. The frontend section names the pages and says which nested data each must show. It is a genuinely thorough specification — the kind you would be pleased to receive as a contractor.
And it contains not one sentence about how any of it should be built.
PLAN.md — the complete plan, verbatim
What came back works
Let me start where the AI genuinely earns its reputation, because the rest of this post is going to be unkind and the credit is real.
It works. The server starts, creates its tables, seeds them, and serves. The React app builds
cleanly with npm run build and produces a production bundle. Every endpoint in the plan exists.
Every page in the plan exists. You can add an author, click through to their books, open a book,
read its reviews, follow a review to the user who wrote it, and delete things — and the lists
refresh correctly afterwards, because the cache invalidation is right.
It is also pleasant to read. Small functions, consistent shape, honest names. The four route modules look so alike that once you have read one you can read them all. Nobody wrote a thousand-line file.
Two of the nine qualities, present and earned, with no instruction given. That is not nothing.
The scorecard
Here is the same application against all nine.
| # | Quality | Verdict |
|---|---|---|
| 1 | Beautiful & readable | ✅ Present |
| 2 | Follows standards | ❌ Missing |
| 3 | Modular | ⚠️ Partial |
| 4 | Bug-free | ⚠️ Partial |
| 5 | Finished | ❌ Missing |
| 6 | Clean | ⚠️ Partial |
| 7 | Documented | ⚠️ Partial |
| 8 | It works | ✅ Present |
| 9 | Tested | ❌ Missing |
Two present, four half-built, three absent. Count a partial as half and it comes to four out of nine.
An application that runs perfectly and scores four out of nine is exactly the situation part 2 warned about, and it is worse than a broken one — because a broken application announces itself, and this one does not. So let us look at what those marks actually mean, in code you can judge for yourself.
Exhibit one: a review rated "oops"
The plan says a rating is an integer from 1 to 5. Here is the check the AI wrote:
Read it quickly and it looks fine. It asks whether the rating is too small, and whether it is too big. What it never asks is whether the rating is a number at all.
So send it a word:
The API answers 201 Created, and hands the record straight back to you:
In JavaScript, comparing the text "oops" against a number is false in both directions. "oops" is
not less than 1. "oops" is not greater than 5. The word walks straight through a guard that looks
like it is doing its job.
SQLite then accepts it too. The column is declared INTEGER NOT NULL, but SQLite only converts text
that actually is a number and stores the rest as text. So the word oops is now sitting in the
ratings column of the database.
Open the page and the review is there, with no stars, labelled (oops/5).
Nothing crashed. Nothing was logged. No test failed, because there are no tests. The application is in exactly the state it reports itself to be in: fine.
This is the shape of the whole problem. Not a crash — a crash would be a kindness. A quiet, permanent wrong answer that the software is entirely satisfied with.
The fix is one line, and it is the line nobody asked for: validate what arrives from outside before you believe it.
Exhibit two: the shape of it
The second quality you cannot see by running the application is its structure. Here is the entire front end:
This is a flat pile grouped by technical role. There is no architecture — not a bad one, none.
Every page reaches directly into api.ts and types.ts; nothing is behind a front door; no
boundary anywhere says "this is the only way in."
At four entities and nine screens it is perfectly navigable. That is the trap. It stops being navigable somewhere around forty, and by then it is a rewrite rather than a refactor. Nothing in the codebase will warn you as you cross that line, because nothing in the codebase knows the line exists.
While we are looking at the tree, notice the rest of what it tells us. Three unused image files left
over from the project template. On the server side, data.db, data.db-wal and data.db-shm — a
live database sitting in the source tree with no .gitignore to stop it being committed. And in
index.html, unchanged since the day the project was scaffolded:
That is quality 5 — finished — failing in one line. The application was built. It was never finished into an application.
Exhibit three: the errors nobody handles
Here is a real page, close to complete:
Three things are wrong here and all three are invisible while the network behaves.
The form calls .unwrap() — which makes a failed request throw — and there is no try/catch
around it. If the request fails, the user sees nothing at all and the page raises an unhandled
promise rejection.
The delete button is worse in the opposite direction: deleteAuthor(a.id) is fired and forgotten.
RTK Query's trigger does not reject unless you unwrap it, so a failed delete is silently
swallowed. The row stays on screen. The user assumes it worked.
And authors! — that exclamation mark is the developer telling the type checker "trust me, this is
never empty." It is the one construct in TypeScript whose entire purpose is to switch off the thing
you installed TypeScript for.
Every one of these is the happy path being mistaken for the whole path.
What was never there at all
The three missing marks are the shortest section to write, because there is nothing to show.
Tested (9): nothing. No test file, no test runner, not even a test script in either
package.json. Every claim in this post about what the application does is a claim about the
minute I checked it, and nothing anywhere is defending it tomorrow.
Standards (2): opted out. The plan asked for TypeScript. The front end got TypeScript — with
strict switched off, which is the single most important setting in the language. The back end is
plain JavaScript, no types at all. There is no formatter. The linter is present with two rules
enabled.
Finished (5): no. Editing works on the author page and nowhere else, even though the update
endpoints for books, users and reviews all exist and are wired up in api.ts. The application
stops halfway through its own feature list.
So what actually happened here
Put the scorecard next to the code and a pattern falls out that I did not expect to be this clean.
The two qualities the AI delivered are the two you can see by running the program. It works, and it reads nicely. Those are visible in a browser and in a diff — which is to say, they are visible in almost every piece of code ever written about, discussed, reviewed, or published. That is what the model learned from.
The seven it did not deliver are the ones that are only visible when you go looking. Nobody sees a missing test suite by opening the app. Nobody sees an absent architecture on a four-entity project. Nobody sees the type check that was never performed on a rating, until a rating arrives that is a word. These qualities live in decisions, in what was deliberately excluded, in scaffolding that never appears on screen — and they are systematically underrepresented in the text the model was trained on, because they are underrepresented in the code the industry writes.
Which brings me to the honest conclusion, and it is not "AI writes bad code."
The AI wrote what it was asked for. The plan described an application in complete detail and said nothing whatsoever about how it should be built — so for everything the plan did not specify, the model fell back on the average of everything it has read. The average is code that runs. The average is not quality code, because most code is not quality code. There was no reason for it to be, and nobody said otherwise.
What you do not specify, you get by accident.
One plan, one model, one application, judged once — an illustration, not a benchmark, and stark enough that it needs no inflating. What it does establish is that the failure is not random: every gap sits in the same place, the part of the work the specification never mentioned.
That is the problem, and this post stops here. You will be thinking review it afterwards and ask it to add tests — and you should. But review finds what is wrong with the code in front of it, and a missing test suite is not wrong code. It is code that is not there, and an absence has no line to comment on.
Asking for the tests afterwards has its own trap: they are written against what the code does rather
than against what it should do, so the rating that accepts oops gets no test that would catch it.
A green test you have never seen fail is a decoration.
The next one runs the identical experiment, with the identical plan, and changes exactly one input.
References
- The Qualities of Quality Code — the nine qualities used as the scorecard
- The Quality of AI Code — the claim this post measures



0 comments
No comments yet — be the first.
Please log in to leave a comment. Log in