The model writes the page. Separate checks decide if it ships.
A real run of a client brief, exactly as the engine logged it. Then the part that matters more: the suite that measures whether those checks actually work, and the four bugs it found.
the run · recorded 17 August 2026, nothing re-enacted
One command. The direction stage names the two reflexes it refuses, picks the typefaces, the pipeline embeds them and generates the photographs, then thirteen checks decide. Terminal output is verbatim from that run; only the pacing is added.
Open the page it produced.
what it rendered · desktop 1440 and mobile 375

page-1440.png

page-375.png
Two photographs generated in the same run and embedded, so the page is still one file with no external requests. 2074kb became 33kb, 1997kb became 20kb. Finished page: 63kb.
deterministic checks · these decide the verdict
| result | check | rule |
| pass | document structure | valid-document |
| pass | single file, no external assets | single-file |
| pass | contact integrity | contact-integrity |
| pass | palette fidelity (3 brief colours) | palette-from-brief |
| pass | type follows direction (Courier Prime, Public Sans) | type-deliberate |
| pass | no unresolved image placeholders | imagery-resolved |
| pass | no single-edge accent stripes | no-side-stripe |
| pass | page weight (249kb of 900kb) | page-weight |
| pass | link audit | cta-specific |
| pass | no horizontal overflow at 375px | responsive |
| pass | axe-core, no serious or critical violations | semantic-html / contrast-aa |
| pass | copy tells (dashes, banned vocabulary) | copy-tells |
| pass | no content hidden by animation, normal and reduced motion | motion-visible |
verdict: pass
claude-sonnet-5
$0.5036
17 August 2026
measuring the gate itself · 14 / 14 caught, 0 false fails
A gate that decides whether a page ships was trusted because good pages passed it. That is not evidence. The suite takes a page that already scored clean, injects exactly one known defect, and asks whether the right check fires, and only that one. Labels are correct by construction, since the harness knows what it broke. It runs with no API key and no network, so CI runs it on every pull request and fails on a stale scorecard.
14 / 14 caught
0 collateral failures
0 false fails on the clean page
39 s, no API key
what the suite found
A check that could not fail.
A 2000px element on a 375px viewport did not fail the responsive check, because the page sets overflow-x: hidden, which clamps scrollWidth. The check read zero overflow on a page overflowing by more than 1600 pixels. The drafting model had written one line of CSS that switched off the check meant to catch its own layout.
scrollWidth → 0px of overflow, on a page 1600px too wide
A defect that came back, and a fix the model routed around.
Two runs produced a tel: href missing one digit, in different positions, while the label read correctly. The advisory review caught the first and missed the second. Comparing digits is exact, so it became a gate. The next run then satisfied the digit comparison by pasting the display string into the href, spaces and all, which is not a valid tel URI. The fix had changed the failure mode rather than removing it.
tel:+6328452210 → tel:+63 2 8845 2210 → tel:+63288452210
Animations broke the screenshots, not the page.
The first animated draft screenshotted its hero part way through a fade, so the artifact a human reviews showed a headline that looked broken while the page itself was correct. Captures wait for animations to land now. The gate that came out of it renders twice, once normally and once under reduced motion, and fails on any text left invisible.
A decoration threw out a paid draft.
When image generation ran out of credit, the failure threw out of the draft stage and left an empty run directory, discarding a page that had already been paid for. Image failures degrade now: the placeholder stays and the gate refuses the page, so the draft survives on disk.
the number that mattered more than the detection rate
Across twelve runs of the same brief in one day, the first-pass verdict was 8 pass and 4 fail. Every one of the failures lands in the last four runs, after the gate had grown from six checks to thirteen. Yield fell as the gate tightened.
That is the gate working, and on its own it is worthless: a page that gets rejected and then re-rolled by hand is not a pipeline. So a rejected page is no longer handed back. The engine reads the failing checks to the drafter, says change only what these name, and re-drafts, up to two repairs. The art direction, the embedded typefaces and the generated photographs are prepared once and reused, so a repair pays for markup and nothing else.
The checks were already saying exactly what was wrong and where. Nothing was listening.
what the art direction stage decided, before any markup existed
Voice.
plain, procedural, unhurried. The object this brand would be if it were printed: the signed quarterly restore-drill report itself.
Reflexes it named and refused.
The navy-and-white enterprise security palette, and the editorial serif-italic layout with tracked small-caps labels and ruled dividers. Those are the first and second training-data defaults for this category, and earlier drafts had landed in both.
What it chose instead.
A committed strategy built from the brief's own three hex values, described as government-ledger cream, bank-vault navy, and rust stamp ink. Courier Prime for display against Public Sans for body, neither on the reject list, both embedded as woff2 data URIs so the page still makes no external request.
the gate could tell a page was wrong. it could not tell it was dull.
a metric that did not survive calibration
The pages kept passing every check and still looking wrong: one left column of heading and paragraph, repeated, with a third of the viewport dead. The obvious move was to measure it, as leftmost-to-rightmost content span over viewport width.
Calibrated against four real runs, the page judged badly composed scored 63.3 percent and the page judged well composed scored 63.3 percent. Full-bleed bands and a wide header mask a stranded column. The metric did not discriminate, so it is not in the gate. Composition moved into the rules file as judgement, where the review can cite it and a person can argue with it.
The split is the point. Eleven deterministic checks decide pass or fail; the model reviews alongside and can only advise. Advisory findings that turn out to be exact comparisons get promoted into the gate, which is how contact integrity and palette fidelity got there. Every number here is read from the run's own qa-report and the committed scorecard, not typed by hand. Repository:
landing-page-engine.