ground truth
The model writes CSS it cannot see. This is how it looks.
An agent writing CSS is working blind. The Probe is a small Playwright harness that gives it eyes: it opens a real browser at a chosen width, signs into the portal, walks a set of routes, and measures what actually rendered — which columns are visible, whether any cell overflows its box, how many lines a title wraps to. The measuring is the agent's job. The design decision stays mine.
The problem
An agent writing CSS is working blind. It can read the stylesheet and it can read the markup, but it cannot tell you whether the invoices table still reads at a laptop width, whether a client name truncates at 1000px, or whether the primary cell has quietly wrapped to three lines. It will tell you the change looks correct, and it will mean it.
I found this out the way everyone does. A column change passed review, passed its tests, and shipped a table where four of the six columns said nothing at the width most people actually use.
What I did about it
So I built it eyes.
.claude/probe/ opens a real Chrome at a chosen viewport, signs into the admin or client portal, walks a list of routes, and reports what rendered. Not "did it compile" — whether any cell overflows its box (scrollWidth > clientWidth), how many lines the primary cell takes at that width, which columns survived, what the computed table-layout resolved to. It writes PNGs at named widths so a session can look at its own work instead of asserting that it looks fine.
The split matters more than the tooling. The probe told me six columns were truncating at 1000px. I decided the fix was removing them rather than shrinking the type, and wrote down why that was a design call and not a cleanup. The harness reports; it does not decide.
Around it sits the rest of the scaffolding that makes a model safe to turn loose in a design system. The token layer has contract tests that fail the build on a dangling var() or a token redeclared in the same scope — one of them caught a dead shadow declaration on the day it was written. CLAUDE.md is tracked in the repo and documents the commit protocol for when more than one session is working at once, which happens, and which buried another session's work more than once before the rule existed. The Copilot instructions were rewritten against the code rather than left aspirational.
Claude writes most of the documentation here and a good share of the commit messages. It is never listed as an author. Authorship is accountability, not credit: git blame exists to answer who to ask about a line, and that has to be someone who can be asked — who remembers the decision, can defend it, and can be found when it breaks. A model is none of those, and putting it in the author field turns blame from a tool into a dead end. The disclosure belongs where it is useful: CLAUDE.md is tracked, the Copilot instructions are maintained against the code, and this page exists. Same rule as the harness — it reports, I decide; it writes, I sign.
The harness is the second half of the job. The first half is deciding what to build, and that is where I reach for the model hardest.
Hedgewitch Horticulture needed a type stack — display, heading, body, accent. Three candidate displays, five headings, two bodies, two accents: sixty combinations. I had been evaluating them the way everyone does, as stacked panels, one candidate per panel. At ten panels it was already unreadable, and ten panels covered one axis of four.
So I had a picker built instead. Four dropdowns, four CSS custom properties on :root, two sample panels that re-render live. The entire engine is four setProperty calls. The design is in the sample content, not the code.
The rules it enforces are not typographic math. They govern what the specimen must contain so the eye cannot be fooled:
- Every heading renders twice, mixed-case and all-caps, stacked. Vintage display faces read radically differently in each treatment, and dropping one variant for cleanness hides exactly the problem you are looking for.
- Two panels on two background tones, because the same pairing reads differently on each.
- Real copy from the client's industry, never lorem. A garden company's H2 is "Native Meadows" — the eye latches onto familiar words, and that is the point.
- License tier verified before a face is allowed into the dropdown at all, with personal-use-only faces shown for benchmarking and flagged unshippable.
That last rule is the one with teeth. The display face I liked best was licensed personal-use only. It sits in the lab labelled as my favourite and marked unusable in the same breath.
The exploration also killed the incumbent. PP Cirka was the locked baseline — already specified, $170 for a web license — and it lost. The final spec retires fifteen faces and notes that dropping it saves the $170. What shipped is free: Otto Attack, Della Respira, Manrope.
Then I generalised the picker into a skill so the next project would not have to re-code it. Three minutes between the working template and the version with the brand stripped out.
Honest about its reach: it has been deployed once. That once is documented end to end — wireframes, lab, decision record, and the tokens sitting in the live variables.css. And the stack kept moving after launch; the spec locked one heading face and production runs another. The lab picked a stack. The site kept editing it. Which is what the harness is for.
What it changes
- An agent can check its own layout work at a real width instead of asserting it looks right
- Truncation, wrap count and column visibility are measured, not eyeballed
- Token drift fails the build — 15 contract test files across 483 tokens in 11 files
- Concurrent agent sessions stopped burying each other's commits
- The judgment stayed mine: the harness reports, I decide, and the reasoning goes in the work log
What is in it
- Playwright harness opening real Chrome at chosen viewport widths
- Authenticated walks through admin and client portal routes
- Per-cell overflow detection (scrollWidth vs clientWidth)
- Line-count measurement on primary table cells
- Computed-style readback: table-layout, resolved widths
- Screenshots written at named breakpoints for visual comparison
- Token contract tests: every name, every literal value, no dangling var(), no same-scope redeclaration
- CLAUDE.md commit protocol for concurrent agent sessions, tracked in the repo
- Copilot instructions rewritten against the code, not aspirational
- A work log recording the design decisions the measurements prompted
- Agent contributions disclosed in tracked instructions, never in the commit author field
- A four-axis type-pairing lab, generalised into a reusable Claude Code skill, with font licensing enforced at the dropdown
Built with