How Our AI Loop Keeps the Skills Honest
Shipping seven agent skills is a claim: hand a coding agent this tooling and it builds better bestax apps than it would cold. Claims want numbers. So we grade cold agents against a frozen rubric, feed every finding back into the skills, and keep the receipts. This post is the receipts.
This is post seven of the catch-up series (tracker). The last post was the what and the why: docs, skills, and a catalog that make the library machine-legible. This one is the proof and the process, how we verify the skills actually improve agentic coding, and how we improve the skills themselves, iteratively, on evidence. None of it is hypothetical. Every number below comes from a committed record you can read.
Skills Are a Shipped Product
Start with the stakes. The first line of the skills directory's contributor contract says it plainly: "Agent Skills are a shipped product." The seven skills reach users two ways, bundled into every create-bestax scaffold and installable with npx skills add, and the contract draws the conclusion: treat changes like library code, because "they get bug reports (#194, #195, #196, #197) and ship to users."
My favorite rule in that file is about repair. When fixing a skill bug, fix "the guidance that produced the bad output, not just the example." A wrong example is a symptom. The skill taught the mistake, so the skill is what gets fixed.
A shipped product with bug reports deserves what every shipped product deserves: proof that it works. Ours is a harness.
The Proof: Graded Cold-Start Runs
How do you grade "agents code better with this"? Cold. The harness lives in eval/skill-loop, and one run is deliberately unsentimental:
- Scaffold a fresh app with the current tooling, the seven skills plus the CLAUDE.md the scaffolder generates.
- Hand a frozen brief to a cold-start
claude -psession: fresh directory, no repo context, empty memory, the library installed from the registry. Exactly what a stranger's agent sees on day one. - Collect mechanized metrics from the result: build pass, type errors, inline styles, raw Bulma classNames, hand-rolled tags, which skill files the agent actually read.
- A grader subagent scores a frozen 100-point rubric: build integrity, component adoption, prop fidelity, hallucination penalty, custom-component conformance, theming, site completeness, skill engagement. The mechanized metrics are ground truth the grader may not contradict.
The rubric's first check is my favorite kind of paranoia. An untouched scaffold typechecks, builds, and reports zero inline styles, zero invented APIs, zero hand-rolled tags. Graded naively, doing nothing collects 50 of 100 points. So before any category is scored, a gate checks whether the builder modified the app at all and zeroes the run if it didn't. Doing nothing is a failed run, not a clean sheet.
The Improvement Loop: 85 to 95.2

Measurement without iteration is trivia. The loop is grade, revise, re-run: read the scorecards, change the tooling (the skills, the generated CLAUDE.md template, the catalog generator), rebuild, and put the next cold agent through the same brief. We ran ten iterations against a fixed SaaS-site brief and wrote the whole thing down: findings in #363, method and per-run evidence in the committed report.
| Baseline (i01) | Revised runs (i02–i10) | |
|---|---|---|
| Rubric score | 85/100 | mean 95.2 (median 96, min 89, max 99) |
| Raw Bulma classNames | 42 | 0 in all nine runs |
| Custom CSS added | 77 lines | 10 to 21, converging on a sanctioned ~10-line pattern |
| Builder cost / turns | $10.55 / 127 | mean $6.00 / 79 |
The score rose while the cost of earning it fell 43% and the turn count fell 38%. Better guidance doesn't just produce better apps; it produces them with less thrashing.
The finding I keep reusing: placement beats content. Facts on always-loaded surfaces (the generated CLAUDE.md, the SKILL.md bodies, the catalog itself) held in every subsequent run. Facts one reference-hop away failed stochastically, even when the pointer to them was read. If you want an agent to know something every time, put it where the agent already is.
That finding is also the honest name for what this loop automates: preserving the model's context. Training data won't carry current bestax, which was the last post's whole argument. The skills are that context, and the loop is how it stays correct and gets better, on a schedule we control, at the pace of a script instead of the pace of organic bug reports.
Two disciplines keep the numbers trustworthy. The yardstick is frozen for the whole loop, same brief, same rubric, same caps, same model, so improvements go into the tooling and never into the test. And the graders get audited: three of the ten scorecards contained a factual error, caught by cross-checking them against the run transcripts before anything acted on them. After the final iteration, a compare-only pass confirmed the gains without shipping anything unvalidated.
The harness is reusable on purpose. Swap the brief, the skills state, or the model, freeze everything else, and the same protocol answers a new question. The first loop was the expensive one; the next ones are cheap.
Captured Agent Runs
Numbers convince maintainers. Seeing convinces everyone else. Storybook has a Skills section where the showcases are agent-generated: a skill's canonical example, built by an agent, rendered live. Each one opens with an ExampleMeta header recording exactly what produced it:
<ExampleMeta
skill="bestax-custom-component"
skillHref={SKILL_HREF}
model="Claude Opus 4.8"
date="2026-06-27"
prompt="Build a ProfileCard component — avatar on top, then name, role, and a short description — following the bestax custom-component skill."
/>
Prompt in, rendered output out, versioned alongside the components. The date tells you which era of the skill produced the example, and the current set lists two different models, Claude Opus 4.8 and Claude Fable 5, because the record keeps what actually ran. It's the same idea as the eval harness at a different altitude: a run you can replay with your eyes.
Between Loops: Review and Gates

An eval loop runs when we run it. Between loops, verification is continuous.
Every skill change is a PR through the same adversarial review as library code, CodeRabbit plus a Claude deep review, with a human doing every merge (the AI-assisted development guide documents the machinery). That's how the bestax-icons skill landed, how dark-mode contrast rules got into the theming and layout skills, how the skills-sync check itself arrived, and how bestax-optimize shipped.
And CI fails outright when skill content drifts from the library:
- The component catalog the custom-component skill leads with is generated from the API docs; CI regenerates it and fails on any diff. The generator itself fails if an exported component lacks an API page, so the 87-entry list an agent reads is complete by construction.
- A skills-sync conformance check: the theming skill's shipped inventories must name every registered
--bulma-*variable and every component with its owncolorprop. Add a themeable component without updating the skill and the build says no. - The contributor checklist makes it a habit, not a heroic act. The heading in the component checklist reads, verbatim: "Skills sync (same PR, always)."
What the Loop Keeps Finding
A working loop's output is a to-do list. The ten runs didn't just raise the score; they filed library bugs and feature gaps that are open right now:
- #367: color props typecheck values that ship no CSS, and
Boxcolor falls through to a text class - #368: the form
labelprop renders a Label but wires nohtmlFor/id - #369: no scheme-aware background route for dark mode
- #370: API-consistency traps agents actually hit, like
isFullWidthvsisFullwidthandTag isLight - #371: scaffold polish, from a
.gitignoregap to a dev-server port fallback
I like that list more than I like the 95.2. The eval's job isn't to certify the skills; it's to find where the skills and the library still let an agent down, faster than waiting for a user to hit it.
The honest limits come straight from the harness docs: one brief, one model, ten runs, and single-run scores swung by six points with identical tooling, so one run is one sample and trends need several. That's fine. The yardstick is built, the baseline is recorded, and the next loop inherits both.
The tracker has the rest of the publishing plan, and the last post has every entrypoint if you want the skills in your own agent. This post just wanted to show you the grading. Skills are a shipped product, so they get what shipped products deserve: proof.
