# 08 — Gaps, open questions and honest tensions

---

## 1. Build-fresh items

| # | Item | Why |
|---|---|---|
| **G1** | **Resolve the `use-sg-playwright` duplicate** | Two copies, 42 words apart, one carrying a claim the code contradicts. The thesis's predicted failure, live. One repo owns it; the other imports |
| **G2** | **A `permissions` field in SKILL.md frontmatter** | Declared, not enforced — one field makes the PBOM computable for skills and starts the identity design with a day's work. `05__` §4 |
| **G3** | **One eval for one skill** | Zero exist. The scoring differentiator (*"provenance, evals, and an auditable graph"*) has two of its three legs; this is the third |
| **G4** | **A `version` field per skill** | The cheapest possible start on the package thesis |
| **G5** | **The authoring guide** | `02__` §§2+4 — five conventions and three rules, currently implicit in eight artefacts |
| **G6** | **A skill review process** | The Feb 2026 trio proved it can be done; nothing since. Even a checklist referencing that precedent is a start |
| **G7** | **The registry v0** | `/catalogue/`, generated. The thesis's npm-shaped hole, started as a static page |

## 2. Open questions

| # | Question | Where it stands |
|---|---|---|
| **Q1** | **What is a skill version?** Semver assumes behavioural compatibility claims; a skill's behaviour depends on the model reading it. Does 1.2.0 mean anything when the interpreter changes underneath? | Unaddressed anywhere. The intent-over-capability claim makes it harder, not easier |
| **Q2** | **How do you test intent?** The thesis demands CI for skills; no one has said what a failing test looks like for *"an opinionated way of approaching something"* | Evals are the obvious answer and the estate has none |
| **Q3** | **Who arbitrates triggering conflicts?** The governance rules say skills compete for triggering — but with no registry, nothing detects that two skills claim one intent | The catalogue could — a build-time overlap check on trigger vocabularies is feasible today |
| **Q4** | **Is the base vault the skill, or its source?** If the SKILL.md is a projection of the vault, which one is versioned, reviewed and sold? | The economy briefs assume the vault; the shipped practice is SKILL.md-only |
| **Q5** | **Does the permission set belong to the skill or the invocation?** The 18 June brief wants moment-of-authorisation — which implies the *invocation* computes it — while the frontmatter proposal declares it statically | Both, probably: declared ceiling, computed grant. Needs a design |
| **Q6** | **Would Tessl-style aggregators accept vault-backed provenance?** The monetisation model rides platforms whose scoring is exactly the *"gameable popularity"* the vault position rejects | A real conversation exists; the answer is not recorded |

## 3. Honest tensions

1. **The theory criticises the practice, and both are yours.** Every shipped skill is the *"static photograph"* the projection paradigm dismisses. Publishing that is the credible move; resolving it is a research programme.
2. **One green cell.** The lifecycle table (`03__` §1) scores the estate's skills against the estate's own checklist and passes only documentation. The site leads with its best artefacts *and* their worst scorecard.
3. **Skills earn their keep on distinct tasks — and this estate ships eight, while the theory imagines thousands.** The governance rule that keeps quality up is in tension with the economy that needs volume.
4. **The intent claim cuts both ways.** If skills describe intent in English, then model drift is dependency drift — and there is no lockfile for the interpreter. Q1 is not pedantry; it is the hard version of the thesis.
5. **The strongest week of thinking produced the least code.** 1–4 June: ~35,000 words, zero commits to `library/skills/`. The site is partly a mechanism for closing exactly that gap.

## 4. Loose ends worth an hour each

- Diff the two `use-sg-playwright` copies and record which claims diverge (one is already known wrong).
- Check whether the sg-playwright skills README's four candidates were ever built elsewhere.
- Confirm whether `skill-creator`-style eval tooling can point at estate skills as-is.
- Ask whether the Tessl conversation is still live before the economy page ships.
- Extract the trigger vocabularies of all 8 skills and run the overlap check (Q3) once — it may already find a collision.

---

This document is released under the Creative Commons Attribution 4.0 International licence (CC BY 4.0).
