Editing as an ordered set of passes, documentation as a typed artefact, and the two headline tests that replace taste with something falsifiable.
A model that cannot write is not the problem any of these solve. The problem is that writing work gets done in the wrong order, judged by taste that nobody can articulate, and shipped in a shape that does not match what the reader came for.
So the skills here carry procedures rather than encouragement. Documentation has four distinct modes with different contracts, and the most common documentation failure is mixing two of them in one page. A style guide is a list of about twenty decisions, and a guide that has not made them will see every one relitigated in every review. A headline can be tested rather than argued about: a stranger should be able to say what the piece is about from the title alone, and the title should name something that exists rather than a feeling about it.
Where this category is honestly weak. Prose quality has no mechanical ground truth, which is exactly the condition under which a skill is hardest to justify. Our published method note is explicit that where no objective spine exists, the only available evidence is blind pairwise preference, which is weaker evidence and is labelled as such wherever we use it. Several skills here are waiting on that, and until they have been through it their pages say so plainly rather than implying a result they do not have.
The exceptions are the ones with real specifications behind them: named plain-language standards, a changelog format with defined categories, and an error message contract with required parts. Those have something to check against, and they are the ones we would test first.
An editing pass order skill was written for this category and deleted before publication. It argued that editing runs in five passes and that doing the sentence work before the structural work wastes most of the sentence work. That is true, and it is also something a capable assistant says without being asked.
Its own pages gave it away. The draft conceded that a strong control already knows the words developmental, structural and copy edit, and that the only real difference was a refusal to descend a level early. A refusal is discipline, not knowledge, and two rounds of measurement on this project say discipline-only skills do not move anything.
It was cut on that basis alone, without being measured, which is a judgement call and we would rather state it than hide it. The gate questions it carried have been folded into the style guide and plain language skills, where they sit next to material that is genuinely checkable.
Written, reviewed and free to take. No run behind them, so no claim about what they do to an output. Each page says so at the top.
Measured means the skill was given a realistic task on real material, then the identical task was run again with the skill removed, five runs each way. Each output was graded alone, against a rubric written by someone who had never seen the skill, by a session that was not told the other arm existed. Whether it passed was decided by a rule written down before any run executed. Those pages carry the worst case, the median, the p value, and what the skill costs you as well as what it buys.
Not yet measured means exactly that. It is written, it has been read, it is free to take, and we have run no experiment on it, so we make no claim about what it does to an output. It is not a skill that failed. Skills that failed are not published at all, in either state, and their numbers are in the results table on the hub.
Measuring one skill properly costs roughly twenty model sessions. We are working down the queue and moving skills from the second group into the first. Back to all skills.