Six skills that produce copy rather than grading it. Each one carries a slot order, a device bank or a set of placement rules, because the model already writes a decent sentence and the ordering is where the work actually is. How we test →
Written, reviewed and free to take. No run behind them, so no claim about what they do to an output. What that means.
Almost every skill in the rest of this directory checks something that already exists. This category deliberately does not. Everything here produces the artifact.
That is a correction, and it is worth saying why out loud, because the cause was our own measurement harness. An audit emits a findings list, a findings list can be scored against mechanical ground truth, and so an audit is cheap to measure. A piece of copy is an artifact whose quality is partly a matter of judgement, so it is expensive to measure. We spent two rounds building a harness, and then, without deciding to, we let what the harness could measure decide what got written. Forty-three of our first fifty-four skills were audits. That is the streetlight effect with a directory attached to it.
Nobody wakes up wanting to audit their landing page. They wake up wanting the page.
So what stops these from being "write good copy" dressed up? The same gate as everywhere else on this site: a skill has to carry something a strong model does not already have. A model writes a competent headline unprompted. What it does not do unprompted is put the damaging admission in the order where the clause after "but" is the one the reader believes, or refuse a simile in the first line because that is the exact place a reader senses a setup coming, or insist that the second most read element on the page is the postscript and that it is allowed to be only one of three things. The value in this category is slot orders, placement rules, device banks and specificity bars. Where a file is only encouragement, it is not here.
The IP line, since this material has a history. Several of the methods behind these files originate in frameworks that are attached to well-known practitioners. Naming a published framework is fine and we do it. Wearing a real person's voice is not, and none of these files does it. There is no roleplay, no fabricated quote, no named individual used as the structure of a skill. Where a source was a voice, we kept the technique and dropped the person.
Measurement. Some of this has an objective spine and some does not, and the split is not where you would guess. An offer page against a fifteen-slot skeleton is mechanically checkable: either the slots are present and in order or they are not, either the damaging admission runs flaw-then-strength or it runs the other way. Whether a social post is actually funny is not checkable at all, and the honest route for that one is blind pairwise preference, which is weaker evidence and is labelled as such when we run it. Nothing here has been measured yet. Each page says which route it is waiting for.
What is not here. Anything that grades copy you already wrote. Those live elsewhere in the directory and they are good at their job: there is a headline test in writing, an offer review and a post review in marketing, and a machine-tell remover in humanize. Use those after these, not instead of them.
Measured means the skill was given a realistic task on real material, then the identical task was run again with the skill removed, five runs each way. Each output was graded alone, against a rubric written by someone who had never seen the skill, by a session that was not told the other arm existed. Whether it passed was decided by a rule written down before any run executed. Those pages carry the worst case, the median, the p value, and what the skill costs you as well as what it buys.
Not yet measured means exactly that. It is written, it has been read, it is free to take, and we have run no experiment on it, so we make no claim about what it does to an output. It is not a skill that failed. Skills that failed are not published at all, in either state, and their numbers are in the results table.
Measuring one skill properly costs roughly twenty model sessions. We are working down the queue and moving skills from the second group into the first. Read the full method, or go back to all skills.