Search them, take them, no account. Every card says whether we measured it or not. How we test →
Run through the harness against a control. The numbers, the p value and the cost of using it are on each page.
Written, reviewed and free to take. No run behind them, so no claim about what they do to an output. What that means.
Skills are organised by the job you are doing, not the job title you have.
Measured. 5 of them. Each was given a realistic task on real material, then the identical task was run again with the skill removed, 5 runs each way. Every output was graded alone against a rubric written by someone who had never seen the skill, by a session that was not told the other arm existed. Whether it passed was decided by a rule written down before a single run executed. Those pages carry the worst case, the median, the p value, and what the skill costs you as well as what it buys.
Not yet measured. 99 of them. Written, read for accuracy, free to take, and we have run no experiment on them, so we make no claim about what they do to an output. Their cards say so in the grid and their pages say so in a grey box above everything else. Each one names the test it is waiting for.
Cut. 4 of them, and you cannot install those, because they are not published. A skill that went through the harness and failed does not get a quiet demotion to the unmeasured pile. It is deleted, and its numbers stay in the results table permanently.
Unmeasured is not the same as failed. Everything in the second group might turn out to work, might turn out to do nothing, and we will tell you which when we have run it. The full method, the results table for every skill we ever tested, and the round we had to throw away.
Locul puts a skill in the right folder for Claude Code, Claude Desktop and Cursor, scrubs any environment variables before they reach you, and keeps the memory those skills read from up to date by itself. Mac and Windows.