Skills/Product management
8 skills · 0 measured · no API keys, no install, no scraping

Claude skills for product managers

Six skills that produce the specification, the tickets, the blueprint and the decision. Each carries an ordering, a gate or a template where getting the sequence wrong is the expensive part. How we test →

All 8 shown

8 skills we have not measured yet

Written, reviewed and free to take. No run behind them, so no claim about what they do to an output. What that means.

Product managementNot yet measured Activation and first-run design Activation is a behaviour that predicts retention, not a screen. The test deciding whether a candidate qualifies is empirical, and the honest branch is marking it provisional. skill · 3,938 words · MIT Read the write-up
Product managementNot yet measured Expert council debate A structured debate method where the lenses are disciplines and published frameworks rather than people, divergence is the output, and the synthesis has to commit to one path anyway. skill · 3,574 words · MIT Read the write-up
Product managementNot yet measured Idea to scaffolded product Runs the unglamorous scaffolding of a new product end to end, and its real content is the autonomy model: where it stops, where it refuses to guess, and how it stays legible while running. skill · 3,567 words · MIT Read the write-up
Product managementNot yet measured Persona set build Demographics predict nothing about a product decision and constraints predict most of it, so the cards carry budget ceilings and incumbents, and the disagreements between them are the finding. skill · 2,876 words · MIT Read the write-up
Product managementNot yet measured Pricing and packaging build The routing tree is ours rather than received wisdom, and the file says so on its first page: four symptoms, four remedy classes, and one branch whose instruction is not to touch the price. skill · 3,761 words · MIT Read the write-up
Product managementNot yet measured Spec to tickets Writes the specification and the tickets, but resolves every field, endpoint and table against the real codebase before any of them reach the document. skill · 3,438 words · MIT Read the write-up
Product managementNot yet measured Unblock decision protocol Most autonomous systems stop when any single condition looks risky, which is why they stall. Here a halt needs all four at once, and everything else becomes a decision plus a deferred item. skill · 3,155 words · MIT Read the write-up
Product managementNot yet measured Viability blueprint A structured pre-launch interview that produces a go-to-market blueprint, built around the rule that you ask one to three questions at a time and never accept a vague answer. skill · 3,809 words · MIT Read the write-up
no API keys, no install, no scraping

The list

Product work is mostly deciding things in an order, and the order is where it goes wrong. A specification with no out-of-scope section is where scope creep enters. A ticket list sorted by priority rather than by dependency has a first ticket nobody can start. A field name invented during speccing gets copied into tickets, then into code, then into a migration. None of those are failures of effort.

That is the shape of everything in this category. Each file carries a sequence, a gate, a template or a hard constraint, and each one names the specific way the sequence gets broken.

These produce artifacts rather than reviewing them, which is a deliberate correction. The rest of this directory leans heavily on audits, and the honest reason is that an audit emits a findings list that a harness can score against mechanical ground truth, so audits are the cheap thing to measure. We let that decide what we wrote for longer than we should have. Forty-three of our first fifty-four skills checked something somebody else had already made. Almost nobody goes looking for that. They go looking for the thing made.

One file in here needed the most care of anything on the site. The convened-council method came from material built around named living practitioners speaking in the first person. That is not publishable and we did not publish it. What we kept is the structure that actually does the work: lenses defined by expertise and by a named published framework, selected by the altitude of the question, forced into three explicit buckets where the disagreement is the deliverable rather than a defect, and then a synthesis that has to commit to one recommendation instead of retreating into "it depends". A lens built on a published framework can be checked. A lens built on a person cannot, and it changes when they do.

Measurement. The split here is unusually clean. A specification either has an out-of-scope section or it does not; its stories either match the required shape or they do not; its ticket order either forms a valid dependency graph or it does not; every field it names either resolves in the repository or it does not. That is an objective spine and those files can be measured under the standard rule. A viability verdict and a council synthesis have no such spine, and the only honest route for them is blind pairwise preference, which is weaker evidence and gets labelled that way. Nothing here has been measured yet, and every page names which route it is waiting for.

What is not here. Roadmapping, prioritisation scoring, and stakeholder management. The first two are well covered by frameworks that need no skill file to explain them, and the third is not a thing a file can teach you. Pricing appears here only as a packaging decision, because the pricing review that grades an existing pricing page lives in marketing.

What the two labels mean

Some of these carry a number. Most do not, and they say so.

Measured means the skill was given a realistic task on real material, then the identical task was run again with the skill removed, five runs each way. Each output was graded alone, against a rubric written by someone who had never seen the skill, by a session that was not told the other arm existed. Whether it passed was decided by a rule written down before any run executed. Those pages carry the worst case, the median, the p value, and what the skill costs you as well as what it buys.

Not yet measured means exactly that. It is written, it has been read, it is free to take, and we have run no experiment on it, so we make no claim about what it does to an output. It is not a skill that failed. Skills that failed are not published at all, in either state, and their numbers are in the results table.

Measuring one skill properly costs roughly twenty model sessions. We are working down the queue and moving skills from the second group into the first. Read the full method, or go back to all skills.