Skills/Research
6 skills · 0 measured · no API keys, no install, no scraping

Claude research skills, aimed at the citation chain

Where a number came from, whether the source is what it appears to be, and the named biases that make a confident analysis wrong.

The failure mode in assisted research is not that the assistant cannot find things. It is that a number arrives with a citation attached, the citation points at an article, the article points at another article, and the chain terminates in something that never said it. That pattern has a shape, it is common, and it is checkable, which makes it exactly the kind of thing a skill can carry.

So the skills here are about the chain rather than the search. What separates a primary source from something that merely cites one. What a preprint is and is not. How to tell a predatory journal from its metadata rather than its design. What a retraction looks like from the outside and where to check. How a statistic gets laundered from a bounded estimate into a headline figure through three intermediate retellings, and how to walk it back.

Two of these are not about literature at all. Survey design carries the named biases that make a result meaningless before a single response arrives, which is the point at which the damage is cheapest to avoid. And the analysis sanity check carries the specific reversals that make a chart say the opposite of what it appears to: base rate neglect, a paradox where every subgroup moves one way and the total moves the other, and survivorship, where the missing data is the finding.

What none of these do. They do not run the search for you, they do not have access to anything behind a paywall, and they cannot verify a claim whose source is offline. They tell you what to check and what a failed check looks like.

The list

The list

6 skills we have not measured yet

Written, reviewed and free to take. No run behind them, so no claim about what they do to an output. Each page says so at the top.

ResearchNot yet measured Analysis sanity check The named reversals that make a confident chart say the opposite of what it appears to, each with the specific check that catches it before anyone presents it. skill · 3,101 words · MIT Read the write-up ResearchNot yet measured Claim provenance trace Traces a number to the earliest document that actually contains it, names what changed at each hop, and treats an empty origin as the finding rather than a dead end. skill · 2,716 words · MIT Read the write-up ResearchNot yet measured Competitive teardown Ranks every claim in a competitive comparison by how it was obtained, and requires each cell to carry a number, a plan name or a named capability rather than a tick. skill · 3,098 words · MIT Read the write-up ResearchNot yet measured Literature search strategy Turns a query into a designed search with a written protocol that can be rerun, plus a scaled-down two-hour version for a commercial question that is not a review. skill · 3,440 words · MIT Read the write-up ResearchNot yet measured Source credibility triage Rates a source and says why, with a hierarchy that inverts for specification claims and a cannot-assess verdict that counts as a real answer rather than a failure. skill · 3,371 words · MIT Read the write-up ResearchNot yet measured Survey design audit Audits a questionnaire while the damage is still cheap to undo, and treats who never answered, rather than how many did, as the dominant threat to the result. skill · 3,273 words · MIT Read the write-up
What the two labels mean

Some of these carry a number. Most do not, and they say so.

Measured means the skill was given a realistic task on real material, then the identical task was run again with the skill removed, five runs each way. Each output was graded alone, against a rubric written by someone who had never seen the skill, by a session that was not told the other arm existed. Whether it passed was decided by a rule written down before any run executed. Those pages carry the worst case, the median, the p value, and what the skill costs you as well as what it buys.

Not yet measured means exactly that. It is written, it has been read, it is free to take, and we have run no experiment on it, so we make no claim about what it does to an output. It is not a skill that failed. Skills that failed are not published at all, in either state, and their numbers are in the results table on the hub.

Measuring one skill properly costs roughly twenty model sessions. We are working down the queue and moving skills from the second group into the first. Back to all skills.