Where a number came from, whether the source is what it appears to be, and the named biases that make a confident analysis wrong.
The failure mode in assisted research is not that the assistant cannot find things. It is that a number arrives with a citation attached, the citation points at an article, the article points at another article, and the chain terminates in something that never said it. That pattern has a shape, it is common, and it is checkable, which makes it exactly the kind of thing a skill can carry.
So the skills here are about the chain rather than the search. What separates a primary source from something that merely cites one. What a preprint is and is not. How to tell a predatory journal from its metadata rather than its design. What a retraction looks like from the outside and where to check. How a statistic gets laundered from a bounded estimate into a headline figure through three intermediate retellings, and how to walk it back.
Two of these are not about literature at all. Survey design carries the named biases that make a result meaningless before a single response arrives, which is the point at which the damage is cheapest to avoid. And the analysis sanity check carries the specific reversals that make a chart say the opposite of what it appears to: base rate neglect, a paradox where every subgroup moves one way and the total moves the other, and survivorship, where the missing data is the finding.
What none of these do. They do not run the search for you, they do not have access to anything behind a paywall, and they cannot verify a claim whose source is offline. They tell you what to check and what a failed check looks like.
Written, reviewed and free to take. No run behind them, so no claim about what they do to an output. Each page says so at the top.
Measured means the skill was given a realistic task on real material, then the identical task was run again with the skill removed, five runs each way. Each output was graded alone, against a rubric written by someone who had never seen the skill, by a session that was not told the other arm existed. Whether it passed was decided by a rule written down before any run executed. Those pages carry the worst case, the median, the p value, and what the skill costs you as well as what it buys.
Not yet measured means exactly that. It is written, it has been read, it is free to take, and we have run no experiment on it, so we make no claim about what it does to an output. It is not a skill that failed. Skills that failed are not published at all, in either state, and their numbers are in the results table on the hub.
Measuring one skill properly costs roughly twenty model sessions. We are working down the queue and moving skills from the second group into the first. Back to all skills.