Locul / Tools / The State of CLAUDE.md
The State of CLAUDE.md: what 6,967 public files contain
We downloaded 6,967 CLAUDE.md files from public GitHub repositories, plus 2,484 AGENTS.md files to compare, and measured what is in them: length, sections, commands, imports, the files that sit next to them, and how often anyone edits them. Below are the numbers, a template built from them, and a checker for your own file.
The short answer
A CLAUDE.md file is the Markdown instruction file Claude Code loads at the start of every session. A typical public one is 100 lines long (about 1,256 tokens), and 22% are longer than Anthropic's 200-line target. Most name real commands (76% list at least one tool such as npm or pytest), but 37% still open with the boilerplate line /init writes, only 12% have a security heading, and only 3.8% use @imports. 38% of files were committed once and never edited again.
The sample: 6,967 root-level CLAUDE.md files from public, non-fork repositories, drawn in proportion to the sizes of the roughly 590,000 such files GitHub's code search indexes. 6,688 of them have real content; the rest are one-line pointers or near-empty. What stands out:
- The median file is 100 lines; 22% are over 200. Anthropic's docs say to "target under 200 lines". 4.4% of files run past 500 lines. See the distribution.
- 37% still carry the /init header. The line "This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository" is left in place. It tells Claude nothing it does not already know.
- Commands are the core. 76% of files name at least one command-line tool, and 73% mention tests somewhere. npm alone appears in 26% of all files. See the tools.
- Security is the gap. 36% mention secrets, keys or
.envanywhere, and only 12% give security its own heading. 7% of repos commit.claude/settings.local.json, a file Claude Code keeps out of git when it creates it. - AGENTS.md is moving in next door. 21% of repos with a CLAUDE.md also have an AGENTS.md, and the most common
@importtarget is AGENTS.md itself. Compare the two. - Most files are written once. 38% have a single commit, the median is 2 commits, and 34% have not been touched in six months. See upkeep.
- Popular repos write longer, more complete files. In repos with 100+ stars the median is 138 lines, 49% have a testing section (all files: 33%) and 82% name a command. See the table.
What the docs and earlier research say
"Size: target under 200 lines per CLAUDE.md file. Longer files consume more context and reduce adherence."
"We conclude that while context files are useful for specifying non-standard coding practices, any attempts to improve performance should be rigorously evaluated before deployment."
"non-functional requirements such as security (14.8%) and performance (14.5%) are rarely specified."
"Never commit secrets" was the most common helpful constraint.
How long is a typical CLAUDE.md file?
The median CLAUDE.md is 100 lines, about 675 words or 1,256 tokens. Half of all files fall between 59 and 180 lines. Anthropic's guidance is to "target under 200 lines per CLAUDE.md file", and 22% of files are over it; 4.4% pass 500 lines. Pointer files and near-empty files are left out of these length numbers.
Lines per CLAUDE.md file
Share of 6,688 files with real content, by line count. Amber bars are over Anthropic's 200-line target.
Lines are counted as they are in the file, blank lines included. Tokens are estimated at 4 characters per token.
| Measure | 10th pct | 25th | Median | 75th | 90th |
|---|---|---|---|---|---|
| Lines | 33 | 59 | 100 | 180 | 334 |
| Words | 204 | 376 | 675 | 1,237 | 2,276 |
| Estimated tokens | 369 | 690 | 1,256 | 2,316 | 4,247 |
| Headings | 4 | 6 | 10 | 16 | 27 |
| Code blocks | 0 | 0 | 2 | 4 | 8 |
All 590,000 indexed root CLAUDE.md files, by size
GitHub's own counts for filename:CLAUDE.md path:/, per size range. Our sample follows these proportions.
GitHub reports these counts as estimates. They cover non-fork public repositories in its code search index.
What do CLAUDE.md files actually contain?
Most CLAUDE.md files contain build and run commands, architecture notes, a project overview and code style rules; testing, git and security sections are far rarer. By heading, the most common are commands (64% of files), architecture (56%) and a project overview (50%). Code style follows at 50%, then structure (41%). Testing gets a heading in 33%, git and pull requests in 19%, gotchas in 15% and security in 12%. The median file has 10 headings; 1.5% have none at all. Switch the group to see how popular repos and each language differ.
Share of files with a heading about…
All 6,688 files with real content. Median length: 100 lines.
A file counts once per category if any heading matches. One heading can match more than one category ("Build and test commands" counts for both).
Topics mentioned anywhere in the file
Not only in headings: any mention in the text.
Claude Code features mentioned
How many files talk about the newer parts of Claude Code.
The 30 most common section headings
Level-2 headings, lowercased, with numbering and emoji stripped. The share is of files with real content.
Which commands and tools do the files name?
76% of files name at least one command-line tool in a shell block or inline code. JavaScript tooling leads: npm appears in 26% of all files and git in 16%. Among the 2,367 files that name a JavaScript package manager, 73% use npm, 20% pnpm, 7.7% bun and 3.7% yarn (a file can name more than one). Beyond running things, 39% mention a linter, 18% type checking and 13% give commit message rules.
Command-line tools named
From shell code blocks and inline code that starts with the tool's name.
Tech the files talk about
Any mention in the file text.
Do popular repositories write them differently?
Yes: files in popular repositories are longer, cover more topics and get edited more often. Using stars as a rough proxy for a project people depend on, files in repos with 100 or more stars (298 files) run longer, at a median of 138 lines against 95 for repos under 10 stars, and they cover more ground: a testing section in 49% (all files: 33%), code style in 63% (all: 50%). They are also edited more: a median of 4 commits against 2 overall. The flip side: they are more likely to pass Anthropic's 200-line target, with 33% over it against 22% of all files.
| Measure | 0 to 9 stars | 10 to 99 | 100 to 999 | 1,000+ | All |
|---|---|---|---|---|---|
| Files in sample | 5,428 | 962 | 257 | 41 | 6,688 |
| Median lines | 95 | 133 | 136 | 161 | 100 |
| Median words | 640 | 921 | 853 | 906 | 675 |
| Over 200 lines | 19% | 35% | 32% | 37% | 22% |
| Names a runnable command | 75% | 80% | 80% | 95% | 76% |
| Has a testing section | 30% | 44% | 46% | 71% | 33% |
| Has a code style section | 47% | 59% | 62% | 71% | 50% |
| Mentions secrets or .env | 37% | 33% | 28% | 20% | 36% |
| Keeps the /init header | 36% | 40% | 38% | 44% | 37% |
| Uses @imports | 3.8% | 4% | 3.5% | 4.9% | 3.8% |
| Repo also has AGENTS.md | 18% | 21% | 24% | 27% | 19% |
| Median commits to the file | 2 | 4 | 4 | 5 | 2 |
Files with real content only (pointer files left out), so shares here can differ slightly from whole-sample figures elsewhere on the page.
Is CLAUDE.md the same as AGENTS.md?
Same idea, different reader. Claude Code reads CLAUDE.md; Codex, Cursor, Copilot and others read AGENTS.md, and Anthropic's docs now say "Claude Code can read AGENTS.md as your project instructions". In practice many repos keep both: 21% of repos with a root CLAUDE.md also have an AGENTS.md, and 35% of repos with an AGENTS.md also have a CLAUDE.md. 3.4% of CLAUDE.md files are pointers ("See AGENTS.md" or a one-line @AGENTS.md), against 1.8% of AGENTS.md files pointing the other way. Next to the file, 32% of repos commit a .claude/ folder, most often for skills (14% of all repos) and shared settings (11%).
Other agent files in the same repo
Share of repos with a root CLAUDE.md that also have…
What is inside the .claude/ folder
32% of these repos commit a .claude/ folder. Share of all repos with each entry:
| Measure | CLAUDE.md | AGENTS.md |
|---|---|---|
| Files in sample | 6,967 | 2,484 |
| Median lines | 100 | 77 |
| Median words | 675 | 520 |
| Over 200 lines | 22% | 19% |
| Pointer files (hand off to another file) | 3.4% | 1.8% |
| Names a runnable command | 76% | 72% |
| Has a testing section | 33% | 41% |
| Mentions secrets or .env | 36% | 32% |
| Repo also has the other file | 21% | 35% |
| Median commits to the file | 2 | 2 |
Both samples were drawn the same way on the same day. Length rows exclude pointer and near-empty files.
How often are CLAUDE.md files updated?
A CLAUDE.md is only as good as its last edit. On the default branch, 38% of files have a single commit, 26% have two or three, 23% have four to ten and 13% have more than ten. The median file was last changed 133 days before we collected the data, and 36% were edited in the last 90 days. Almost all of these files are new: 81% were first committed in 2026.
When each file was first committed
Month of the first commit that touched the file, for files with up to 50 commits. Hover a bar for the share.
Based on 6,865 files whose full history we could read. October 2026 is a partial month (data collected on the 8th).
Do people shout at Claude? IMPORTANT, MUST and NEVER
Some files read like a terms-of-service page: IMPORTANT, MUST, NEVER in capitals. Fewer than you might think do it. 16% of files use at least one shouted word and only 3.6% use five or more. MUST is the most common (8.2% of files), then NEVER (6.6%), CRITICAL (4.7%), IMPORTANT (3.8%), ALWAYS (3.9%) and DO NOT (1.9%). Code blocks are not counted.
Anthropic's advice points the other way from volume: "The more specific and concise your instructions, the more consistently Claude follows them." A specific rule ("Use 2-space indentation") does more than a loud vague one.
How does your CLAUDE.md compare?
Paste your file (or drop it in) and see where it sits against the 6,688 public files. It is measured with the same rules as the study, in your browser. Nothing is uploaded.
A CLAUDE.md template built from the data
We took the sections that files in repos with 100+ stars use most and put them in the order they usually appear. Each heading below shows how many of those files have it. The template is short on purpose: fill in what is true for your repo, delete the rest, and you will land well under Anthropic's 200-line target.
- Project overview58% of 100+ star repos
- Commands (build, run, dev)72% of 100+ star repos
- Architecture61% of 100+ star repos
- Code style and conventions63% of 100+ star repos
- Testing49% of 100+ star repos
- Git, commits and PRs29% of 100+ star repos
- Rules and constraints39% of 100+ star repos
- Gotchas and troubleshooting22% of 100+ star repos
# CLAUDE.md <!-- Template from the State of CLAUDE.md study (locul.ai/research/state-of-claude-md). Replace every <...>, delete what does not apply, and delete this line. Aim to stay under 200 lines. --> ## Project overview <One or two sentences: what this repo is, who uses it, and the main stack.> Example: "REST API and admin dashboard for <product>. TypeScript, Next.js 15, Postgres via Prisma." ## Commands Swap in your own tools; keep the exact flags. ```bash pnpm install # install dependencies pnpm dev # dev server on http://localhost:3000 pnpm test # full test suite pnpm test -- path/to/file.test.ts # one test file pnpm lint # lint (must pass before a PR) pnpm typecheck # type check pnpm build # production build ``` ## Architecture - `<dir>/`: <what lives here and why> - `<dir>/`: <what lives here and why> - <How a request or job flows through the system, in one or two lines.> - <Where config and environment variables are read.> ## Code style - <The conventions a linter does not enforce: naming, file layout, error handling.> - <Preferred patterns, with one real example path to copy from: see `<path/to/good/example>`.> - <Anything Claude tends to get wrong here, stated as a concrete rule.> ## Testing - Run `pnpm test` before saying a change is done. Run a single file while iterating. - New code gets a test next to it in `<test dir or naming pattern>`. - <Test data, fixtures or services the tests need, and how to start them.> ## Git and pull requests - Branch from `<main branch>`; never push to it directly. - Commit messages: <format, e.g. Conventional Commits: feat:, fix:, chore:>. - <PR checklist: tests pass, lint clean, screenshots for UI changes.> ## Rules - Never commit secrets. Keys live in `<.env.local or secret manager>`, which is gitignored. - Do not edit `<generated or vendored paths>`; regenerate them with `<command>`. - Ask before adding a dependency or changing the database schema. ## Gotchas - <The non-obvious thing that costs a new contributor an hour.> - <Known flaky test, required service, or OS-specific step.>
# CLAUDE.md <!-- Template from the State of CLAUDE.md study (locul.ai/research/state-of-claude-md). Replace every <...>, delete what does not apply, and delete this line. Aim to stay under 200 lines. --> ## Project overview <One or two sentences: what this repo is, who uses it, and the main stack.> Example: "REST API and admin dashboard for <product>. TypeScript, Next.js 15, Postgres via Prisma." ## Commands Swap in your own tools; keep the exact flags. ```bash pnpm install # install dependencies pnpm dev # dev server on http://localhost:3000 pnpm test # full test suite pnpm test -- path/to/file.test.ts # one test file pnpm lint # lint (must pass before a PR) pnpm typecheck # type check pnpm build # production build ``` ## Architecture - `<dir>/`: <what lives here and why> - `<dir>/`: <what lives here and why> - <How a request or job flows through the system, in one or two lines.> - <Where config and environment variables are read.> ## Code style - <The conventions a linter does not enforce: naming, file layout, error handling.> - <Preferred patterns, with one real example path to copy from: see `<path/to/good/example>`.> - <Anything Claude tends to get wrong here, stated as a concrete rule.> ## Testing - Run `pnpm test` before saying a change is done. Run a single file while iterating. - New code gets a test next to it in `<test dir or naming pattern>`. - <Test data, fixtures or services the tests need, and how to start them.> ## Git and pull requests - Branch from `<main branch>`; never push to it directly. - Commit messages: <format, e.g. Conventional Commits: feat:, fix:, chore:>. - <PR checklist: tests pass, lint clean, screenshots for UI changes.> ## Rules - Never commit secrets. Keys live in `<.env.local or secret manager>`, which is gitignored. - Do not edit `<generated or vendored paths>`; regenerate them with `<command>`. - Ask before adding a dependency or changing the database schema. ## Gotchas - <The non-obvious thing that costs a new contributor an hour.> - <Known flaky test, required service, or OS-specific step.>
Real CLAUDE.md examples worth reading
The most-starred repositories in the sample whose CLAUDE.md has real content and no /init boilerplate. Each card shows the file's length when we measured it and the first sections it covers. Links open the current version on GitHub.
If you have not written one yet, Anthropic's three-minute video shows where the file lives and how /init drafts it. Then trim what /init writes: in our data, 37% of files never removed its opening line. To check your own file against this study, use the checker above; for a line-by-line audit, the CLAUDE.md Analyzer. If the context your agent needs lives in notes and docs rather than the repo, Locul can serve that to Claude Code as memory over MCP, which keeps CLAUDE.md for the repo-specific rules.
How was this measured?
What we collected
GitHub's code search reported about 586,432 files named CLAUDE.md at the root of public, non-fork repositories on 8 October 2026 (query filename:CLAUDE.md path:/). The search API returns at most 1,000 results per query, so we split the population into 14 file-size ranges, read GitHub's count for each, and took a quota from each range in proportion to its share. That keeps the sample's size distribution the same as the population's. We downloaded each file at the exact commit GitHub had indexed, then read repository metadata, the file's commit history on the default branch (up to 50 commits) and the presence of other agent files through the GitHub GraphQL API.
The result is 6,967 repositories, one root CLAUDE.md each. 6,688 have real content; the rest are pointer files (a file under 400 characters that hands off to AGENTS.md, README.md or an @import) or have fewer than 20 words. Length, section and tone numbers use the 6,688 files with content; sibling-file, folder and history numbers use all 6,967. A comparison sample of 2,484 root AGENTS.md files was drawn the same way.
How we measured
- Lines are counted as written, blank lines included. Tokens are estimated at 4 characters per token; exact counts vary by model.
- Sections come from Markdown headings outside code blocks, matched to 16 categories by keyword (for example, "Build and test commands" counts for both commands and testing).
- Commands are tools at the start of a line in a shell code block, or at the start of an inline code span with arguments (
`pnpm test`). - Mentions match anywhere in the file. Shouted words count capitalised IMPORTANT, MUST, NEVER, ALWAYS, CRITICAL and DO NOT outside code blocks.
- @imports are
@pathreferences to a file with an extension such as .md or .json, outside code blocks.
The checker on this page runs the same rules in JavaScript. We tested it against the Python pipeline on 1,000 random files from the sample, and every measured field matched on all 1,000.
What this does not tell you
- GitHub orders results by its own relevance, not at random, so within each size range the sample leans towards what GitHub ranks first. The search index also lags and skips some repositories.
- Only root-level
CLAUDE.mdfiles are counted. Files in subfolders,.claude/CLAUDE.md, personal~/.claude/CLAUDE.mdfiles and private repositories are not. - Keyword matching misses headings phrased in unusual ways, and 7.5% of files are written mostly in a non-Latin script, which the keyword rules mostly miss.
- Nothing here measures whether a file makes Claude better at its job. For that, see the ETH Zurich study below.
Earlier studies
| Study | Sample | Focus |
|---|---|---|
| Agent READMEs, Chatlatanagulchai et al., 2025 | 2,303 context files from 1,925 repos, including 922 CLAUDE.md | Content categories, words, edits |
| Agentic coding manifests, Chatlatanagulchai et al., 2025 | 253 CLAUDE.md files from 242 repos | Structure and content |
| Context engineering in open source, Mohsenimofidi et al., 2025 | 466 projects | Adoption of AGENTS.md |
| Evaluating AGENTS.md, Gloaguen et al., 2026 | Agent runs on SWE-bench Lite and 138 new tasks | Whether context files help agents |
| GitHub Blog, Nigh, 2025 | 2,500+ agents.md files | Practitioner lessons |
| This study, 2026 | 6,967 CLAUDE.md + 2,484 AGENTS.md, sampled by size | Length in lines, sections, commands, imports, sibling files, .claude/ folder, edit history |
Questions people ask about CLAUDE.md
How long is too long for a CLAUDE.md?
Anthropic's own line is the best answer we have: "target under 200 lines per CLAUDE.md file. Longer files consume more context and reduce adherence." In our sample the median file is 100 lines and 22% of files are over 200. Claude Code only refuses a file outright when it is over 4 MiB, so length is a quality problem long before it is a hard limit.
How many lines should a CLAUDE.md have?
Most public files land between 59 and 180 lines, with a median of 100. Files in repositories with 100 or more stars have a median of 138 lines. A file that lists the commands, the layout, the conventions a linter cannot enforce and the few rules that matter usually fits well under 200 lines; the template on this page is under 70.
What should go into a CLAUDE.md?
What Claude cannot work out from the code. In public files the most common headings cover commands (64% of files), architecture (56%), a project overview (50%) and code style (50%). Testing has its own heading in 33% and security in only 12%. Exact commands ("pnpm test -- path") beat descriptions ("we use Jest").
Is CLAUDE.md the same as AGENTS.md?
They do the same job for different tools. CLAUDE.md is the file Claude Code reads; AGENTS.md is the shared format used by Codex, Cursor, Copilot and others. Anthropic's docs say "Claude Code can read AGENTS.md as your project instructions". In our sample 21% of repos with a CLAUDE.md also have an AGENTS.md, and 3.4% of CLAUDE.md files are pointers that hand off to another file, usually AGENTS.md.
Is CLAUDE.md really useful?
For the right content, yes. A 2026 ETH Zurich study found that "context files are useful for specifying non-standard coding practices" but that "providing context files does not generally improve task success rates, while increasing inference cost by over 20% on average." The lesson for your file: keep what is specific to your repo (commands, odd conventions, traps) and cut the generic overview.
Does Claude always read CLAUDE.md?
Claude Code loads the CLAUDE.md files in the directory tree above where you start it at launch, and files in subdirectories when it works there. Anthropic is clear that "Claude treats them as context, not enforced configuration", so a rule in the file is a strong hint, not a guarantee. Shorter, specific rules are followed more reliably.
Where should CLAUDE.md live?
For a project, at the repository root as ./CLAUDE.md or ./.claude/CLAUDE.md, committed so the whole team shares it. Personal rules for every project go in ~/.claude/CLAUDE.md. This study only measures root-level CLAUDE.md files, the most common place.
What is the difference between CLAUDE.md and CLAUDE.local.md?
CLAUDE.md is shared through version control. CLAUDE.local.md holds your private notes for one project; Anthropic's docs say: "For private per-project preferences that shouldn't be checked into version control, create a CLAUDE.local.md at the project root." Because it is meant to stay out of git, public data cannot measure it. Its settings cousin leaks more often: 7% of repos in our sample commit .claude/settings.local.json, a file meant to stay personal.
Download or cite this study
The data is free to reuse under CC BY 4.0. Link back to this page so readers find the method.
Cite it
Khalid, J. (2026). The State of CLAUDE.md: what 6,967 public CLAUDE.md files contain. Data collected 8 October 2026. Locul. https://locul.ai/research/state-of-claude-md/
Download
- state-of-claude-md.csv one row per repository: 6,967 rows, 57 measured columns
- data.json every aggregate on this page, plus the AGENTS.md sample
- claude-md-template.md the template above
The CSV holds measurements only, not file contents. Each row links to the public file on GitHub.