Repository navigation
feat(tools): make the llms.txt generator write context, not a link list - #118
Merged
Merged
Conversation
The file the generator produced was a correct llms.txt and a thin one: a name, a one-line summary, and a bullet per page in whatever order the model emitted them. An assistant reading it learned which URLs exist, which is what the sitemap already told it. Five changes, four of them deterministic: - **A company overview.** The model now returns two or three factual sentences that render between the blockquote and the first heading, where the convention reserves free-form context. This is the part an assistant reads before it opens a single link. - **The pages that matter, first.** `urlScore` is lifted out of `rankUrls` (lib/crawl/sitemap.ts) and exported, so the ranking that decides what to crawl now also decides what to list first — which page is worth reading first and which is worth listing first are the same question. Sections are ordered About → Services → Products → Pricing → Work → Contact → unrecognised → Optional, rather than by the order the model happened to emit them, because a reader under a context limit reads top-down. - **Descriptions that say something.** The prompt asks for what a reader finds on that page — offerings, audience, named facts — bans filler openers and title restatement, and both are enforced: a description that normalises to its own title, or runs under five words, is dropped. - **Figures checked against the site.** The evidence quote now has to actually appear in the page text (verbatim after normalisation, or 80% of its words), where before any non-empty string satisfied the guard. And every number in a description, in the summary and in each overview sentence has to appear in the source, or that unit is dropped. A hallucinated "300+ clients" in a file a customer publishes is the failure with the longest tail. - **Archive furniture left out.** Paginated archives, tag and category indexes, author pages, on-site search, feeds and trailing-slash duplicates are excluded before the model call. They are excluded, not skipped: an archive URL is not evidence that a site fails to describe itself, so counting it in the finding would inflate the one number this tool reports. The response now carries `read`, `pages`, `described`, `skipped` and `excluded`, and `read = described + skipped + excluded` holds on screen. The post-processing is split into `buildDoc`/`selectPages` so every guard is testable without a network call — the same split render.ts already had. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DAz5uKk9XGWWEP8kimVW3s
|
Your team has run out of review credits, so I couldn't finish this review. |
Against a real 20-page site the model's reply was cut at ~10,500 characters, mid-array, and `JSON.parse` failed with "Expected ',' or ']' after array element". That reads as a model that cannot write JSON, but it is the `max_tokens` ceiling: the reply is one document covering every page, and adding the overview plus a per-page evidence quote pushed it past 3000 tokens. Every tier in the chain truncates at the same place, so the fallback cannot recover it and the visitor gets a 503. Two halves: - The budget goes to 6000, about 2x the observed reply. - `evidence` is capped at 12 words. It is verification-only and never rendered, so a full sentence per page is output we pay for in the ceiling and then throw away. Named in the architecture doc, with the upgrade path if it ever recurs: salvage the complete array elements out of the cut reply rather than raising the number again — a partial file beats an error. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DAz5uKk9XGWWEP8kimVW3s
Five gaps against what a useful llms.txt has to do, on top of the overview and ordering already in this branch. **Categorisation.** Documentation and Support join the section order, which is now About, Products, Services, Pricing, Work, Documentation, Support, Contact, anything the model named itself, then Optional. The list is exported once from render.ts and interpolated into the prompt, so the sections we ask for are exactly the ones the renderer can order. **Curation.** `Optional` is capped at five links. A content-heavy site comes back mostly blog posts, and twelve of them under one heading turns a map of the company into a feed. The pages are rank-ordered already, so the cap keeps the best few; the rest count as excluded, not skipped, because they were describable and we chose not to list them. No other section is capped — trimming About or Services would cut the pages the file exists to point at. **Repetition.** A description identical to one already used is dropped and its page joins the finding, which is the honest place for it: two pages that describe identically do not distinguish themselves. An overview sentence that only restates the blockquote above it is dropped too — pure restatement only, since a sentence that contains the summary and goes on is elaboration, which is the whole job of the overview. **Stable information.** The prompt asks for facts that will still be true in a year over awards, campaigns, funding rounds and follower counts. No guard behind it, deliberately: marketing language is the site's own voice, not model invention, and a superlative filter would drop a customer's real copy and inflate the finding with it. Guards are for what the model made up. **Parseable structure.** Every model-written string now goes through `plain()` before rendering — leading list and heading markers, [text](url) pairs, emphasis and backticks are stripped. The file's value is that a machine can read its hierarchy without guessing, and a description opening "- " reads as a nested list item. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DAz5uKk9XGWWEP8kimVW3s
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
The generator produced a correct llms.txt and a thin one — a name, a one-line summary, and a bullet per page in whatever order the model emitted them. An assistant reading it learned which URLs exist, which is what the sitemap already told it. This makes the file worth reading.
The five changes
A company overview. The model returns two or three factual sentences, rendered between the blockquote and the first heading — the slot the convention reserves for free-form context. It is the part an assistant reads before it opens a single link.
The pages that matter, first.
urlScoreis lifted out ofrankUrls(lib/crawl/sitemap.ts) and exported, so the ranking that decides what to crawl also decides what to list first. Sections render About → Services → Products → Pricing → Work → Contact → unrecognised →Optional, rather than in model-emission order: a reader under a context limit reads top-down and may stop early.Descriptions that say something. The prompt asks for what a reader finds on that page — offerings, audience, named facts — and bans filler openers and title restatement. Both enforced: a description that normalises to its own title, or runs under five words, is dropped.
Figures checked against the site. Two new guards:
evidencequote now has to actually be in the page text (verbatim after normalising case/punctuation/whitespace, falling back to 80% of its words). Before, any non-empty string satisfied the guard — which madeevidencea formality a model could satisfy by writing a sentence it liked the sound of.Archive furniture left out. Paginated archives, tag/category indexes, author pages, on-site search, feeds and trailing-slash duplicates are excluded before the model call. Excluded, not skipped — an archive URL is not evidence that a site fails to describe itself, so counting it in the finding would inflate the one number this tool exists to report.
Numbers on screen now add up
The response carries
read,pages,described,skipped,excluded, andread = described + skipped + excluded. The result window shows all four; the gap section names the excluded ones so nothing goes missing between two counts.Shape
Post-processing is split into pure
buildDoc/selectPages, so every guard above is testable without a network call — the same splitrender.tsalready had. Newgenerate.test.ts(11 cases) pins each rejection;render.test.tsgains overview and section-ordering cases.Checks
pnpm test(196 passed) ·tsc --noEmit·pnpm lint(no new warnings) ·opennextjs-cloudflare build.Docs:
docs/llms-txt-architecture.md§2.2 rewritten (three guards → five, plus §2.2.1 on ordering and exclusions) and the §4 response contract updated.Not done
No date signal exists in the corpus, so "outdated" is handled as junk-and-duplicate URL removal rather than freshness — flagging that explicitly rather than implying we detect stale pages. Worth a follow-up if
lastmodfrom the sitemap turns out to be reliable enough to carry.🤖 Generated with Claude Code
https://claude.ai/code/session_01DAz5uKk9XGWWEP8kimVW3s