Skip to content

feat(tools): make the llms.txt generator write context, not a link list - #118

Merged
harshit-epyc merged 3 commits into
mainfrom
feat/llms-txt-richer-context
Sep 9, 2026
Merged

harshit-epyc merged 3 commits into
mainfrom
feat/llms-txt-richer-context

Conversation

@harshit-epyc

Copy link
Copy Markdown
Collaborator

What

The generator produced a correct llms.txt and a thin one — a name, a one-line summary, and a bullet per page in whatever order the model emitted them. An assistant reading it learned which URLs exist, which is what the sitemap already told it. This makes the file worth reading.

The five changes

A company overview. The model returns two or three factual sentences, rendered between the blockquote and the first heading — the slot the convention reserves for free-form context. It is the part an assistant reads before it opens a single link.

The pages that matter, first. urlScore is lifted out of rankUrls (lib/crawl/sitemap.ts) and exported, so the ranking that decides what to crawl also decides what to list first. Sections render About → Services → Products → Pricing → Work → Contact → unrecognised → Optional, rather than in model-emission order: a reader under a context limit reads top-down and may stop early.

Descriptions that say something. The prompt asks for what a reader finds on that page — offerings, audience, named facts — and bans filler openers and title restatement. Both enforced: a description that normalises to its own title, or runs under five words, is dropped.

Figures checked against the site. Two new guards:

  • The evidence quote now has to actually be in the page text (verbatim after normalising case/punctuation/whitespace, falling back to 80% of its words). Before, any non-empty string satisfied the guard — which made evidence a formality a model could satisfy by writing a sentence it liked the sound of.
  • Every number in a description, in the summary, and in each overview sentence has to appear in the source text, or that unit is dropped. A hallucinated "300+ clients" in a file a customer publishes is the failure with the longest tail.

Archive furniture left out. Paginated archives, tag/category indexes, author pages, on-site search, feeds and trailing-slash duplicates are excluded before the model call. Excluded, not skipped — an archive URL is not evidence that a site fails to describe itself, so counting it in the finding would inflate the one number this tool exists to report.

Numbers on screen now add up

The response carries read, pages, described, skipped, excluded, and read = described + skipped + excluded. The result window shows all four; the gap section names the excluded ones so nothing goes missing between two counts.

Shape

Post-processing is split into pure buildDoc / selectPages, so every guard above is testable without a network call — the same split render.ts already had. New generate.test.ts (11 cases) pins each rejection; render.test.ts gains overview and section-ordering cases.

Checks

pnpm test (196 passed) · tsc --noEmit · pnpm lint (no new warnings) · opennextjs-cloudflare build.

Docs: docs/llms-txt-architecture.md §2.2 rewritten (three guards → five, plus §2.2.1 on ordering and exclusions) and the §4 response contract updated.

Not done

No date signal exists in the corpus, so "outdated" is handled as junk-and-duplicate URL removal rather than freshness — flagging that explicitly rather than implying we detect stale pages. Worth a follow-up if lastmod from the sitemap turns out to be reliable enough to carry.

🤖 Generated with Claude Code

https://claude.ai/code/session_01DAz5uKk9XGWWEP8kimVW3s

The file the generator produced was a correct llms.txt and a thin one: a
name, a one-line summary, and a bullet per page in whatever order the model
emitted them. An assistant reading it learned which URLs exist, which is
what the sitemap already told it.

Five changes, four of them deterministic:

- **A company overview.** The model now returns two or three factual
  sentences that render between the blockquote and the first heading, where
  the convention reserves free-form context. This is the part an assistant
  reads before it opens a single link.

- **The pages that matter, first.** `urlScore` is lifted out of `rankUrls`
  (lib/crawl/sitemap.ts) and exported, so the ranking that decides what to
  crawl now also decides what to list first — which page is worth reading
  first and which is worth listing first are the same question. Sections are
  ordered About → Services → Products → Pricing → Work → Contact →
  unrecognised → Optional, rather than by the order the model happened to
  emit them, because a reader under a context limit reads top-down.

- **Descriptions that say something.** The prompt asks for what a reader
  finds on that page — offerings, audience, named facts — bans filler
  openers and title restatement, and both are enforced: a description that
  normalises to its own title, or runs under five words, is dropped.

- **Figures checked against the site.** The evidence quote now has to
  actually appear in the page text (verbatim after normalisation, or 80% of
  its words), where before any non-empty string satisfied the guard. And
  every number in a description, in the summary and in each overview
  sentence has to appear in the source, or that unit is dropped. A
  hallucinated "300+ clients" in a file a customer publishes is the failure
  with the longest tail.

- **Archive furniture left out.** Paginated archives, tag and category
  indexes, author pages, on-site search, feeds and trailing-slash duplicates
  are excluded before the model call. They are excluded, not skipped: an
  archive URL is not evidence that a site fails to describe itself, so
  counting it in the finding would inflate the one number this tool reports.
  The response now carries `read`, `pages`, `described`, `skipped` and
  `excluded`, and `read = described + skipped + excluded` holds on screen.

The post-processing is split into `buildDoc`/`selectPages` so every guard is
testable without a network call — the same split render.ts already had.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DAz5uKk9XGWWEP8kimVW3s
@xhawk-ai

xhawk-ai Bot commented Sep 9, 2026

Copy link
Copy Markdown

⚠️ XHawk Review paused — out of credits

Your team has run out of review credits, so I couldn't finish this review.
Top up to resume: → https://app.xhawk.ai/billing

harshit-epyc and others added 2 commits September 9, 2026 17:50
Against a real 20-page site the model's reply was cut at ~10,500 characters,
mid-array, and `JSON.parse` failed with "Expected ',' or ']' after array
element". That reads as a model that cannot write JSON, but it is the
`max_tokens` ceiling: the reply is one document covering every page, and
adding the overview plus a per-page evidence quote pushed it past 3000
tokens. Every tier in the chain truncates at the same place, so the fallback
cannot recover it and the visitor gets a 503.

Two halves:

- The budget goes to 6000, about 2x the observed reply.
- `evidence` is capped at 12 words. It is verification-only and never
  rendered, so a full sentence per page is output we pay for in the ceiling
  and then throw away.

Named in the architecture doc, with the upgrade path if it ever recurs:
salvage the complete array elements out of the cut reply rather than raising
the number again — a partial file beats an error.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DAz5uKk9XGWWEP8kimVW3s
Five gaps against what a useful llms.txt has to do, on top of the overview
and ordering already in this branch.

**Categorisation.** Documentation and Support join the section order, which
is now About, Products, Services, Pricing, Work, Documentation, Support,
Contact, anything the model named itself, then Optional. The list is
exported once from render.ts and interpolated into the prompt, so the
sections we ask for are exactly the ones the renderer can order.

**Curation.** `Optional` is capped at five links. A content-heavy site comes
back mostly blog posts, and twelve of them under one heading turns a map of
the company into a feed. The pages are rank-ordered already, so the cap
keeps the best few; the rest count as excluded, not skipped, because they
were describable and we chose not to list them. No other section is capped —
trimming About or Services would cut the pages the file exists to point at.

**Repetition.** A description identical to one already used is dropped and
its page joins the finding, which is the honest place for it: two pages that
describe identically do not distinguish themselves. An overview sentence
that only restates the blockquote above it is dropped too — pure restatement
only, since a sentence that contains the summary and goes on is elaboration,
which is the whole job of the overview.

**Stable information.** The prompt asks for facts that will still be true in
a year over awards, campaigns, funding rounds and follower counts. No guard
behind it, deliberately: marketing language is the site's own voice, not
model invention, and a superlative filter would drop a customer's real copy
and inflate the finding with it. Guards are for what the model made up.

**Parseable structure.** Every model-written string now goes through
`plain()` before rendering — leading list and heading markers, [text](url)
pairs, emphasis and backticks are stripped. The file's value is that a
machine can read its hierarchy without guessing, and a description opening
"- " reads as a nested list item.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DAz5uKk9XGWWEP8kimVW3s
@harshit-epyc
harshit-epyc merged commit d62aed6 into main Sep 9, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant