llms.txt best practices: facts, canonical URLs, what to leave out
An /llms.txt file serves as a lightweight, plain-text index hosted at your canonical root domain. It provides AI crawlers and automated retrieval tools with verified entity definitions, canonical resource links, and factual assertions backed by public HTML pages. This guide covers file formatting, HTTP serving requirements, pre-deployment validation, common maintenance pitfalls, and how to automate file generation from a single source of truth.
1. The spec: purpose and scope
The /llms.txt standard (proposed via llmstxt.org and tracked in the Answer.AI repository) defines a simple convention: a Markdown-formatted index placed at the website root to provide language models with a concise map of core resources and verified facts.
Core characteristics of the standard:
- Root location: Must be served directly at
https://example.com/llms.txtrather than within subdirectories. - Format: Clean Markdown without HTML markup, front matter, or script tags.
- Audience: Consumed by web crawlers and developer tooling that check for structured discovery files.
- Adoption status: An emerging convention; search engines and model providers treat it as an optional discovery aid rather than a guaranteed indexing trigger.
- Optional companion: An optional
/llms-full.txtfile can aggregate expanded documentation text, but/llms.txtremains the compact entry point.
When crawlers encounter complex client-rendered web applications, they may fail to extract key product descriptions before timing out. A lightweight text index gives automated fetchers an accessible fallback containing canonical links and core product assertions.
Public HTML pages remain the authoritative source of truth. The llms.txt file is an index referencing those pages. Any discrepancy between the two means the index is outdated.
Sources
- llmstxt.org — The proposed /llms.txt standard for plain-text fact indexes at the site root.
- Answer.AI llms-txt repository — Schema specifications, parsers, and reference implementations.
- Schema.org Organization — Ensuring entity attributes align between public HTML, structured data, and plain-text indexes.
- MDN: Content-Type — HTTP header compatibility for text/plain plain-text indexes.
2. Markdown structure and formatting
While the specification allows standard Markdown syntax, following a consistent structural schema ensures predictable parsing across different LLM ingestion tools:
# Brand Name
> One-sentence, verifiable description that matches the homepage definition.
## Official resources
- [Product](https://example.com/): Primary product overview.
- [Docs](https://example.com/docs): Technical documentation.
- [Pricing](https://example.com/pricing): Tier comparisons and billing terms.
- [Contact](https://example.com/contact): Support and sales inquiries.
## Verified facts
- Acme provides enterprise analytics in the US and EU — see https://example.com/about.
- SOC 2 Type II compliance report available for Enterprise tiers — see https://example.com/security.
## Optional
- [Blog](https://example.com/blog): Engineering updates and case studies.
- [Status](https://status.example.com/): Real-time system availability.
Formatting guidelines for reliable parsing:
- Single H1 header: Use the official brand name, matching your
<title>andOrganization.nameschema. - Blockquote definition: Place a single concise description immediately beneath the H1, aligned with your homepage meta description.
- Flat H2 sections: Group items under clear H2 headings without deep nesting.
- Explicit markdown links: Format entries as
- [Label](https://example.com/path): brief contextual hint.Use absolute HTTPS URLs on your canonical domain. - Traceable factual claims: Format verifiable statements with direct URLs:
- Statement — see https://example.com/target-page.Ensure the referenced page contains the supporting text. - Encoding: Use UTF-8 with standard LF line endings and no byte-order mark (BOM).
Omit visual tables, code blocks, images, or interactive scripts. The file should remain cleanly readable when converted to plain unformatted text.
3. Hosting and HTTP headers
Host the file exclusively on your primary canonical domain. If https://example.com is your canonical URL, configure https://www.example.com/llms.txt to issue an HTTP 301 redirect to the apex host.
- Serve over HTTPS with automatic redirects from HTTP.
- Ensure the endpoint allows anonymous GET requests without bot challenges, CAPTCHAs, or authentication walls.
- Return an exact file response at
/llms.txtwithout trailing slashes.
Content-Type configuration
Serve the file with text/plain; charset=utf-8 to maximize compatibility across automated clients and proxy caches:
Content-Type: text/plain; charset=utf-8
Content-Length: <bytes>
Cache-Control: public, max-age=3600, stale-while-revalidate=86400
X-Content-Type-Options: nosniff
Deploy the file as a static asset alongside your web application build to ensure immediate updates during releases.
4. What belongs in the file
Every line in llms.txt should reference publicly verifiable information that is already live on your website:
- Entity identity: Official brand name and a clear one-sentence summary of what the product does.
- Core navigation links (4–8 URLs): Product overviews, documentation, pricing tables, contact points, and system status pages.
- Key factual assertions (3–6 bullet points): Regional availability, compliance certifications, primary platform integrations, or architecture properties, each linked to a supporting page.
Keep the complete file under 5 KB. A concise, link-dense document is easier for crawlers to fetch, cache, and process than an exhaustive promotional page.
5. Maintenance anti-patterns
Common mistakes that degrade the accuracy and usefulness of an llms.txt index:
Hardcoded pricing figures
Copying exact price numbers into llms.txt often creates discrepancies when pricing plans are updated. Instead of writing raw figures, link directly to your pricing page:
- Pricing: https://example.com/pricing — Current tier details and billing options.
Other common pitfalls
- Unsubstantiated marketing claims: Avoid vague claims like "#1 rated platform" unless backed by a cited public benchmark.
- Private or authenticated routes: Do not include staging links, internal admin URLs, or pages behind authentication walls.
- Unlinked keyword lists: Avoid lists of raw buzzwords without supporting URLs.
- Sitemap duplication: Do not dump hundreds of URLs into the file; use
sitemap.xmlfor exhaustive crawling. - Regional fragmentation: Maintain a single canonical
/llms.txtfile referencing localized subdirectories rather than creating multiple regional root files.
Validation checks
Incorporate basic validation into your deployment pipeline:
- Verify the file starts with a single H1 and contains valid H2 headings.
- Confirm all URLs use absolute HTTPS formats and return HTTP 200 responses.
- Check that every factual assertion references an active source URL.
- Ensure total file size remains below 5 KB.
6. Generation and CI validation
Rather than editing llms.txt manually, generate it as part of your automated build pipeline using verified metadata from your CMS or product repository:
- Extract metadata: Load brand details, canonical links, and approved factual assertions from a structured configuration file.
- Verify links: Validate that all referenced URLs resolve with HTTP 200 status codes during the build step.
- Render template: Output clean Markdown using string interpolation:
# ${brand.name} > ${brand.tagline} ## Official resources ${resources.map(r => `- [${r.label}](${r.url}): ${r.hint}`).join("\n")} ## Verified facts ${facts.map(f => `- ${f.claim} — see ${f.source}`).join("\n")} ## Optional ${optional.map(r => `- [${r.label}](${r.url}): ${r.hint}`).join("\n")} - Emit static artifact: Write
dist/llms.txtwith UTF-8 encoding and deploy atomically with your site assets.
Deployment verification
After deployment, test the live endpoint using curl:
# Inspect headers and content
curl -i https://example.com/llms.txt
# Verify all referenced URLs return HTTP 200
grep -oE 'https://[^) ]+' /tmp/llms.txt | xargs -n1 curl -s -o /dev/null -w '%{url_effective}: %{http_code}\n'
7. How CiteAura manages llms.txt
CiteAura workspace projects generate reviewed llms.txt assets directly from your verified brand entity registry. In Execution → Assets, teams can select approved resource links, verify supporting citations, and export production-ready Markdown files for deployment. CiteAura provides the validated artifact for your engineering team to publish to your hosting infrastructure.
Generate valid, pre-verified llms.txt markdown assets from your brand entity library in CiteAura.
Start free trialScope and realistic expectations
An accurate /llms.txt file provides clean reference material for retrieval tools, but it does not guarantee automatic AI citations or search ranking changes. Real-world model visibility depends on overall crawl accessibility, content extractability, and search relevance. Treat llms.txt as a clear, maintainable index that complements well-structured HTML markup.
FAQ: publishing llms.txt
What are the official llms.txt syntax conventions and best practices?
According to llmstxt.org conventions: serve the file at /llms.txt as text/plain; charset=utf-8 directly at the site root, begin with an H1 brand title, include a concise blockquote summary, and structure canonical documentation links with clean Markdown unordered lists.
What should not go in llms.txt?
Do not include private internal endpoints, staging environments, login-protected content, or unsubstantiated claims. Avoid hardcoding raw pricing figures; link to your live pricing page instead.
Does llms.txt guarantee AI mentions?
No. The format provides structured discovery for AI crawlers, but models evaluate many factors including crawl accessibility, HTML structure, and prompt context.
Where should llms.txt be published?
Serve the file at https://example.com/llms.txt as text/plain; charset=utf-8 directly from your canonical root domain.
How do I validate llms.txt before deploy?
Check that the file contains clean Markdown headings, absolute HTTPS links, reachable source URLs for all factual claims, and valid UTF-8 encoding.