Discovery file — Engineering guide

llms.txt best practices: facts, canonical URLs, what to leave out

An /llms.txt file serves as a lightweight, plain-text index hosted at your canonical root domain. It provides AI crawlers and automated retrieval tools with verified entity definitions, canonical resource links, and factual assertions backed by public HTML pages. This guide covers file formatting, HTTP serving requirements, pre-deployment validation, common maintenance pitfalls, and how to automate file generation from a single source of truth.

Updated 14 September 2026 · 9 min read · Technical

Editorial Transparency: Fact-checked against search engine specifications and empirical crawling tests. Our Editorial Standards →

1. The spec: purpose and scope

The /llms.txt standard (proposed via llmstxt.org and tracked in the Answer.AI repository) defines a simple convention: a Markdown-formatted index placed at the website root to provide language models with a concise map of core resources and verified facts.

Core characteristics of the standard:

When crawlers encounter complex client-rendered web applications, they may fail to extract key product descriptions before timing out. A lightweight text index gives automated fetchers an accessible fallback containing canonical links and core product assertions.

Public HTML pages remain the authoritative source of truth. The llms.txt file is an index referencing those pages. Any discrepancy between the two means the index is outdated.

Sources

2. Markdown structure and formatting

While the specification allows standard Markdown syntax, following a consistent structural schema ensures predictable parsing across different LLM ingestion tools:

# Brand Name

> One-sentence, verifiable description that matches the homepage definition.

## Official resources

- [Product](https://example.com/): Primary product overview.
- [Docs](https://example.com/docs): Technical documentation.
- [Pricing](https://example.com/pricing): Tier comparisons and billing terms.
- [Contact](https://example.com/contact): Support and sales inquiries.

## Verified facts

- Acme provides enterprise analytics in the US and EU — see https://example.com/about.
- SOC 2 Type II compliance report available for Enterprise tiers — see https://example.com/security.

## Optional

- [Blog](https://example.com/blog): Engineering updates and case studies.
- [Status](https://status.example.com/): Real-time system availability.

Formatting guidelines for reliable parsing:

Omit visual tables, code blocks, images, or interactive scripts. The file should remain cleanly readable when converted to plain unformatted text.

3. Hosting and HTTP headers

Host the file exclusively on your primary canonical domain. If https://example.com is your canonical URL, configure https://www.example.com/llms.txt to issue an HTTP 301 redirect to the apex host.

Content-Type configuration

Serve the file with text/plain; charset=utf-8 to maximize compatibility across automated clients and proxy caches:

Content-Type: text/plain; charset=utf-8
Content-Length: <bytes>
Cache-Control: public, max-age=3600, stale-while-revalidate=86400
X-Content-Type-Options: nosniff

Deploy the file as a static asset alongside your web application build to ensure immediate updates during releases.

4. What belongs in the file

Every line in llms.txt should reference publicly verifiable information that is already live on your website:

Keep the complete file under 5 KB. A concise, link-dense document is easier for crawlers to fetch, cache, and process than an exhaustive promotional page.

5. Maintenance anti-patterns

Common mistakes that degrade the accuracy and usefulness of an llms.txt index:

Hardcoded pricing figures

Copying exact price numbers into llms.txt often creates discrepancies when pricing plans are updated. Instead of writing raw figures, link directly to your pricing page:

- Pricing: https://example.com/pricing — Current tier details and billing options.

Other common pitfalls

Validation checks

Incorporate basic validation into your deployment pipeline:

  1. Verify the file starts with a single H1 and contains valid H2 headings.
  2. Confirm all URLs use absolute HTTPS formats and return HTTP 200 responses.
  3. Check that every factual assertion references an active source URL.
  4. Ensure total file size remains below 5 KB.

6. Generation and CI validation

Rather than editing llms.txt manually, generate it as part of your automated build pipeline using verified metadata from your CMS or product repository:

  1. Extract metadata: Load brand details, canonical links, and approved factual assertions from a structured configuration file.
  2. Verify links: Validate that all referenced URLs resolve with HTTP 200 status codes during the build step.
  3. Render template: Output clean Markdown using string interpolation:
    # ${brand.name}
    
    > ${brand.tagline}
    
    ## Official resources
    ${resources.map(r => `- [${r.label}](${r.url}): ${r.hint}`).join("\n")}
    
    ## Verified facts
    ${facts.map(f => `- ${f.claim} — see ${f.source}`).join("\n")}
    
    ## Optional
    ${optional.map(r => `- [${r.label}](${r.url}): ${r.hint}`).join("\n")}
    
  4. Emit static artifact: Write dist/llms.txt with UTF-8 encoding and deploy atomically with your site assets.

Deployment verification

After deployment, test the live endpoint using curl:

# Inspect headers and content
curl -i https://example.com/llms.txt

# Verify all referenced URLs return HTTP 200
grep -oE 'https://[^) ]+' /tmp/llms.txt | xargs -n1 curl -s -o /dev/null -w '%{url_effective}: %{http_code}\n'

7. How CiteAura manages llms.txt

CiteAura workspace projects generate reviewed llms.txt assets directly from your verified brand entity registry. In Execution → Assets, teams can select approved resource links, verify supporting citations, and export production-ready Markdown files for deployment. CiteAura provides the validated artifact for your engineering team to publish to your hosting infrastructure.

Generate valid, pre-verified llms.txt markdown assets from your brand entity library in CiteAura.

Start free trial

Scope and realistic expectations

An accurate /llms.txt file provides clean reference material for retrieval tools, but it does not guarantee automatic AI citations or search ranking changes. Real-world model visibility depends on overall crawl accessibility, content extractability, and search relevance. Treat llms.txt as a clear, maintainable index that complements well-structured HTML markup.

FAQ: publishing llms.txt

What are the official llms.txt syntax conventions and best practices?

According to llmstxt.org conventions: serve the file at /llms.txt as text/plain; charset=utf-8 directly at the site root, begin with an H1 brand title, include a concise blockquote summary, and structure canonical documentation links with clean Markdown unordered lists.

What should not go in llms.txt?

Do not include private internal endpoints, staging environments, login-protected content, or unsubstantiated claims. Avoid hardcoding raw pricing figures; link to your live pricing page instead.

Does llms.txt guarantee AI mentions?

No. The format provides structured discovery for AI crawlers, but models evaluate many factors including crawl accessibility, HTML structure, and prompt context.

Where should llms.txt be published?

Serve the file at https://example.com/llms.txt as text/plain; charset=utf-8 directly from your canonical root domain.

How do I validate llms.txt before deploy?

Check that the file contains clean Markdown headings, absolute HTTPS links, reachable source URLs for all factual claims, and valid UTF-8 encoding.

About the Author: CiteAura Research Team

Search Systems Engineers & Evaluation Analysts

The CiteAura Research Team is the engineering working group behind the open-source CiteAura GEO engine, empirical crawler audits, and verifiable citation benchmarks across generative AI search systems.

Explore Team & Research → Editorial Standards →