How to write an llms.txt, and what it will and won't do

An llms.txt is a short Markdown file that tells a language model what your site is and where the important pages are. It is quick to write. Whether AI crawlers use it is a separate and much less certain question.

Published 2 October 2026 · 5 min read

llms.txt is a proposal, published at llmstxt.org in 2024, for a single Markdown file at the root of a website that gives a language model a concise map of the site: what it is, and which pages matter. The idea is that a model working with a limited context window should not have to wade through navigation, scripts and markup to find out what your product does.

This post covers the format, a worked example, how to write each part well, and then the part most guides skip: what the file will and will not do.

The format

The file lives at /llms.txt and is written in Markdown. The proposal defines its structure in this order:

  1. An H1 with the name of the site or project. This is the only required part.
  2. A blockquote with a short summary: the key information someone needs to understand the rest of the file.
  3. Optionally, more paragraphs or lists with detail. No headings in this part.
  4. Optionally, any number of sections headed with H2, each containing a list of links. Each item is a Markdown link, optionally followed by a colon and a short note about the page.

One section name has a special meaning. A section headed ## Optional contains links that can be skipped when a shorter context is needed, so it is the place for secondary pages.

The proposal also suggests that pages offer a clean Markdown version at the same URL with .md added. That part is optional, and plenty of sites publish an llms.txt without it.

A worked example

Here is a complete file for an invented product, a bookkeeping tool called Tallybook:

# Tallybook

> Tallybook turns the CSV file your bank exports into a categorised
> spreadsheet an accountant will accept. It is a web app for freelancers
> in the EU, free for up to 200 transactions a month.

Tallybook does not connect to bank accounts; you upload an export.
It reads CSV and OFX files and exports CSV and XLSX. It does not file
taxes or produce invoices.

## Docs

- [Getting started](https://tallybook.example/docs/start.md): upload a file, review the categories, export
- [Supported banks and formats](https://tallybook.example/docs/formats.md): which exports work and which do not
- [Categories](https://tallybook.example/docs/categories.md): the default categories and how to change them

## Pricing

- [Pricing](https://tallybook.example/pricing): the free tier's limits and the paid plan

## Optional

- [Changelog](https://tallybook.example/changelog): dated list of changes
- [About](https://tallybook.example/about): who builds Tallybook
An invented example. The .example domain is reserved and resolves nowhere.

How to write each part

The summary is the most important line. Write it statement-first, so that it still makes sense if a model quotes only that sentence: what the product is, who it is for and the one fact that distinguishes it. Avoid the stock vocabulary that fills so many landing pages; “seamless all-in-one platform” tells a model as little as it tells a person. The post on stock phrases has a list.

Put your limits in the detail section. What the product does not do is exactly the kind of fact that gets lost when a model summarises a marketing page. Saying it plainly makes it more likely that an answer about your product is accurate rather than flattering, which is better for you in the long run.

Link to the pages that answer real questions. How it works, pricing, supported platforms, limits, a FAQ. Not every page in the sitemap: the file is supposed to be a curated map, and a list of two hundred URLs is just a sitemap in a different format.

Write a note after each link. The note is what lets a model decide whether a page is worth reading without fetching it. “Docs” is a useless note; “which file formats work and which do not” is a useful one.

Keep it current. A file that describes last year's product is worse than no file, because it is confidently wrong. Put a date in it, or keep a short changelog section, and update it when the product changes.

Serve it as plain text at the root of your domain, and open it in a browser afterwards to check it is actually reachable and not swallowed by a single-page app's catch-all route. If you want a starting point, the free llms.txt generator drafts one in this format from your homepage, sitemap and up to twelve of your pages. Its output is a draft to edit, not a finished file. IntentGrid's own file is at /llms.txt, if you want to see one with a product-status paragraph and a section of stated limits.

What it will do

  • Give any tool or person who fetches it a short, accurate description of your site in a format a model reads easily. That includes developers who paste a project's llms.txt into an AI coding assistant, and tools that look for the file deliberately.
  • Make you write a clear, quotable description of your own product, which usually improves your homepage and your meta descriptions too.
  • Cost very little: a few hundred words and a link list.

What it will not do

This is the part to be clear-eyed about. llms.txt is a proposal, not a standard adopted by search engines or AI companies. Whether the crawlers behind AI assistants and answer engines fetch it, and whether anything they fetch from it affects their answers, is uncertain and largely undocumented. Publishing one is not guaranteed to change anything.

  • It will not improve your search rankings. Search engines rank pages, and this file is not a ranking signal.
  • It will not make an AI assistant cite you. Being cited depends on whether your pages are useful, accessible and trusted, not on a summary file.
  • It will not replace good HTML. If your pages render their text only in the browser, a crawler that does not run JavaScript may still see an empty shell, with or without an llms.txt.
  • It will not control who crawls your site. That is a different file.

Before relying on it for anything, check the current documentation of the specific AI products you care about, rather than taking any blog post's word for it, this one included.

llms.txt is not robots.txt

The two are easy to confuse because both sit at the root and both mention AI. They do different jobs.

robots.txt tells crawlers which paths they may fetch. Most AI companies document the user agents their crawlers use, such as GPTBot and OAI-SearchBot for OpenAI, ClaudeBot for Anthropic and PerplexityBot for Perplexity, and Google uses a separate robots.txt token, Google-Extended, to let sites opt out of some AI uses of their content. Several companies distinguish crawlers that collect training data from crawlers that fetch pages to answer a user's question, so blocking one may not block the other. robots.txt is a request that well-behaved crawlers honour, not an enforcement mechanism.

llms.txt allows and blocks nothing. It is a description, offered to whoever chooses to read it. Decide what you want crawled in robots.txt, and use llms.txt to describe what is there.

For the wider picture of how answer engines read a site, and how that differs from classic SEO, see does website personalization hurt SEO?, which covers GEO, generative engine optimization, in its own section.