Half of B2B software buyers now start their search inside an AI assistant instead of Google. That means a large share of the people evaluating your company will never see your homepage. They will see a two-line summary of it, written by a model that read your site in a few seconds. This post explains what those models actually fetch, in what order, and the five files that decide whether they describe you correctly.
First, there is no single “AI crawler”
Each company runs several bots with different jobs, and blocking the wrong one is a common mistake. A training crawler like GPTBot improves a future model. A live-fetch agent such as ChatGPT-User retrieves a page because someone asked a question right now. Blocking the second can erase a brand from active answers. contently
The bots that matter in 2026, grouped by job:
| Company | Training | Search index | Live fetch (user asked) |
|---|---|---|---|
| OpenAI | GPTBot | OAI-SearchBot | ChatGPT-User |
| Anthropic | ClaudeBot | Claude-SearchBot | Claude-User |
| Google-Extended (a robots.txt signal, not a crawler) | Googlebot | Googlebot | |
| Perplexity | PerplexityBot | PerplexityBot | Perplexity-User |
The practical reading: if you want to be cited in answers, the search and live-fetch bots must be allowed. Training bots are a separate business decision. GPTBot is only responsible for grabbing data to train the model. What really determines whether you will appear in the ChatGPT answers is another crawler called OAI-SearchBot. The same split applies to Anthropic and Google. tenten
And these visits are not a rounding error anymore. In Cloudflare’s May 2026 crawler data, Googlebot was the single largest AI-adjacent bot at 27.26 percent of requests, followed by GPTBot at 11.48 percent and ClaudeBot at 9.73 percent. geotoolbox
1. robots.txt: the door
This is the first file every well-behaved bot reads. It is a request, not a wall; a bot that ignores it needs a firewall rule, not a text file. A sensible default for a company that wants to be found and is comfortable with training:
User-agent: OAI-SearchBot User-agent: ChatGPT-User User-agent: Claude-SearchBot User-agent: Claude-User User-agent: PerplexityBot Allow: / User-agent: GPTBot User-agent: ClaudeBot User-agent: Google-Extended Allow: / Sitemap: https://yourdomain.com/sitemap.xml
If you would rather opt out of training and keep citations, change the second block to Disallow: /. Either way, the Sitemap: line stays. It is the one place every bot is guaranteed to look for it.
2. sitemap.xml: the map
A sitemap does for a bot what a table of contents does for a reader. It lists every page you want indexed, with a lastmod date for each. Live-fetch bots rarely crawl your whole site; they fetch one or two pages that seem relevant. The search-index bots do crawl, and they use lastmod to decide what to revisit. Keep the dates honest. A sitemap where every page claims to have changed today is quietly ignored.
3. JSON-LD: the facts, in a form a machine trusts
Models are good at reading prose and bad at being certain about it. Structured data removes the guessing. A small Organization block in your site layout states your name, founding year, address, and official profiles once, in a format that Google, Bing, and every model built on their indexes already parse.
{
"@context": "https://schema.org",
"@type": "Organization",
"name": "Northline Software",
"url": "https://northlinesoftware.com",
"foundingDate": "2017",
"address": { "@type": "PostalAddress", "addressLocality": "Austin", "addressRegion": "TX", "addressCountry": "US" },
"sameAs": ["https://github.com/northline-software", "https://www.linkedin.com/company/northline-software"]
}
On individual pages, Service and FAQPage do the most work. FAQ schema in particular mirrors the exact question a person types into an assistant, which is why answer engines quote FAQ blocks more than any other section. One rule: never put a question in JSON-LD that is not visible on the page. Search engines treat that as spam, and so should you.
4. llms.txt: the honest one-page brief
llms.txt is a plain Markdown file at your site root: what your company is, a one-line summary of each important page, and links. It was proposed in 2024 as a “here is the short version” for models with limited context.
Here is the part most guides skip: as of this writing, none of OpenAI, Anthropic, or Google has confirmed that their bots read it. Some agencies sell it as a control file. It is not. It cannot block anything and it does not guarantee a citation. What it does cost is twenty minutes, and it gives any tool that does read it (several coding assistants and smaller answer engines do) a clean summary you wrote instead of one it guessed. Publish it, keep it short, keep it current, and do not expect miracles from it.
5. The page itself: the first 60 words
After all the files, the bot reads the page. It does not scroll. It weights the title tag, the H1, and the opening paragraph heavily, and most of them handle JavaScript poorly or not at all. So:
- State the answer in the first 60 words. Explanation comes after.
- Make every H2 a question or a claim that reads correctly on its own.
- Render your main content as HTML on the server. Content that appears only after a client-side script runs is invisible to most of these bots.
- Use the same entity names everywhere: website, GitHub, LinkedIn, Clutch, npm. Consistency is how a model learns that five profiles are one company.
The checklist
- robots.txt allows the search and live-fetch bots, names your sitemap, and reflects a deliberate decision on training.
- sitemap.xml is complete with honest
lastmoddates. OrganizationJSON-LD sits in the layout;ServiceandFAQPagesit on the pages that earn them.- llms.txt exists, is short, and matches the site.
- Every key page answers its question in the opening paragraph, in server-rendered HTML.
Then test it. Once a quarter, ask ChatGPT, Claude, Gemini, and Perplexity the three questions you most want to be the answer to. Record whether you appear and what they say. Fix the pages that got you wrong. The bots are reading anyway. The only choice is whether they read something you wrote on purpose.

