Before a model's crawler reads your homepage it reads three small files, and it decides from them whether your site is worth reading at all. Most founders have never looked at these files. This is what they are, what to put in them, and how to check yours.
robots.txt: the door
At yourdomain.com/robots.txt. It tells crawlers what they may read. The trap: some hosts and some "block AI" plugins add rules that refuse the AI crawlers by name (GPTBot, ClaudeBot, Google-Extended and others). If you want to be in AI answers, those crawlers need to be allowed. Open the file and look for Disallow lines under those names. If your site is blocking them, you are invisible to those engines, whatever else you do.
The other thing that belongs in robots is the address of your sitemap, as a Sitemap: line.
sitemap.xml: the map
At yourdomain.com/sitemap.xml. A list of every page you want read, with the date it last changed. Crawlers use it to find pages that are not linked from the homepage, and the dates tell them what is new. If a page is not in the sitemap and not linked, it is often not read.
Check: does it exist, does it list your new pages, are the dates real. Submit it to Google Search Console once; Google will re-read it on its own after that.
llms.txt: the note to the models
At yourdomain.com/llms.txt. A newer convention: a plain text file that tells a model, in words, what the site is, who it is for, and where the important pages are. It is not read by every engine yet, but it costs nothing, and the engines that read it use it to decide which pages to fetch first.
What to put in it: one paragraph on what you do and for whom, then a short list of your most important pages with a one-line description each: pricing, the comparison pages, the documentation, the honest "who it is not for" page.
Schema: the facts, in a form a machine trusts
Schema is structured data inside a page, invisible to readers, that tells a machine what the page is and what facts it holds. Four kinds matter for AI visibility.
- Organization on the homepage: your name, logo, site, the same name you use everywhere. This is how a model knows the brand on this site is the brand named elsewhere.
- Article on posts and guides: title, date, author. Dates matter; recent pages win ties.
- FAQ on any page that answers questions: the question and the answer, exactly as on the page. Models lift these as ready-made question and answer pairs.
- Product or Offer wherever there is a price: the price, the currency, the plan name. A price in schema becomes a fact in an answer. A price in prose becomes an opinion.
Check with a schema validator; it will show what a machine sees.
Readability: can it read the words at all
One more check that is not a file. Some sites render everything with JavaScript and show a crawler an empty page. Fetch your page with a tool that does not run JavaScript, or view the page source, and look for your text. If it is not there, the model is not reading it, and that is the first thing to fix.
Five minutes, once a week
Robots allows the AI crawlers. Sitemap exists and lists the new pages. llms.txt exists and points at the pages that matter. Schema on the pages that carry facts. Text in the HTML. That is the whole check, and it is the technical half of AI visibility. ZwayRank's technical round runs it every week and creates the missing files on connected sites, but you can do it by hand in five minutes, and you should at least once.
The first scan is free: what ChatGPT, Gemini and Claude say about you, where Google ranks you, and what the crew would do first.
Start free for 7 days