ON THE RECORD · NO. 16 · THE CRAFT

Getting AI to read your website

What we ship on every build so the machines can read the site, what Google says you don't need, and where the truth moved this year.

Getting AI to read your website
"OAI-SearchBot, rendered faithfully," 1566
GIUSEPPE ARCIMBOLDO · THE LIBRARIAN · SKOKLOSTER CASTLE, SWEDEN

Every build we ship carries a machine-file inventory. A robots.txt that admits the AI crawlers by name, an llms.txt that regenerates itself on publish, the sitemap, the feeds left on because aggregators and AI crawlers still read them, schema built from real stored fields, and a security.txt for the scanner that always comes. Each one costs minutes. This note is what we’ve learned shipping them, on our own site and on client builds this quarter.

The robots.txt work is mostly bookkeeping, with one trap. The bot list grew, and the bots are different animals. Training crawlers, search crawlers, and the user agents that fetch a page live when someone asks an assistant about you are separate visitors with separate names, and admitting one class is a different policy call from admitting another. We allow the search and user bots everywhere, and leave training admission as the owner’s recorded decision. Two field lessons. Hosts can quietly write blocks for you, a privacy toggle flipped in a dashboard can close the door while you’re optimizing the welcome mat, so read the robots.txt your host actually serves. And when you add a named group for a new bot, copy your existing Disallow rules into it, because a crawler obeys the most specific group that matches it and ignores the rest. A new group with no Disallows just opened your admin paths to that bot.

On llms.txt, the honest report is that Google told everyone not to bother, and Google is right about Google. No AI system currently fetches the file, the server logs confirm it, and Google’s own guidance groups it with markdown page variants and AI-specific rewriting as things its AI surfaces don’t need. We ship it anyway, with eyes open. The module costs nothing to maintain, the spec is alive in the tooling world around the assistants, and a bet priced in minutes is worth holding. What it is not is the strategy. The strategy is being readable at all, server-rendered pages that work without JavaScript, facts stated in prose a model can lift, claims attached to sources. The sites that parse cleanly for the oldest crawler parse cleanly for the newest one.

The pattern we built on a client engagement this quarter is a page written for the machine reader. One URL that states what the company is, what it does, and the numbers that matter, each with its source, in plain declarative prose. It sits in the sitemap and stays out of the navigation, canonical and indexable, with one binding rule carried over from ordinary SEO. It must never create a duplicate-content problem for the pages humans read. For Google it’s neutral. It exists for the assistants whose user agents fetch pages live mid-conversation, because when that fetch happens, you want it landing on the page you wrote for exactly that reader.

The identity layer is schema, and it only works where the facts exist. Every post on a build carries a real byline wired to a real bio page, and the markup underneath, Article pointing to Person pointing to Organization, is how a machine resolves who wrote something and whether the name checks out. Machines assess credibility the way a skeptical editor would. A named human with a role, a bio, and a live profile that matches gives them something to verify, and “by the marketing team” gives them nothing. The markup is ten minutes of module code. The verifiable people behind it are the actual asset.

FAQs come with the honest caveat attached. Google stopped showing FAQ rich results for most sites in 2023 and retired them entirely this year, so the markup no longer buys a prettier search listing, and anyone selling FAQ schema as a rankings move is selling last decade. The format is what earns its keep now. A real question as the heading, the answer leading in its first sentence, 40 to 80 words, numbers only with a source. That’s the shape assistants lift whole into answers. On our builds one registry renders the visible FAQ and its schema together so the two can’t drift apart, and the writing rule does the heavy lifting. Write the answer for a human in a hurry and you’ve written it for the machine as well.

And one correction to our own standard, which is what a versioned standard is for. We used to treat Bing Webmaster registration as the ChatGPT submission, because ChatGPT’s search rode Bing’s index. Researchers watching ChatGPT’s own retrieval streams this summer found OpenAI’s in-house index now does the work, with Bing practically absent from the consumer path. Bing registration still earns its click, Copilot rides that index and the setup takes a minute from Search Console. The ChatGPT door, though, is OpenAI’s own crawler reading your actual site. Which means the submission is the one you control entirely, a robots.txt that lets OAI-SearchBot in and a site worth fetching.

The way to know any of this is working costs nothing. Read your access logs. The AI crawlers announce themselves by name, GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, and the user agents that fetch live during conversations show up the same way. Crawlers appearing in the logs is the first indicator on the whole chain, ahead of impressions, ahead of citations, ahead of any tool that charges for the answer. Five minutes in the logs tells you whether the door you opened is being used, and a fetch arriving thirty seconds after you asked an assistant about your own company is the whole system demonstrating itself.

None of this earns a citation. The record other people write about you does that, and the measurement problem that follows is its own discipline. This is the floor underneath both, and the floor costs an afternoon.

You can’t be quoted if you can’t be read.

Katie Clark signature
SIGNED & DATED · SEP 28, 2026 · WORK NO. 16
EXHIBITED SINCE SEP 2026 · REVISED OCT 2026
EVERY POST IS A SIGNED WORK
Get the next signed work first.
One piece a month. No filler, no reheated takes.
MORE FROM THE GALLERY
VIEW THE COLLECTION →
THANKS FOR VISITING. TIME FOR FIKA. THE SWEETSHOP IS OPEN

Discover more from JXT Collective

Subscribe now to keep reading and get access to the full archive.

Continue reading