Setting Up Web Crawlers for AI

Have you ever given ChatGPT or Gemini the address of a document made with ManualWorks, only to hear "I found the page but couldn't read its content"? Then it's time to check your web crawler settings.

Why can't it read the page?

ManualWorks sends a different kind of page depending on the User-Agent of the request.

The web viewer is the right choice when people read documents in a web browser. But a crawler that doesn't run JavaScript only sees an empty page when it gets the web viewer. This is one of the main reasons an AI finds the address but can't read the content.

There are two kinds of crawlers

AI services usually use two kinds.

ManualWorks treats a request as a crawler when its User-Agent contains bot, crawl, or spider. Crawlers with those strings are recognized without any settings, but crawlers without them, such as ChatGPT-User, Google-Agent, and MistralAI-Index, must be added as user patterns.

The problem described above is directly related to user-triggered fetchers, which are called when a user pastes an address into an AI.

Gemini doesn't fetch in real time

Gemini answers from search data that Google has already collected. So even when you give it a document address, no request reaches ManualWorks at that moment, and nothing shows up in the access log.

This means that even with the right crawler settings, Gemini can't read the latest content right away. What matters first is whether Google has indexed the document.

However, if you add .md to the end of the address, the request reaches ManualWorks right then. To have it read the current content without waiting for indexing, give it the Markdown address.

Update to the latest version

Starting with ManualWorks 6.0.24, the major AI crawlers come built in as system patterns. They are recognized without any settings, so if you are on 6.0.23 or earlier, updating is the simplest fix.


Crawler names keep growing and changing. Keeping up with the latest version is easier than managing the list yourself. Only when you need a crawler that isn't on the list, add it as a user pattern in the <Admin | Preferences | Web Crawler> menu.

In 6.0.23 and earlier, always enter user patterns in lowercase. Only the incoming User-Agent is converted to lowercase, not the saved patterns, so a pattern with uppercase letters never matches. Starting with 6.0.24, patterns are converted to lowercase when saved, so case doesn't matter.

Among the built-in names, Google's are used for different purposes. Google-Agent is used when an agent running on Google infrastructure accesses the web at a user's request, and GoogleAgent-URLContext is used when Gemini reads the content of a given address. Google-GeminiNotebook is used to fetch addresses that a user enters as sources in Gemini Notebook. GoogleOther also covers GoogleOther-Image and GoogleOther-Video.

headlesschrome is there to catch crawlers that don't identify themselves. Grok fetches pages with headless Chrome, so the browser signature is the only way to tell. However, the same signature also appears in automation built with Puppeteer or Playwright, monitoring tools, and build scripts. If your company captures or checks documents with headless Chrome, those tools are also treated as crawlers: they receive static HTML, and they get a 404 for documents with “Disable the access of Web Crawler.” selected.

Some fetchers that aren't AI crawlers but fetch pages at a user's request are recognized as well. They don't have bot in their names either, so they wouldn't be caught otherwise.

Google-Site-Verification belongs to the same group, but it's better not to add it. If you verify Search Console ownership with a meta tag, the static HTML for crawlers doesn't include that tag, so verification can fail.

Major AI services

Service

User-Agent

Needs setup in 6.0.23 and earlier

ChatGPT

GPTBot, OAI-SearchBot

No

ChatGPT

ChatGPT-User

Yes

Google Search, Gemini

Googlebot

No

Google agents

Google-Agent

Yes

Gemini URL reading

GoogleAgent-URLContext

Yes

Gemini Notebook

Google-GeminiNotebook

Yes

Other Google services

GoogleOther, GoogleOther-Image, GoogleOther-Video

Yes

Claude

ClaudeBot, Claude-SearchBot

No

Claude

Claude-User

Yes

Perplexity

PerplexityBot

No

Perplexity

Perplexity-User

Yes

Copilot

bingbot

No

Meta AI

meta-externalagent, meta-externalfetcher

Yes

Mistral

MistralAI-User, MistralAI-Index, MistralAI-Training

Yes

Grok

HeadlessChrome

Yes

Apple Intelligence

Applebot

No

Amazon

Amazonbot

No

ByteDance

Bytespider

No

Common Crawl

CCBot

No

robots.txt is a separate matter

Google-Extended and Applebot-Extended look like crawler names, but they are control tokens used only in robots.txt. They never appear in the User-Agent of an HTTP request, so adding them to web crawler patterns has no effect. To block or allow the use of crawled content for AI model training, set it in robots.txt.

ManualWorks serves the content you create as the robots type in the Web Page feature at /robots.txt.

What changes when a request is recognized as a crawler

When a request is recognized as a crawler, three things change together.

Checking that it works

On the Web Crawler screen, you can enter a User-Agent string and check right away whether it is recognized as a crawler.

From outside the server, check with the following command.

curl -A "ChatGPT-User" https://example.com/manual/chapter

If you get HTML that contains the content, it's set up correctly. If you get only JavaScript code without the content, the request isn't recognized as a crawler yet.

If it still can't read the page

If the AI still can't read the page after you've set up the crawlers, look at the address itself. The fetch features of external AI services generally require the following.

The usual setup is to put a web server in front and pass requests on to ManualWorks. Also check that the document is published publicly so it opens without logging in.

Finally

Each company sometimes changes its crawler names or adds new ones. The surest way to know which names are coming in now is to check the actual User-Agent in your access log.