Have you ever given ChatGPT or Gemini the address of a document made with ManualWorks, only to hear "I found the page but couldn't read its content"? Then it's time to check your web crawler settings.
Why can't it read the page?
ManualWorks sends a different kind of page depending on the User-Agent of the request.
If it decides the request comes from a crawler, it sends static HTML that contains the content, with no JavaScript. It also includes the canonical URL and Open Graph information.
Otherwise, it sends the web viewer, which loads the content with JavaScript.
The web viewer is the right choice when people read documents in a web browser. But a crawler that doesn't run JavaScript only sees an empty page when it gets the web viewer. This is one of the main reasons an AI finds the address but can't read the content.
There are two kinds of crawlers
AI services usually use two kinds.
Indexing crawlers. They explore the web in advance to build training data or search indexes. Their names often contain
bot, such asGPTBot,OAI-SearchBot,Googlebot, andClaudeBot.User-triggered fetchers. They fetch a page when a user asks, "Read this address." Their names often contain
-User, such asChatGPT-UserandClaude-User, but there are exceptions likeGoogle-Agent.
ManualWorks treats a request as a crawler when its User-Agent contains bot, crawl, or spider. Crawlers with those strings are recognized without any settings, but crawlers without them, such as ChatGPT-User, Google-Agent, and MistralAI-Index, must be added as user patterns.
The problem described above is directly related to user-triggered fetchers, which are called when a user pastes an address into an AI.
Gemini doesn't fetch in real time
Gemini answers from search data that Google has already collected. So even when you give it a document address, no request reaches ManualWorks at that moment, and nothing shows up in the access log.
This means that even with the right crawler settings, Gemini can't read the latest content right away. What matters first is whether Google has indexed the document.
However, if you add .md to the end of the address, the request reaches ManualWorks right then. To have it read the current content without waiting for indexing, give it the Markdown address.
Update to the latest version
Starting with ManualWorks 6.0.24, the major AI crawlers come built in as system patterns. They are recognized without any settings, so if you are on 6.0.23 or earlier, updating is the simplest fix.
Crawler names keep growing and changing. Keeping up with the latest version is easier than managing the list yourself. Only when you need a crawler that isn't on the list, add it as a user pattern in the <Admin | Preferences | Web Crawler> menu.
In 6.0.23 and earlier, always enter user patterns in lowercase. Only the incoming User-Agent is converted to lowercase, not the saved patterns, so a pattern with uppercase letters never matches. Starting with 6.0.24, patterns are converted to lowercase when saved, so case doesn't matter.
Among the built-in names, Google's are used for different purposes. Google-Agent is used when an agent running on Google infrastructure accesses the web at a user's request, and GoogleAgent-URLContext is used when Gemini reads the content of a given address. Google-GeminiNotebook is used to fetch addresses that a user enters as sources in Gemini Notebook. GoogleOther also covers GoogleOther-Image and GoogleOther-Video.
headlesschrome is there to catch crawlers that don't identify themselves. Grok fetches pages with headless Chrome, so the browser signature is the only way to tell. However, the same signature also appears in automation built with Puppeteer or Playwright, monitoring tools, and build scripts. If your company captures or checks documents with headless Chrome, those tools are also treated as crawlers: they receive static HTML, and they get a 404 for documents with “Disable the access of Web Crawler.” selected.
Some fetchers that aren't AI crawlers but fetch pages at a user's request are recognized as well. They don't have bot in their names either, so they wouldn't be caught otherwise.
google-inspectiontoolis used when Search Console inspects indexing status.google-read-aloudis used to read a page aloud.google-pinpointis used to fetch addresses that a user sets as sources in Pinpoint.googlemessagesis used to create previews of addresses shared in chats.feedfetcher-googleis used to fetch RSS or Atom feeds.
Google-Site-Verification belongs to the same group, but it's better not to add it. If you verify Search Console ownership with a meta tag, the static HTML for crawlers doesn't include that tag, so verification can fail.
Major AI services
Service | User-Agent | Needs setup in 6.0.23 and earlier |
|---|---|---|
ChatGPT | GPTBot, OAI-SearchBot | No |
ChatGPT | ChatGPT-User | Yes |
Google Search, Gemini | Googlebot | No |
Google agents | Google-Agent | Yes |
Gemini URL reading | GoogleAgent-URLContext | Yes |
Gemini Notebook | Google-GeminiNotebook | Yes |
Other Google services | GoogleOther, GoogleOther-Image, GoogleOther-Video | Yes |
Claude | ClaudeBot, Claude-SearchBot | No |
Claude | Claude-User | Yes |
Perplexity | PerplexityBot | No |
Perplexity | Perplexity-User | Yes |
Copilot | bingbot | No |
Meta AI | meta-externalagent, meta-externalfetcher | Yes |
Mistral | MistralAI-User, MistralAI-Index, MistralAI-Training | Yes |
Grok | HeadlessChrome | Yes |
Apple Intelligence | Applebot | No |
Amazon | Amazonbot | No |
ByteDance | Bytespider | No |
Common Crawl | CCBot | No |
robots.txt is a separate matter
Google-Extended and Applebot-Extended look like crawler names, but they are control tokens used only in robots.txt. They never appear in the User-Agent of an HTTP request, so adding them to web crawler patterns has no effect. To block or allow the use of crawled content for AI model training, set it in robots.txt.
ManualWorks serves the content you create as the robots type in the Web Page feature at /robots.txt.
What changes when a request is recognized as a crawler
When a request is recognized as a crawler, three things change together.
The crawler receives static HTML that contains the content.
Documents with “Disable the access of Web Crawler.” selected in their settings return a 404. You can block documents you don't want to expose by selecting this option.
Visit statistics count the request as a crawler, so it isn't included in human visits.
Checking that it works
On the Web Crawler screen, you can enter a User-Agent string and check right away whether it is recognized as a crawler.
From outside the server, check with the following command.
curl -A "ChatGPT-User" https://example.com/manual/chapter
If you get HTML that contains the content, it's set up correctly. If you get only JavaScript code without the content, the request isn't recognized as a crawler yet.
If it still can't read the page
If the AI still can't read the page after you've set up the crawlers, look at the address itself. The fetch features of external AI services generally require the following.
A domain name. Many refuse IP addresses.
HTTPS. Many refuse unencrypted http.
A standard port. Exposing ManualWorks' default port, 1975, as is won't get through.
The usual setup is to put a web server in front and pass requests on to ManualWorks. Also check that the document is published publicly so it opens without logging in.
Finally
Each company sometimes changes its crawler names or adds new ones. The surest way to know which names are coming in now is to check the actual User-Agent in your access log.