Skip to content
Technical

AI crawlers read the HTML you send, not the page your visitors see

Most AI crawlers do not run JavaScript. Why a client-rendered pricing page can be blank to ChatGPT, Claude and Perplexity, and how to test yours.

Obility Editorial · · 4 min read

Most AI crawlers stop at the first response

When a person opens a modern website, the server often sends a thin HTML shell and the browser runs JavaScript to fetch and draw the actual content: the plans, the prices, the feature table. A crawler that does not run that JavaScript never sees the finished page. It reads the shell and leaves.

The best public evidence on which AI crawlers fall into that group comes from Vercel and the consultancy MERJ, who published an analysis of crawler traffic across Vercel’s network in December 2024. They checked nextjs.org and two job board sites built on different stacks. None of the major AI crawlers they measured rendered JavaScript: OpenAI’s OAI-SearchBot, ChatGPT-User and GPTBot, Anthropic’s ClaudeBot, PerplexityBot, and the crawlers from Meta and ByteDance. OpenAI’s and Anthropic’s crawlers did download JavaScript files, about 11.5 and 24 percent of their requests respectively, but did not execute them. Vercel added one useful caveat: content already present in the initial response, such as JSON data embedded in the page, may still be read.

Google renders, because AI Overviews use its search index

Google is the exception, and the reason is plumbing rather than a special AI bot. Google’s documentation on AI features says a page needs to be indexed and eligible to show with a snippet to appear as a supporting link in AI Overviews or AI Mode, with no additional technical requirements. Indexing goes through Googlebot, which Google says works in three phases: crawl, render and index. Pages wait in a rendering queue, sometimes for seconds and sometimes longer, until a headless Chromium runs the JavaScript. Vercel’s data showed Gemini using the same Googlebot infrastructure.

Some guides circulating now describe Google-Extended as a separate crawler that renders JavaScript. Google’s crawler documentation says Google-Extended has no user agent of its own; it is a robots.txt token that controls whether crawled content is used for Gemini training and grounding, and it does not affect inclusion in Search. The practical consequence is the trap: a client-rendered site can look healthy in Google’s AI answers while being close to empty for ChatGPT, Claude and Perplexity. A team that only spot-checks Google will never notice. Google’s own JavaScript guide still recommends server-side rendering or pre-rendering, partly because not every bot can run JavaScript.

The facts buyers ask about are often the ones drawn late

The parts of a site most likely to depend on JavaScript tend to be the parts buyers ask assistants about. Pricing tables that switch between monthly and annual billing and pull plans from an API. Integration directories with filters. Comparison tables inside tabs. FAQ answers that load when someone clicks. Review widgets supplied by a third party. Lists of features behind a “load more” button.

Take an illustrative case. A software company’s pricing page returns a heading, a navigation bar and a loading spinner in its initial HTML; the plans arrive from an API a moment later. A buyer asks ChatGPT whether the company has a free plan. If the engine’s crawler ever fetched that page, it got no plan information from it, so the answer leans on whatever else it can find: a review site’s pricing summary from last year, a forum thread, or nothing. The company wrote the correct answer and published it where only browsers could read it. Links have the same problem: navigation built with JavaScript click handlers instead of real links with an href gives a non-rendering crawler no path to the pages behind them.

Test the response, not the browser

The check takes ten minutes. Request your important pages without a browser, with curl or any tool that shows the raw response, or open View Source rather than the inspector, which shows the page after JavaScript has run. Then search that raw HTML for the exact sentences you need an assistant to repeat: your plan names, your prices, the integrations you support, the one line that explains what you do. If a fact is missing from the raw response, assume ChatGPT, Claude and Perplexity cannot see it on that page. Turning JavaScript off in a browser gives a quicker, rougher view of the same thing.

Your server logs add the second half: which pages AI crawlers actually request and what status codes they get back. Vercel found that about 35 percent of fetches by OpenAI’s and Anthropic’s crawlers hit 404 pages, many of them outdated asset URLs, so redirects and a current sitemap matter more than they do for Googlebot. Obility’s Agent Analytics treats this evidence carefully: a crawler request shows a page was requested, not that it was used in an answer, and a user agent string is a useful signal rather than verified identity.

Fix the server response, then expect a lag

The fix is to put the content in the first response. Google recommends server-side rendering, static rendering or hydration, and describes dynamic rendering, where crawlers get a separate pre-rendered version, as a workaround it no longer recommends. Most modern frameworks support server or static rendering per page, so this rarely means a rebuild. Start with pricing, product, comparison and documentation pages. Client-side code is still fine for enhancements such as chat widgets, view counters and interactive filters, as long as the facts exist without them.

Know where this stops. The Vercel analysis is nearly two years old, and none of OpenAI, Anthropic or Perplexity documents whether its crawlers render pages, so behaviour can change without notice. Test your own pages rather than relying on anyone’s summary, including this one. Readable HTML makes your content eligible to be used; it does not make an engine cite you. And Vercel observed assistants answering questions about fresh documentation without any matching fetch in its logs, relying on cached or training data instead. A fixed page only helps once it is fetched again, and no engine publishes when that will be.

Share this article

Put the workflow into practice.

Bring your questions and content to an Obility demo.

Book a demo