CoreCited
Technical

Next.js SEO for AI search: what crawlers receive, and the metadata catch

Next.js SEO for AI crawlers: what they receive from server and client components, why streamed metadata lands in <body>, and the config that fixes it.

AlexFounder & Developer9 min read

Next.js SEO is mostly solved by the framework’s defaults: pages render on the server, metadata has a proper API, and sitemaps and robots files are built in. The gaps appear with AI crawlers, which do not run JavaScript, and with one newer behaviour, streaming metadata, that puts your title and description somewhere they may not look. This site runs on Next.js, so everything below was checked on a real deployment.

Key takeaways
1App Router pages are Server Components by default, so their content is in the HTML every crawler receives.
2Major AI crawlers do not run JavaScript. Anything that only appears after scripts run is invisible to them.
3Since Next.js 15.2, request-time metadata can be streamed into <body>. GPTBot and ClaudeBot are not on the list that gets it in <head>.
4Prerender where you can; where you cannot, add AI crawlers to htmlLimitedBots.

What crawlers receive from a Next.js page

In the App Router, “layouts and pages are Server Components” by default[4]. Even components marked 'use client' are prerendered to HTML on the first load, which Next.js uses “to immediately show a fast non-interactive preview of the route”. So the common fear that Next.js hides content from crawlers is mostly wrong. The exception is content fetched in the browser after load, in a useEffect or a client-side data library: that never reaches the server HTML.

That exception matters more now. Google renders JavaScript, but its own guidance says server rendering “is still a great idea” because “not all bots can run JavaScript”[9]. Vercel measured which ones: the major AI crawlers, from OpenAI, Anthropic, Meta, ByteDance and Perplexity, do not render JavaScript, while Gemini (through Googlebot) and AppleBot do[8].

Who sees what on a Next.js page
Rendering from Vercel's crawler study and Google's documentation.
Server-rendered contentContent loaded by client JavaScript
GooglebotYesYes, after rendering
GPTBot, OAI-SearchBotYesNo
ClaudeBotYesNo
PerplexityBotYesNo

The streaming metadata catch

Next.js 15.2 introduced streaming for generateMetadata. When metadata is resolved at request time, Next.js can send the page without waiting and “the resulting metadata tags are appended to the <body> tag”. Next.js says it verified this works for bots that run JavaScript, such as Googlebot. For “HTML-limited bots” it keeps the old behaviour and puts metadata in the head, detected by user agent[1].

The default list of HTML-limited bots covers link previewers and several search engines, such as facebookexternalhit, Bingbot and Twitterbot. It does not include GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot or PerplexityBot[3]. We checked what that means in practice on this site, which runs Next.js 15.5 on Vercel. Its dynamic /unsubscribe page served the <title> inside <body> to GPTBot, ClaudeBot, Googlebot and Chrome, and inside <head> to facebookexternalhit. Try a user agent:

Where does this crawler get your metadata?
For a page whose generateMetadata runs at request time, on Next.js 15.2 or later with the default settings.
<body>. Next.js streams the page first and appends the metadata to the body when it resolves. Next.js says it verified Googlebot reads it there; crawlers that do not run JavaScript may not.
What this does and does not prove
The tags are still in the HTML, just not in the head. No AI company has said its crawler ignores metadata in the body. But a parser that does not run JavaScript and reads the head for the title, description and canonical will not find them there. For pages you want cited, there is no reason to take the risk.

Fixing it

The best fix is to not need it: prerender. If generateMetadata only uses route params and build-time data, and the route is listed in generateStaticParams, the metadata is in the initial HTML for every crawler. Note that you “must always return an array” from generateStaticParams, even an empty one, or the route renders dynamically[5].

Where pages must render per request, set htmlLimitedBots. Your value replaces the default list[2], so keep the defaults and add the AI crawlers:

next.config.ts
// next.config.ts — the default list plus AI crawlers.
// Setting htmlLimitedBots replaces the default, so the default is included.
const nextConfig = {
  htmlLimitedBots: /[\w-]+-Google|Google-[\w-]+|Chrome-Lighthouse|Slurp|DuckDuckBot|baiduspider|yandex|sogou|bitlybot|tumblr|vkShare|quora link preview|redditbot|ia_archiver|Bingbot|BingPreview|applebot|facebookexternalhit|facebookcatalog|Twitterbot|LinkedInBot|Slackbot|Discordbot|WhatsApp|SkypeUriPreview|Yeti|googleweblight|GPTBot|OAI\-SearchBot|ChatGPT\-User|ClaudeBot|Claude\-User|Claude\-SearchBot|PerplexityBot|Perplexity\-User|CCBot/i,
};

export default nextConfig;

Next.js also documents htmlLimitedBots: /.*/ to switch streaming off entirely, with the warning that it “could lead to longer response times”[1].

Myth
Next.js sites are invisible to crawlers because React runs in the browser.
What is true
App Router pages render on the server by default. Only content fetched in the browser after load is missing.
Myth
If Googlebot sees it, AI crawlers see it.
What is true
Googlebot renders JavaScript. The major AI crawlers do not.
Myth
generateMetadata always puts tags in <head>.
What is true
Since 15.2, request-time metadata is streamed into <body> for bots not on the HTML-limited list.

The rest of Next.js SEO, briefly

  • Metadata. Export metadata or generateMetadata from each page, with a unique title, a description and alternates.canonical. It only works in Server Components, because metadata “must be resolved on the server”[1].
  • Sitemap. app/sitemap.ts generates the XML and is cached by default; split large sites with generateSitemaps[7].
  • robots.txt. app/robots.ts accepts separate rules per user agent, which is how you treat GPTBot and OAI-SearchBot differently[6]. The robots.txt generator writes the rules.
  • Share images. An opengraph-image.tsx next to a page generates its card at build time.
  • Structured data. Render JSON-LD in a server component so it is in the HTML. The schema generator builds it.
Test with curl, not with a browser. A browser shows you what JavaScript builds. An AI crawler gets what the server sends.
Next.js SEO checklist for AI search
0/7
Check a page as GPTBot
curl -s -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot" \
  https://example.com/page | grep -o "<title>.*</title>\|</head>"

If the title appears after </head>, metadata is being streamed. If nothing prints at all, the page may be redirecting, and one that redirects in circles is covered in fixing ERR_TOO_MANY_REDIRECTS. For which AI crawlers actually visit, see log file analysis for AI crawlers.

Questions people ask

Is Next.js good for SEO?

Yes, if pages are rendered on the server. In the App Router, pages and layouts are Server Components by default, so their content is in the HTML a crawler receives. The risks come from content that only appears after client-side JavaScript runs, and from metadata resolved at request time.

How do I do SEO in Next.js?

Export a metadata object or generateMetadata from each page for title, description, canonical and Open Graph; add app/sitemap.ts and app/robots.ts; prerender pages where you can with generateStaticParams; keep important content out of client-only components; and add an opengraph-image for share cards.

Can AI crawlers read Next.js sites?

They can read whatever is in the server HTML. Vercel found that the major AI crawlers, including OpenAI's and Anthropic's, do not execute JavaScript, so text that only appears after scripts run is invisible to them. Server-rendered and prerendered pages are fine.

What is streaming metadata in Next.js?

Since Next.js 15.2, when generateMetadata runs at request time, Next.js can send the page before the metadata is ready and append the tags to the body afterwards. Bots on its HTML-limited list get blocking metadata in the head instead. AI crawlers are not on the default list.

Do I need htmlLimitedBots?

Only if pages resolve metadata at request time and you want AI crawlers to get it in the head. Prerendered pages already have metadata in the head for everyone. If you set htmlLimitedBots, include the default list too, because your value replaces it.

Keep reading

Sources

[1]generateMetadata — Next.js documentation, updated 25 August 2026
[2]htmlLimitedBots — Next.js documentation, read 26 September 2026
[3]html-bots.ts (default HTML-limited bot list) — Next.js source, GitHub, read 26 September 2026
[4]Server and Client Components — Next.js documentation, updated 25 August 2026
[5]generateStaticParams — Next.js documentation, updated 25 August 2026
[6]robots.txt — Next.js documentation, read 26 September 2026
[7]sitemap.xml — Next.js documentation, updated 25 August 2026
[8]The rise of the AI crawler — Vercel, 17 December 2024
[9]Understand JavaScript SEO basics — Google Search Central, updated 4 March 2026

Find out where you actually stand

One real question, real AI engines, and the answer they gave — including who was named in it. No account, no card.