AI爬虫抓取验证工具

模拟 GPTBot、ClaudeBot、CCBot、Kimi、Bytespider 等 36 款主流 AI 爬虫UA访问 · robots.txt / Meta robots / HTTP 状态三级检测 · 覆盖中美主要 AI 厂商

返回 Favicon & ICO 在线生成器 您还可以使用 蜘蛛大全 HTTP状态码大全 全国DNS大全 一键提交外链工具

内置 36 款主流 AI 爬虫(美国厂商 24 款 / 中国厂商 11 款),按高 / 中 / 低三档优先级并发检测,下方清单可查看每款爬虫的完整 UA 与屏蔽方法。

快速体验: ico5.net example.com

支持带或不带 http/https 前缀;服务端以各爬虫的真实 UA 发起 GET 请求,只读取页面前 160KB,单 IP 每分钟限 6 次。

🤖
正在以 GPTBot 身份访问…

内置 AI 爬虫清单(36 款,2026-09 核查版)

地区标识:中国厂商 美国厂商(其他地区不标注); 优先级: = 主流 AI 搜索/聊天产品、基础模型训练或高抓取量; = 有明确用途但流量或资料完整度相对较低; = 特定场景补充。 带 非官方 标注的爬虫,其 UA 来自日志观测或第三方记录,厂商尚未正式确认。蜘蛛名称可点击跳转蜘蛛大全对应详情页。

优先级 名称 厂商 说明 完整 User-Agent robots.txt 屏蔽方法
高优先级
GPTBot OpenAI OpenAI GPT 模型训练数据抓取,放行与否直接影响内容是否进入 GPT 语料。 Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot
User-agent: GPTBot
Disallow: /
OAI-SearchBot OpenAI ChatGPT Search 搜索索引抓取(非训练用途),放行有助于内容被 ChatGPT 搜索引用。 Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot
User-agent: OAI-SearchBot
Disallow: /
ChatGPT-User用户触发型 OpenAI 用户在 ChatGPT 中触发实时联网回答时的访问代理。 Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot
User-agent: ChatGPT-User
Disallow: /
ClaudeBot Anthropic Anthropic Claude 模型训练数据抓取。 Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ClaudeBot/1.0; +claudebot@anthropic.com)
User-agent: ClaudeBot
Disallow: /
Claude-SearchBot Anthropic Anthropic 搜索索引抓取,用于 Claude 搜索结果与内容引用。 Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Claude-SearchBot/1.0; +https://www.anthropic.com)
User-agent: Claude-SearchBot
Disallow: /
Claude-User用户触发型 Anthropic 用户在 Claude 中触发实时联网时的访问代理。 Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Claude-User/1.0; +Claude-User@anthropic.com)
User-agent: Claude-User
Disallow: /
Googlebot Google Google 搜索主爬虫,同时是 AI Overviews、AI Mode 与 Gemini 实时引用的数据来源。 Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)
User-agent: Googlebot
Disallow: /
Google-ExtendedToken 型 · 不发请求 Google 控制内容能否用于 Gemini / Vertex AI 训练的 robots.txt Token,不发起独立请求,声明即生效。 Mozilla/5.0 (compatible; Google-Extended)
User-agent: Google-Extended
Disallow: /
Google-GeminiNotebook用户触发型 Google Gemini Notebook(原 NotebookLM)用户添加来源 URL 时的抓取器。 Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/137.0.0.0 Safari/537.36 (compatible; Google-GeminiNotebook; +https://developers.google.com/crawling/docs/crawlers-fetchers/google-gemininotebook)
User-agent: Google-GeminiNotebook
Disallow: /
Google-Agent(桌面)用户触发型 Google Google 托管 Agent(如 Project Mariner)桌面端访问代理。 Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko; compatible; Google-Agent; +https://developers.google.com/crawling/docs/crawlers-fetchers/google-agent) Chrome/W.X.Y.Z Safari/537.36
User-agent: Google-Agent
Disallow: /
Google-Agent(移动)用户触发型 Google Google 托管 Agent 移动端访问代理。 Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/W.X.Y.Z Mobile Safari/537.36 (compatible; Google-Agent; +https://developers.google.com/crawling/docs/crawlers-fetchers/google-agent)
User-agent: Google-Agent
Disallow: /
PerplexityBot Perplexity Perplexity 答案引擎的索引抓取,用于搜索结果与内容引用。 Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)
User-agent: PerplexityBot
Disallow: /
Perplexity-User用户触发型 Perplexity 用户在 Perplexity 触发实时检索时的访问代理。 Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user)
User-agent: Perplexity-User
Disallow: /
CCBot Common Crawl Common Crawl 开源网页存档爬虫,众多大模型预训练语料的上游数据源,屏蔽它等于退出公共语料。 CCBot/2.0 (https://commoncrawl.org/faq/)
User-agent: CCBot
Disallow: /
bingbot Microsoft 同时驱动 Bing 搜索与 Copilot AI 回答的基础爬虫,GEO 检测不可缺。 Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/116.0.1938.76 Safari/537.36
User-agent: bingbot
Disallow: /
KimiBot 月之暗面 Kimi 月之暗面(Moonshot AI)Kimi 模型训练数据抓取,遵守 robots.txt。 Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; KimiBot/1.0; +https://www.kimi.com/policies/kimi-crawlers
User-agent: KimiBot
Disallow: /
Kimi-SearchBot 月之暗面 Kimi Kimi AI 搜索索引抓取。 Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; Kimi-SearchBot/1.0; +https://www.kimi.com/policies/kimi-crawlers
User-agent: Kimi-SearchBot
Disallow: /
Kimi-User用户触发型 月之暗面 Kimi 用户向 Kimi 提问需要实时检索时按需抓取。 Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; Kimi-User/1.0; +https://www.kimi.com/policies/kimi-crawlers
User-agent: Kimi-User
Disallow: /
DeepSeekBot非官方 DeepSeek DeepSeek AI 内容抓取,官方尚无爬虫文档,UA 来自访问日志观测。 Mozilla/5.0 (compatible; DeepSeekBot/1.0; +https://www.deepseek.com/bot)
User-agent: DeepSeekBot
Disallow: /
Bytespider 字节跳动 字节跳动(豆包等)LLM 训练数据抓取,实测抓取量大,robots.txt 合规性存在争议。 Mozilla/5.0 (Linux; Android 5.0) AppleWebKit/537.36 (KHTML, like Gecko) Mobile Safari/537.36 (compatible; Bytespider; spider-feedback@bytedance.com)
User-agent: Bytespider
Disallow: /
Baiduspider 百度 百度搜索主爬虫,百度 AI 搜索与文心一言联网回答的基础数据来源。 Mozilla/5.0 (compatible; Baiduspider/2.0; +http://www.baidu.com/search/spider.html)
User-agent: Baiduspider
Disallow: /
中优先级
Gemini-Deep-Research Google Gemini Deep Research 深度研究功能的资料收集代理。 Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Gemini-Deep-Research; +https://gemini.google/overview/deep-research/) Chrome/135.0.0.0 Safari/537.36
User-agent: Gemini-Deep-Research
Disallow: /
Google-CloudVertexBot Google Vertex AI Agent Builder 访问代理,仅应站长请求抓取,不影响 Google 搜索。 Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/141.0.7390.122 Mobile Safari/537.36 (compatible; Google-CloudVertexBot; +https://cloud.google.com/enterprise-search)
User-agent: Google-CloudVertexBot
Disallow: /
meta-externalagent Meta Meta LLM(Llama 等)训练数据抓取与内容索引,实测抓取量大。 meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler)
User-agent: meta-externalagent
Disallow: /
meta-externalfetcher用户触发型 Meta Meta AI 用户触发的实时网页抓取。 meta-externalfetcher/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler)
User-agent: meta-externalfetcher
Disallow: /
Meta-WebIndexer Meta 用于改进 Meta AI 搜索的索引爬虫。 meta-webindexer/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler)
User-agent: Meta-WebIndexer
Disallow: /
Amazonbot Amazon Alexa 及 Amazon AI 训练数据抓取,实测抓取量大。 Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amazonbot/0.1; +https://developer.amazon.com/support/amazonbot) Chrome/119.0.6045.214 Safari/537.36
User-agent: Amazonbot
Disallow: /
Applebot Apple Apple 搜索 / Siri 索引爬虫,为 Apple Intelligence 提供实时上下文。 Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.4 Safari/605.1.15 (Applebot/0.1; +http://www.apple.com/go/applebot)
User-agent: Applebot
Disallow: /
Applebot-ExtendedToken 型 · 不发请求 Apple 控制 Applebot 抓取内容能否用于 Apple 基础模型(Apple Intelligence)训练的 Token,不发起独立请求。 Applebot-Extended
User-agent: Applebot-Extended
Disallow: /
QwenBot非官方 阿里·通义千问 疑似通义千问内容抓取,阿里无正式爬虫文档,UA 为社区流传。 Mozilla/5.0 (compatible; QwenBot/1.0; +https://tongyi.aliyun.com/bot)
User-agent: QwenBot
Disallow: /
ChatGLM-Spider非官方 智谱AI 智谱 AI 内容抓取,可能用于模型训练,UA 为日志观测所得。 Mozilla/5.0 (compatible; ChatGLM-Spider/1.0; +https://chatglm.cn/)
User-agent: ChatGLM-Spider
Disallow: /
MinimaxBot非官方 MiniMax 可能用于 MiniMax 模型训练,第三方记录的实验性 UA。 MinimaxBot
User-agent: MinimaxBot
Disallow: /
Minimax-User非官方用户触发型 MiniMax 可能为 MiniMax 用户触发抓取,第三方记录的实验性 UA。 Minimax-User
User-agent: Minimax-User
Disallow: /
PetalBot 华为 华为 Petal Search 搜索索引爬虫,为华为搜索与 AI 产品提供网页索引。 Mozilla/5.0 (compatible;PetalBot;+https://webmaster.petalsearch.com/site/petalbot)
User-agent: PetalBot
Disallow: /
低优先级
MistralAI-User用户触发型 Mistral AI Mistral Le Chat 的实时引用抓取。 Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-User/1.0; +https://docs.mistral.ai/robots)
User-agent: MistralAI-User
Disallow: /
DuckAssistBot DuckDuckGo DuckDuckGo AI 回答(DuckAssist)的索引抓取。 DuckAssistBot/1.2; (+http://duckduckgo.com/duckassistbot.html)
User-agent: DuckAssistBot
Disallow: /

检测结果怎么读?三级判定逻辑

1先看 robots.txt

工具按 robots.txt 协议匹配 UA 分组(最长 token 优先,组内最长路径规则优先,支持 * 通配与 $ 锚定)。被 Disallow 的爬虫会遵守规则不再抓取——这是站点级开关,改完即刻生效,但只对守规矩的爬虫有效。

2再看 HTTP 状态码

robots 允许但返回 403/401/429/5xx,说明 CDN、WAF 或主机防火墙在 UA 层面拦截(软屏蔽)。这类拦截后台看不出来,是 GEO 优化最常见的翻车点,需要到 CDN/防火墙白名单里放行对应爬虫 UA。

3最后看 Meta robots

页面能抓到之后,<meta name="robots" content="noindex"> 决定是否收录引用——这是页面级开关。注意 Disallow 与 noindex 互斥:robots 里屏蔽的页面爬虫根本抓不到,noindex 写了也白写。

4按用途决定放行策略

想进 AI 搜索引用:放行 OAI-SearchBot、PerplexityBot、Claude-SearchBot、bingbot、Googlebot 等搜索索引类;不想进训练语料:Disallow GPTBot、ClaudeBot、Bytespider 等训练类;Token 型的 Google-Extended、Applebot-Extended 写进 robots.txt 即生效,不产生任何抓取。

如何放行或屏蔽 AI 爬虫?

robots.txt:站点级声明

放在网站根目录,按 UA 逐个声明。示例——放行 ChatGPT 搜索、屏蔽训练采集:

User-agent: OAI-SearchBot
Allow: /

User-agent: GPTBot
Disallow: /

User-agent: Bytespider
Disallow: /

User-agent: *
Allow: /

Meta robots:页面级声明

写在页面 <head> 里,对所有合规爬虫生效:

<!-- 允许抓取与收录 -->
<meta name="robots" content="index, follow">

<!-- 抓取但不收录 -->
<meta name="robots" content="noindex, follow">

<!-- 只对 Google 生效 -->
<meta name="googlebot" content="noindex">

Nginx:UA 层硬拦截

对付不守规矩的爬虫(如无视 robots.txt 的采集器),在 server 块按 UA 直接拒绝:

if ($http_user_agent ~* (Bytespider|FakeBot)) {
    return 403;
}

别忘了 CDN 和宝塔防火墙

很多站长 robots.txt 写得没问题,却被 CDN 的「防 AI 爬虫」开关或宝塔 Nginx 防火墙的 UA 规则误杀。如果本工具检测出「HTTP 拒绝」,先检查 CDN/WAF 的 Bot 管理设置,把需要放行的爬虫 UA 加入白名单。

AI爬虫抓取验证常见问题 (FAQ)

为什么要检测 AI 爬虫能否抓取我的网站?

AI 搜索(ChatGPT 搜索、Perplexity、Kimi、豆包等)和大模型训练都靠各自的爬虫抓取网页内容。如果 robots.txt、Meta robots 或服务器防火墙在无意中拦掉了这些爬虫,你的内容就进不了 AI 的语料和答案引用,等于主动放弃了 AI 时代的曝光入口。反过来,如果你不希望内容被拿去训练,也需要先弄清每个爬虫的身份和规则,再逐个声明禁止。本工具把这两类排查一次做完。

GPTBot 和 OAI-SearchBot 有什么区别?

GPTBot 负责 OpenAI 模型训练数据抓取,放行意味着内容可能进入 GPT 的训练语料;OAI-SearchBot 负责 ChatGPT Search 的搜索索引(非训练用途),放行它有助于页面出现在 ChatGPT 搜索结果里。两者是独立的 robots.txt token,需要分别声明。类似地,ClaudeBot 与 Claude-SearchBot、KimiBot 与 Kimi-SearchBot 也是「训练」与「搜索索引」的分工。

robots.txt 屏蔽和 Meta robots 屏蔽有什么区别?

robots.txt 是站点级开关:被 Disallow 的 URL,合规爬虫根本不会来抓,所以 meta 标签也就无从谈起。Meta robots 是页面级开关:页面可以被抓取,但 content 里的 noindex 告诉爬虫不要收录/引用。想让内容完全消失用 robots.txt;想让内容被抓取但不收录用 noindex;两者同时 Disallow + noindex 是矛盾配置,爬虫看不到 noindex,页面仍可能被引用。

「HTTP 拒绝(软屏蔽)」是什么意思?

指 robots.txt 和 Meta robots 都允许抓取,但服务器对爬虫 UA 返回了 403/401 等拒绝状态。常见原因:CDN 或主机防火墙开启了「防恶意爬虫」规则、WAF 按 UA 特征误杀、或宝塔 Nginx 防火墙的 UA 黑名单里有相关关键词。这种屏蔽在网站后台看不出来,只能靠实测发现,也是 GEO 优化中最常被忽视的问题。

如何让网站进入 ChatGPT 搜索、Perplexity 等 AI 搜索结果?

至少放行对应产品的搜索索引爬虫:OAI-SearchBot(ChatGPT 搜索)、PerplexityBot(Perplexity)、Claude-SearchBot(Claude 搜索)、bingbot(Copilot)、Googlebot(AI Overviews/AI Mode),国内产品放行 Kimi-SearchBot、Baiduspider 等。同时确保这些 UA 实际访问返回 200,页面没有被 noindex,核心内容是静态 HTML 直出而非 JS 渲染。检测工具会逐项核对这些条件。

检测会对我的网站造成压力吗?

一次检测会对目标 URL 发起 34 个并发请求(Token 型的两个爬虫不发请求),每个请求只读取页面前 160KB(够解析 meta robots 即止),与 36 位访客同时打开一次页面相当,正常网站毫无压力。服务端对单 IP 限流每分钟 6 次,避免工具被滥用。Token 型的 Google-Extended 与 Applebot-Extended 不发起任何请求,robots.txt 声明即生效。