模拟 GPTBot、ClaudeBot、CCBot、Kimi、Bytespider 等 36 款主流 AI 爬虫UA访问 · robots.txt / Meta robots / HTTP 状态三级检测 · 覆盖中美主要 AI 厂商
返回 Favicon & ICO 在线生成器 您还可以使用 蜘蛛大全 HTTP状态码大全 全国DNS大全 一键提交外链工具内置 36 款主流 AI 爬虫(美国厂商 24 款 / 中国厂商 11 款),按高 / 中 / 低三档优先级并发检测,下方清单可查看每款爬虫的完整 UA 与屏蔽方法。
地区标识:★中国厂商 ★美国厂商(其他地区不标注); 优先级:高 = 主流 AI 搜索/聊天产品、基础模型训练或高抓取量;中 = 有明确用途但流量或资料完整度相对较低;低 = 特定场景补充。 带 非官方 标注的爬虫,其 UA 来自日志观测或第三方记录,厂商尚未正式确认。蜘蛛名称可点击跳转蜘蛛大全对应详情页。
| 优先级 | 名称 | 厂商 | 说明 | 完整 User-Agent | robots.txt 屏蔽方法 |
|---|---|---|---|---|---|
| 高优先级 | |||||
| 高 | ★GPTBot | OpenAI | OpenAI GPT 模型训练数据抓取,放行与否直接影响内容是否进入 GPT 语料。 | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot |
User-agent: GPTBot Disallow: / |
| 高 | ★OAI-SearchBot | OpenAI | ChatGPT Search 搜索索引抓取(非训练用途),放行有助于内容被 ChatGPT 搜索引用。 | Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot |
User-agent: OAI-SearchBot Disallow: / |
| 高 | ★ChatGPT-User用户触发型 | OpenAI | 用户在 ChatGPT 中触发实时联网回答时的访问代理。 | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot |
User-agent: ChatGPT-User Disallow: / |
| 高 | ★ClaudeBot | Anthropic | Anthropic Claude 模型训练数据抓取。 | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ClaudeBot/1.0; +claudebot@anthropic.com) |
User-agent: ClaudeBot Disallow: / |
| 高 | ★Claude-SearchBot | Anthropic | Anthropic 搜索索引抓取,用于 Claude 搜索结果与内容引用。 | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Claude-SearchBot/1.0; +https://www.anthropic.com) |
User-agent: Claude-SearchBot Disallow: / |
| 高 | ★Claude-User用户触发型 | Anthropic | 用户在 Claude 中触发实时联网时的访问代理。 | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Claude-User/1.0; +Claude-User@anthropic.com) |
User-agent: Claude-User Disallow: / |
| 高 | ★Googlebot | Google 搜索主爬虫,同时是 AI Overviews、AI Mode 与 Gemini 实时引用的数据来源。 | Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html) |
User-agent: Googlebot Disallow: / |
|
| 高 | ★Google-ExtendedToken 型 · 不发请求 | 控制内容能否用于 Gemini / Vertex AI 训练的 robots.txt Token,不发起独立请求,声明即生效。 | Mozilla/5.0 (compatible; Google-Extended) |
User-agent: Google-Extended Disallow: / |
|
| 高 | ★Google-GeminiNotebook用户触发型 | Gemini Notebook(原 NotebookLM)用户添加来源 URL 时的抓取器。 | Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/137.0.0.0 Safari/537.36 (compatible; Google-GeminiNotebook; +https://developers.google.com/crawling/docs/crawlers-fetchers/google-gemininotebook) |
User-agent: Google-GeminiNotebook Disallow: / |
|
| 高 | ★Google-Agent(桌面)用户触发型 | Google 托管 Agent(如 Project Mariner)桌面端访问代理。 | Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko; compatible; Google-Agent; +https://developers.google.com/crawling/docs/crawlers-fetchers/google-agent) Chrome/W.X.Y.Z Safari/537.36 |
User-agent: Google-Agent Disallow: / |
|
| 高 | ★Google-Agent(移动)用户触发型 | Google 托管 Agent 移动端访问代理。 | Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/W.X.Y.Z Mobile Safari/537.36 (compatible; Google-Agent; +https://developers.google.com/crawling/docs/crawlers-fetchers/google-agent) |
User-agent: Google-Agent Disallow: / |
|
| 高 | ★PerplexityBot | Perplexity | Perplexity 答案引擎的索引抓取,用于搜索结果与内容引用。 | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot) |
User-agent: PerplexityBot Disallow: / |
| 高 | ★Perplexity-User用户触发型 | Perplexity | 用户在 Perplexity 触发实时检索时的访问代理。 | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user) |
User-agent: Perplexity-User Disallow: / |
| 高 | ★CCBot | Common Crawl | Common Crawl 开源网页存档爬虫,众多大模型预训练语料的上游数据源,屏蔽它等于退出公共语料。 | CCBot/2.0 (https://commoncrawl.org/faq/) |
User-agent: CCBot Disallow: / |
| 高 | ★bingbot | Microsoft | 同时驱动 Bing 搜索与 Copilot AI 回答的基础爬虫,GEO 检测不可缺。 | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/116.0.1938.76 Safari/537.36 |
User-agent: bingbot Disallow: / |
| 高 | ★KimiBot | 月之暗面 Kimi | 月之暗面(Moonshot AI)Kimi 模型训练数据抓取,遵守 robots.txt。 | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; KimiBot/1.0; +https://www.kimi.com/policies/kimi-crawlers |
User-agent: KimiBot Disallow: / |
| 高 | ★Kimi-SearchBot | 月之暗面 Kimi | Kimi AI 搜索索引抓取。 | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; Kimi-SearchBot/1.0; +https://www.kimi.com/policies/kimi-crawlers |
User-agent: Kimi-SearchBot Disallow: / |
| 高 | ★Kimi-User用户触发型 | 月之暗面 Kimi | 用户向 Kimi 提问需要实时检索时按需抓取。 | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; Kimi-User/1.0; +https://www.kimi.com/policies/kimi-crawlers |
User-agent: Kimi-User Disallow: / |
| 高 | ★DeepSeekBot非官方 | DeepSeek | DeepSeek AI 内容抓取,官方尚无爬虫文档,UA 来自访问日志观测。 | Mozilla/5.0 (compatible; DeepSeekBot/1.0; +https://www.deepseek.com/bot) |
User-agent: DeepSeekBot Disallow: / |
| 高 | ★Bytespider | 字节跳动 | 字节跳动(豆包等)LLM 训练数据抓取,实测抓取量大,robots.txt 合规性存在争议。 | Mozilla/5.0 (Linux; Android 5.0) AppleWebKit/537.36 (KHTML, like Gecko) Mobile Safari/537.36 (compatible; Bytespider; spider-feedback@bytedance.com) |
User-agent: Bytespider Disallow: / |
| 高 | ★Baiduspider | 百度 | 百度搜索主爬虫,百度 AI 搜索与文心一言联网回答的基础数据来源。 | Mozilla/5.0 (compatible; Baiduspider/2.0; +http://www.baidu.com/search/spider.html) |
User-agent: Baiduspider Disallow: / |
| 中优先级 | |||||
| 中 | ★Gemini-Deep-Research | Gemini Deep Research 深度研究功能的资料收集代理。 | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Gemini-Deep-Research; +https://gemini.google/overview/deep-research/) Chrome/135.0.0.0 Safari/537.36 |
User-agent: Gemini-Deep-Research Disallow: / |
|
| 中 | ★Google-CloudVertexBot | Vertex AI Agent Builder 访问代理,仅应站长请求抓取,不影响 Google 搜索。 | Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/141.0.7390.122 Mobile Safari/537.36 (compatible; Google-CloudVertexBot; +https://cloud.google.com/enterprise-search) |
User-agent: Google-CloudVertexBot Disallow: / |
|
| 中 | ★meta-externalagent | Meta | Meta LLM(Llama 等)训练数据抓取与内容索引,实测抓取量大。 | meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler) |
User-agent: meta-externalagent Disallow: / |
| 中 | ★meta-externalfetcher用户触发型 | Meta | Meta AI 用户触发的实时网页抓取。 | meta-externalfetcher/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler) |
User-agent: meta-externalfetcher Disallow: / |
| 中 | ★Meta-WebIndexer | Meta | 用于改进 Meta AI 搜索的索引爬虫。 | meta-webindexer/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler) |
User-agent: Meta-WebIndexer Disallow: / |
| 中 | ★Amazonbot | Amazon | Alexa 及 Amazon AI 训练数据抓取,实测抓取量大。 | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amazonbot/0.1; +https://developer.amazon.com/support/amazonbot) Chrome/119.0.6045.214 Safari/537.36 |
User-agent: Amazonbot Disallow: / |
| 中 | ★Applebot | Apple | Apple 搜索 / Siri 索引爬虫,为 Apple Intelligence 提供实时上下文。 | Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.4 Safari/605.1.15 (Applebot/0.1; +http://www.apple.com/go/applebot) |
User-agent: Applebot Disallow: / |
| 中 | ★Applebot-ExtendedToken 型 · 不发请求 | Apple | 控制 Applebot 抓取内容能否用于 Apple 基础模型(Apple Intelligence)训练的 Token,不发起独立请求。 | Applebot-Extended |
User-agent: Applebot-Extended Disallow: / |
| 中 | ★QwenBot非官方 | 阿里·通义千问 | 疑似通义千问内容抓取,阿里无正式爬虫文档,UA 为社区流传。 | Mozilla/5.0 (compatible; QwenBot/1.0; +https://tongyi.aliyun.com/bot) |
User-agent: QwenBot Disallow: / |
| 中 | ★ChatGLM-Spider非官方 | 智谱AI | 智谱 AI 内容抓取,可能用于模型训练,UA 为日志观测所得。 | Mozilla/5.0 (compatible; ChatGLM-Spider/1.0; +https://chatglm.cn/) |
User-agent: ChatGLM-Spider Disallow: / |
| 中 | ★MinimaxBot非官方 | MiniMax | 可能用于 MiniMax 模型训练,第三方记录的实验性 UA。 | MinimaxBot |
User-agent: MinimaxBot Disallow: / |
| 中 | ★Minimax-User非官方用户触发型 | MiniMax | 可能为 MiniMax 用户触发抓取,第三方记录的实验性 UA。 | Minimax-User |
User-agent: Minimax-User Disallow: / |
| 中 | ★PetalBot | 华为 | 华为 Petal Search 搜索索引爬虫,为华为搜索与 AI 产品提供网页索引。 | Mozilla/5.0 (compatible;PetalBot;+https://webmaster.petalsearch.com/site/petalbot) |
User-agent: PetalBot Disallow: / |
| 低优先级 | |||||
| 低 | MistralAI-User用户触发型 | Mistral AI | Mistral Le Chat 的实时引用抓取。 | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-User/1.0; +https://docs.mistral.ai/robots) |
User-agent: MistralAI-User Disallow: / |
| 低 | ★DuckAssistBot | DuckDuckGo | DuckDuckGo AI 回答(DuckAssist)的索引抓取。 | DuckAssistBot/1.2; (+http://duckduckgo.com/duckassistbot.html) |
User-agent: DuckAssistBot Disallow: / |
工具按 robots.txt 协议匹配 UA 分组(最长 token 优先,组内最长路径规则优先,支持 * 通配与 $ 锚定)。被 Disallow 的爬虫会遵守规则不再抓取——这是站点级开关,改完即刻生效,但只对守规矩的爬虫有效。
robots 允许但返回 403/401/429/5xx,说明 CDN、WAF 或主机防火墙在 UA 层面拦截(软屏蔽)。这类拦截后台看不出来,是 GEO 优化最常见的翻车点,需要到 CDN/防火墙白名单里放行对应爬虫 UA。
页面能抓到之后,<meta name="robots" content="noindex"> 决定是否收录引用——这是页面级开关。注意 Disallow 与 noindex 互斥:robots 里屏蔽的页面爬虫根本抓不到,noindex 写了也白写。
想进 AI 搜索引用:放行 OAI-SearchBot、PerplexityBot、Claude-SearchBot、bingbot、Googlebot 等搜索索引类;不想进训练语料:Disallow GPTBot、ClaudeBot、Bytespider 等训练类;Token 型的 Google-Extended、Applebot-Extended 写进 robots.txt 即生效,不产生任何抓取。
放在网站根目录,按 UA 逐个声明。示例——放行 ChatGPT 搜索、屏蔽训练采集:
User-agent: OAI-SearchBot Allow: / User-agent: GPTBot Disallow: / User-agent: Bytespider Disallow: / User-agent: * Allow: /
写在页面 <head> 里,对所有合规爬虫生效:
<!-- 允许抓取与收录 --> <meta name="robots" content="index, follow"> <!-- 抓取但不收录 --> <meta name="robots" content="noindex, follow"> <!-- 只对 Google 生效 --> <meta name="googlebot" content="noindex">
对付不守规矩的爬虫(如无视 robots.txt 的采集器),在 server 块按 UA 直接拒绝:
if ($http_user_agent ~* (Bytespider|FakeBot)) {
return 403;
}
很多站长 robots.txt 写得没问题,却被 CDN 的「防 AI 爬虫」开关或宝塔 Nginx 防火墙的 UA 规则误杀。如果本工具检测出「HTTP 拒绝」,先检查 CDN/WAF 的 Bot 管理设置,把需要放行的爬虫 UA 加入白名单。
AI 搜索(ChatGPT 搜索、Perplexity、Kimi、豆包等)和大模型训练都靠各自的爬虫抓取网页内容。如果 robots.txt、Meta robots 或服务器防火墙在无意中拦掉了这些爬虫,你的内容就进不了 AI 的语料和答案引用,等于主动放弃了 AI 时代的曝光入口。反过来,如果你不希望内容被拿去训练,也需要先弄清每个爬虫的身份和规则,再逐个声明禁止。本工具把这两类排查一次做完。
GPTBot 负责 OpenAI 模型训练数据抓取,放行意味着内容可能进入 GPT 的训练语料;OAI-SearchBot 负责 ChatGPT Search 的搜索索引(非训练用途),放行它有助于页面出现在 ChatGPT 搜索结果里。两者是独立的 robots.txt token,需要分别声明。类似地,ClaudeBot 与 Claude-SearchBot、KimiBot 与 Kimi-SearchBot 也是「训练」与「搜索索引」的分工。
robots.txt 是站点级开关:被 Disallow 的 URL,合规爬虫根本不会来抓,所以 meta 标签也就无从谈起。Meta robots 是页面级开关:页面可以被抓取,但 content 里的 noindex 告诉爬虫不要收录/引用。想让内容完全消失用 robots.txt;想让内容被抓取但不收录用 noindex;两者同时 Disallow + noindex 是矛盾配置,爬虫看不到 noindex,页面仍可能被引用。
指 robots.txt 和 Meta robots 都允许抓取,但服务器对爬虫 UA 返回了 403/401 等拒绝状态。常见原因:CDN 或主机防火墙开启了「防恶意爬虫」规则、WAF 按 UA 特征误杀、或宝塔 Nginx 防火墙的 UA 黑名单里有相关关键词。这种屏蔽在网站后台看不出来,只能靠实测发现,也是 GEO 优化中最常被忽视的问题。
至少放行对应产品的搜索索引爬虫:OAI-SearchBot(ChatGPT 搜索)、PerplexityBot(Perplexity)、Claude-SearchBot(Claude 搜索)、bingbot(Copilot)、Googlebot(AI Overviews/AI Mode),国内产品放行 Kimi-SearchBot、Baiduspider 等。同时确保这些 UA 实际访问返回 200,页面没有被 noindex,核心内容是静态 HTML 直出而非 JS 渲染。检测工具会逐项核对这些条件。
一次检测会对目标 URL 发起 34 个并发请求(Token 型的两个爬虫不发请求),每个请求只读取页面前 160KB(够解析 meta robots 即止),与 36 位访客同时打开一次页面相当,正常网站毫无压力。服务端对单 IP 限流每分钟 6 次,避免工具被滥用。Token 型的 Google-Extended 与 Applebot-Extended 不发起任何请求,robots.txt 声明即生效。
如果本站工具对你有帮助,欢迎请作者喝杯咖啡