- 视策略而定
GPTBot
OpenAI
GPTBot
用于模型训练与检索的数据抓取。若希望 ChatGPT 能引用你的内容,通常需要放行;若不希望内容进入训练集,可拦截。
- 建议放行
ChatGPT-User
OpenAI
ChatGPT-User
用户在 ChatGPT 里点开链接时的实时抓取,影响「用户问 AI 时能不能打开你的页」。
- 视策略而定
ClaudeBot
Anthropic
ClaudeBot · anthropic-ai
Claude 侧爬虫。与 GPTBot 类似,放行有助于被引用,拦截可限制进入训练/索引。
- 建议放行
PerplexityBot
Perplexity
PerplexityBot
Perplexity 检索爬虫,拦截会直接影响在 Perplexity 答案里的可见度。
- 视策略而定
Google-Extended
Google
Google-Extended
控制内容是否用于 Gemini 等谷歌 AI 产品训练,与 Google 搜索爬虫 Googlebot 不同。
- 视策略而定
Bytespider
ByteDance
Bytespider
字节系检索爬虫,海外站点可按品牌策略决定是否放行。
- 视策略而定
cohere-ai
Cohere
cohere-ai
Cohere 侧爬虫,量级相对小,可按需配置 robots。
- 视策略而定
Meta-ExternalAgent
Meta
Meta-ExternalAgent · FacebookBot
Meta 侧外部抓取,与 Facebook 分享预览抓取不同,按业务需要配置。
共收录 8 个常见 AI 爬虫;实际 UA 可能随厂商更新而变化。
