robots.txt
robots.txt 是站点根目录的爬虫指令文件,用 User-agent / Allow / Disallow 控制谁可以抓取哪些路径。GEO 时代必须显式 Allow GPTBot、PerplexityBot、Google-Extended,否则 AI 引擎不会引用你。
robots.txt 是站点根目录的爬虫指令文件,用 User-agent / Allow / Disallow 控制谁可以抓取哪些路径。GEO 时代必须显式 Allow GPTBot、PerplexityBot、Google-Extended,否则 AI 引擎不会引用你。
什么是robots.txt?
robots.txt 是站点根目录的爬虫指令文件,用 User-agent / Allow / Disallow 控制谁可以抓取哪些路径。GEO 时代必须显式 Allow GPTBot、PerplexityBot、Google-Extended,否则 AI 引擎不会引用你。
细节与示例
示例片段: ``` User-agent: GPTBot Allow: / User-agent: PerplexityBot Allow: / User-agent: Google-Extended Allow: / ```
相关术语
llms.txt、GEO(生成式引擎优化)、Crawl Budget(抓取预算)。
robots.txt · 术语速览
| 项目 | 内容 |
|---|---|
| 术语 | robots.txt |
| 类别 | 技术 SEO |
| 相关术语 | llms.txt、GEO(生成式引擎优化)、Crawl Budget(抓取预算) |
常见问题
什么是robots.txt?
robots.txt 是站点根目录的爬虫指令文件,用 User-agent / Allow / Disallow 控制谁可以抓取哪些路径。GEO 时代必须显式 Allow GPTBot、PerplexityBot、Google-Extended,否则 AI 引擎不会引用你。
参考资料
作者:23SEOGEO 团队 · 最近更新: