Python 爬虫-Robots协议

时间：2017-07-25 22:08:00 阅读：501 评论：0 收藏：0 [点我收藏+]

2017-07-25 21:08:16

一、网络爬虫的规模

二、网络爬虫的限制

? 来源审查：判断User‐Agent进行限制
　　检查来访HTTP协议头的User‐Agent域，只响应浏览器或友好爬虫的访问
? 发布公告：Robots协议
　　告知所有爬虫网站的爬取策略，要求爬虫遵守

三、Robots 协议

作用：网站告知网络爬虫哪些页面可以抓取，哪些不行
形式：在网站根目录下的robots.txt文件

如果网站不提供Robots协议则表示该网站允许任意爬虫爬取任意次数。

类人类行为原则上可以不遵守Robots协议

https://www.baidu.com/robots.txt
http://news.sina.com.cn/robots.txt

举例：

https://www.jd.com/robots.txt

User‐agent: *
Disallow: /?*
Disallow: /pop/*.html
Disallow: /pinpai/*.html?*
User‐agent: EtaoSpider
Disallow: /
User‐agent: HuihuiSpider
Disallow: /
User‐agent: GwdangSpider
Disallow: /
User‐agent: WochachaSpider
Disallow: /

# 注释，*代表所有，/代表根目录
User‐agent: *
Disallow: /

Python 爬虫-Robots协议

原文：http://www.cnblogs.com/TIMHY/p/7236670.html

踩

(0)

评论一句话评论（0）

分享档案

更多>

2021年09月23日 (328)
2021年09月24日 (313)
2021年09月17日 (191)
2021年09月15日 (369)
2021年09月16日 (411)
2021年09月13日 (439)
2021年09月11日 (398)
2021年09月12日 (393)
2021年09月10日 (160)
2021年09月08日 (222)