Scrapy实战：爬取http://quotes.toscrape.com网站数据

需要学习的地方：

1.Scrapy框架流程梳理，各文件的用途等

2.在Scrapy框架中使用MongoDB数据库存储数据

3.提取下一页链接，回调自身函数再次获取数据

重点：从当前页获取下一页的链接，传给函数自身继续发起请求

next = response.css('.pager .next a::attr(href)').extract_first() # 获取下一页的相对链接
url = response.urljoin(next) # 生成完整的下一页链接
yield scrapy.Request(url=url, callback=self.parse) # 把下一页的链接回调给自身再次请求

站点:http://quotes.toscrape.com

该站点网页结构比较简单，需要的数据都在div标签中

操作步骤：

1.创建项目

# scrapy startproject quotetutorial

此时目录结构如下：

2.生成爬虫文件

# cd quotetutorial
# scrapy genspider quotes quotes.toscrape.com # 若是有多个爬虫多次操作该命令即可

3.编辑items.py文件，获取需要输出的数据

import scrapy

class QuoteItem(scrapy.Item):

    # define the fields for your item here like:

    # name = scrapy.Field()

    text = scrapy.Field()

    author = scrapy.Field()

    tags = scrapy.Field()

4.编辑quotes.py文件，爬取网站数据

# -*- coding: utf-8 -*-

import scrapy

from quotetutorial.items import QuoteItem

class QuotesSpider(scrapy.Spider):

    name = 'quotes'

    allowed_domains = ['quotes.toscrape.com']

    start_urls = ['http://quotes.toscrape.com/']

    def parse(self, response):

        # print(response.status) # 200

        quotes = response.css('.quote')

        for quote in quotes:

            item = QuoteItem()

            text = quote.css('.text::text').extract_first()

            author = quote.css('.author::text').extract_first()

            tags = quote.css('.tags .tag::text').extract()

            item['text'] = text

            item['author'] = author

            item['tags'] = tags

            yield item

        next = response.css('.pager .next a::attr(href)').extract_first()  # 获取下一页的相对链接

        url = response.urljoin(next)  # 生成完整的下一页链接

        yield scrapy.Request(url=url, callback=self.parse)  # 把下一页的链接回调给自身再次请求

5.编写pipelines.py文件，进一步处理item数据，保存到mongodb数据库

# -*- coding: utf-8 -*-

# Define your item pipelines here

#

# Don't forget to add your pipeline to the ITEM_PIPELINES setting

# See: https://doc.scrapy.org/en/latest/topics/item-pipeline.html

# 使用的话需要在settings文件中设置

import pymongo as pymongo

from scrapy.exceptions import DropItem

class TextPipeline(object):

    """对输出的item进行进一步的处理"""

    def __init__(self):

        self.limit = 50

    def process_item(self, item, spider):

        if item['text']:

            if len(item['text']) > self.limit:

                item['text'] = item['text'][0:self.limit].rstrip() + '......'

            return item

        else:

            return DropItem('Missing Text!')

class MongoPipeline(object):

    """把输出的item保存到MongoDB数据库"""

    def __init__(self, mongo_url, mongo_db):

        self.mongo_uri = mongo_url

        self.mongo_db = mongo_db

    @classmethod

    def from_crawler(cls, crawler):

        """从settings文件获取配置信息"""

        return cls(

            mongo_url=crawler.settings.get('MONGO_URI'),

            mongo_db=crawler.settings.get('MONGO_DB')

        )

    def open_spider(self, spider):

        """初始化mongodb"""

        self.client = pymongo.MongoClient(self.mongo_uri)

        self.db = self.client[self.mongo_db]  # 为啥用[],而不是()

    def process_item(self, item, spider):

        name = item.__class__.__name__  # 获取item的名称用作表名,也就是QuoteItem

        self.db[name].insert(dict(item))  # 为啥要用dict(item)

        return item

    def close_spider(self, spider):

        self.client.close()

6.编辑配置文件，增加mongodb数据库参数，以及使用的pipeline管道参数

ITEM_PIPELINES = {

   # 'quotetutorial.pipelines.TextPipeline': 300,

   'quotetutorial.pipelines.MongoPipeline': 400,

}

MONGO_URI = 'localhost'

MONGO_DB = 'quotestutorial'

7.执行程序

# scrapy crawl quotes

8.保存到文件

# scrapy crawl quotes -o quotes.json # 保存成json文件
# scrapy crawl quotes -o quotes.csv # 保存成csv文件
# scrapy crawl quotes -o quotes.xml # 保存成xml文件
# scrapy crawl quotes -o quotes.jl # 保存成jl文件
# scrapy crawl quotes -o quotes.pickle # 保存成pickle文件
# scrapy crawl quotes -o quotes.marshal # 保存成marshal文件
# scrapy crawl quotes -o ftp://user:password@ftp.example.com/path/quotes.csv # 生成csv文件保存到远程FTP上

效果：

源码下载地址：https://files.cnblogs.com/files/sanduzxcvbnm/quotetutorial.7z

Scrapy实战：爬取http://quotes.toscrape.com网站数据的更多相关文章

简单的scrapy实战:爬取腾讯招聘北京地区的相关招聘信息
简单的scrapy实战:爬取腾讯招聘北京地区的相关招聘信息简单的scrapy实战:爬取腾讯招聘北京地区的相关招聘信息系统环境:Fedora22(昨天已安装scrapy环境) 爬取的开始URL:ht ...
教程+资源,python scrapy实战爬取知乎最性感妹子的爆照合集(12G)!
一.出发点: 之前在知乎看到一位大牛(二胖)写的一篇文章:python爬取知乎最受欢迎的妹子(大概题目是这个,具体记不清了),但是这位二胖哥没有给出源码,而我也没用过python,正好顺便学一学,所以 ...
scrapy实战--爬取最新美剧
现在写一个利用scrapy爬虫框架爬取最新美剧的项目. 准备工作: 目标地址:http://www.meijutt.com/new100.html 爬取项目:美剧名称.状态.电视台.更新时间 1.创建 ...
<scrapy爬虫>爬取quotes.toscrape.com
1.创建scrapy项目 dos窗口输入: scrapy startproject quote cd quote 2.编写item.py文件(相当于编写模板,需要爬取的数据在这里定义) import ...
Scrapy Learning笔记（四）- Scrapy双向爬取
摘要:介绍了使用Scrapy进行双向爬取(对付分类信息网站)的方法. 所谓的双向爬取是指以下这种情况,我要对某个生活分类信息的网站进行数据爬取,譬如要爬取租房信息栏目,我在该栏目的索引页看到如下页面, ...
第三百三十节，web爬虫讲解2—urllib库爬虫—实战爬取搜狗微信公众号—抓包软件安装Fiddler4讲解
第三百三十节,web爬虫讲解2—urllib库爬虫—实战爬取搜狗微信公众号—抓包软件安装Fiddler4讲解封装模块 #!/usr/bin/env python # -*- coding: utf- ...
使用scrapy框架爬取自己的博文（2）
之前写了一篇用scrapy框架爬取自己博文的博客,后来发现对于中文的处理一直有问题- - 显示的时候 [u'python\u4e0b\u722c\u67d0\u4e2a\u7f51\u9875\u76 ...
如何提高scrapy的爬取效率
提高scrapy的爬取效率增加并发: 默认scrapy开启的并发线程为32个,可以适当进行增加.在settings配置文件中修改CONCURRENT_REQUESTS = 100值为100,并发设置 ...
九 web爬虫讲解2—urllib库爬虫—实战爬取搜狗微信公众号—抓包软件安装Fiddler4讲解
封装模块 #!/usr/bin/env python # -*- coding: utf-8 -*- import urllib from urllib import request import j ...

随机推荐

java中inputstream的使用
java中的inputstream是一个面向字节的流抽象类,其依据详细应用派生出各种详细的类. 比方FileInputStream就是继承于InputStream,专门用来读取文件流的对象,其详细继承 ...
There was a conflict between
解读,首先搜索到第一个5>的开头的那一行,确认是在编译哪一个项目. 那么后面的冲突,就是在和这个项目冲突. There was a conflict between "log4net, ...
Tarjan Algorithm
List Tarjan Algorithm List Knowledge 基本知识基本概念复杂度有向图 Code 缩点 Code 用途无向图 Articulation Point-割顶与连通度 ...
[模板] manacher(教程)
还是不会马拉车啊.今天又学了一遍,在这里讲一下. 其实就是一个很妙的思路,就是设置一个辅助的数组len,记录每个点的最大对称长度,然后再存一个mx记录最大的对称子串的右端点.先开二倍数组,然后一点点扩 ...
codevs1557 热浪（堆优化dijkstra）
1557 热浪时间限制: 1 s 空间限制: 256000 KB 题目等级 : 钻石 Diamond 题解查看运行结果题目描述 Description 德克萨斯纯朴的民眾们这个夏 ...
[Swift通天遁地]二、表格表单-(15)自定义表单文本框内容的格式
★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★★➤微信公众号:山青咏芝(shanqingyongzhi)➤博客园地址:山青咏芝(https://www.cnblogs. ...
memcache缓存系统
一.缓存系统静态web页面: 1.在静态Web程序中,客户端使用Web浏览器(IE.FireFox等)经过网络(Network)连接到服务器上,使用HTTP协议发起一个请求(Request),告诉服 ...
【洛谷4219】[BJOI2014]大融合（线段树分治）
题目: 洛谷4219 分析: 很明显,查询的是删掉某条边后两端点所在连通块大小的乘积. 有加边和删边,想到LCT.但是我不会用LCT查连通块大小啊.果断弃了有加边和删边,还跟连通性有关,于是开始yy ...
JAVA FORK JOIN EXAMPLE--转
http://www.javacreed.com/java-fork-join-example/ Java 7 introduced a new type of ExecutorService (Ja ...
C# Autofac 出现尝试创建“XXController”类型的控制器时出错。请确保控制器具有无参数公共构造函数错误解决方案
出现以下错误: 总结解决方案: 本项目采用构造函数方法进行依赖注入,由于个人原因在业务层相互注入了接口,导致交叉:报错

Scrapy实战：爬取http://quotes.toscrape.com网站数据

Scrapy实战：爬取http://quotes.toscrape.com网站数据的更多相关文章

随机推荐

热门专题