Scrapy基础01

一、Scarpy简介

Scrapy基于事件驱动网络框架 Twisted 编写。（Event-driven networking）

因此，Scrapy基于并发性考虑由非阻塞(即异步)的实现。

二、爬取chouti.com新闻示例

# chouti.py

# -*- coding: utf-8 -*-

import scrapy

from scrapy.http import Request

from scrapy.selector import HtmlXPathSelector

from ..items import Day24SpiderItem

# For windows:

import sys,io

sys.stdout=io.TextIOWrapper(sys.stdout.buffer,encoding='gb18030')

class ChoutiSpider(scrapy.Spider):

    name = 'chouti'

    allowed_domains = ['chouti.com']

    start_urls = ['http://chouti.com/']

    def parse(self, response):

        # print(response.body)

        # print(response.text)

        hxs = HtmlXPathSelector(response)

        item_list = hxs.xpath('//div[@id="content-list"]/div[@class="item"]')

        # 找到首页所有消息的连接、标题、作业信息然后yield给pipeline进行持久化

        for item in item_list:

            link = item.xpath('./div[@class="news-content"]/div[@class="part1"]/a/@href').extract_first()

            title = item.xpath('./div[@class="news-content"]/div[@class="part2"]/@share-title').extract_first()

            author = item.xpath('./div[@class="news-content"]/div[@class="part2"]/a[@class="user-a"]/b/text()').extract_first()

            yield Day24SpiderItem(link=link,title=title,author=author)

        # 找到第二页、第三页、、、第十页的消息，全部爬取下来做持久化

        # hxs.xpath('//div[@id="dig_lcpage"]//a/@href').extract()

        '''或者用正则精确匹配'''

        page_url_list = hxs.xpath('//div[@id="dig_lcpage"]//a[re:test(@href,"/all/hot/recent/\d+")]/@href').extract()

        for url in page_url_list:

            url = "http://dig.chouti.com" + url

            print(url)

            yield Request(url, callback=self.parse, dont_filter=False)

# pipelines.py

# -*- coding: utf-8 -*-

# Define your item pipelines here

#

# Don't forget to add your pipeline to the ITEM_PIPELINES setting

# See: http://doc.scrapy.org/en/latest/topics/item-pipeline.html

class Day24SpiderPipeline(object):

    def __init__(self,file_path):

        self.file_path = file_path  # 文件路径

        self.file_obj = None        # 文件对象：用于读写操作

    @classmethod

    def from_crawler(cls, crawler):

        """

        初始化时候，用于创建pipeline对象

        :param crawler:

        :return:

        """

        val = crawler.settings.get('STORAGE_CONFIG')

        return cls(val)

    def process_item(self, item, spider):

        print(">>>> ",item)

        if 'chouti' == spider.name:

            self.file_obj.write(item.get('link') + "\n" + item.get('title') + "\n" + item.get('author') + "\n\n")

        return item

    def open_spider(self, spider):

        """

        爬虫开始执行时，调用

        :param spider:

        :return:

        """

        if 'chouti' == spider.name:

            # 如果不加：encoding='utf-8' 会导致文件里中文乱码

            self.file_obj = open(self.file_path,mode='a+',encoding='utf-8')

    def close_spider(self, spider):

        """

        爬虫关闭时，被调用

        :param spider:

        :return:

        """

        if 'chouti' == spider.name:

            self.file_obj.close()

# items.py

# -*- coding: utf-8 -*-

# Define here the models for your scraped items

#

# See documentation in:

# http://doc.scrapy.org/en/latest/topics/items.html

import scrapy

class Day24SpiderItem(scrapy.Item):

    link = scrapy.Field()

    title = scrapy.Field()

    author = scrapy.Field()

# settings.py

# -*- coding: utf-8 -*-

# Scrapy settings for day24spider project

#

# For simplicity, this file contains only settings considered important or

# commonly used. You can find more settings consulting the documentation:

#

#     http://doc.scrapy.org/en/latest/topics/settings.html

#     http://scrapy.readthedocs.org/en/latest/topics/downloader-middleware.html

#     http://scrapy.readthedocs.org/en/latest/topics/spider-middleware.html

BOT_NAME = 'day24spider'

SPIDER_MODULES = ['day24spider.spiders']

NEWSPIDER_MODULE = 'day24spider.spiders'

# Crawl responsibly by identifying yourself (and your website) on the user-agent

#USER_AGENT = 'day24spider (+http://www.yourdomain.com)'

# Obey robots.txt rules

ROBOTSTXT_OBEY = True

# Configure maximum concurrent requests performed by Scrapy (default: 16)

#CONCURRENT_REQUESTS = 32

# Configure a delay for requests for the same website (default: 0)

# See http://scrapy.readthedocs.org/en/latest/topics/settings.html#download-delay

# See also autothrottle settings and docs

#DOWNLOAD_DELAY = 3

# The download delay setting will honor only one of:

#CONCURRENT_REQUESTS_PER_DOMAIN = 16

#CONCURRENT_REQUESTS_PER_IP = 16

# Disable cookies (enabled by default)

#COOKIES_ENABLED = False

# Disable Telnet Console (enabled by default)

#TELNETCONSOLE_ENABLED = False

# Override the default request headers:

#DEFAULT_REQUEST_HEADERS = {

#   'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8',

#   'Accept-Language': 'en',

#}

# Enable or disable spider middlewares

# See http://scrapy.readthedocs.org/en/latest/topics/spider-middleware.html

#SPIDER_MIDDLEWARES = {

#    'day24spider.middlewares.Day24SpiderSpiderMiddleware': 543,

#}

# Enable or disable downloader middlewares

# See http://scrapy.readthedocs.org/en/latest/topics/downloader-middleware.html

#DOWNLOADER_MIDDLEWARES = {

#    'day24spider.middlewares.MyCustomDownloaderMiddleware': 543,

#}

# Enable or disable extensions

# See http://scrapy.readthedocs.org/en/latest/topics/extensions.html

#EXTENSIONS = {

#    'scrapy.extensions.telnet.TelnetConsole': None,

#}

# Configure item pipelines

# See http://scrapy.readthedocs.org/en/latest/topics/item-pipeline.html

ITEM_PIPELINES = {

   'day24spider.pipelines.Day24SpiderPipeline': 300,

}

# Enable and configure the AutoThrottle extension (disabled by default)

# See http://doc.scrapy.org/en/latest/topics/autothrottle.html

#AUTOTHROTTLE_ENABLED = True

# The initial download delay

#AUTOTHROTTLE_START_DELAY = 5

# The maximum download delay to be set in case of high latencies

#AUTOTHROTTLE_MAX_DELAY = 60

# The average number of requests Scrapy should be sending in parallel to

# each remote server

#AUTOTHROTTLE_TARGET_CONCURRENCY = 1.0

# Enable showing throttling stats for every response received:

#AUTOTHROTTLE_DEBUG = False

# Enable and configure HTTP caching (disabled by default)

# See http://scrapy.readthedocs.org/en/latest/topics/downloader-middleware.html#httpcache-middleware-settings

#HTTPCACHE_ENABLED = True

#HTTPCACHE_EXPIRATION_SECS = 0

#HTTPCACHE_DIR = 'httpcache'

#HTTPCACHE_IGNORE_HTTP_CODES = []

#HTTPCACHE_STORAGE = 'scrapy.extensions.httpcache.FilesystemCacheStorage'

STORAGE_CONFIG = "chouti.json"

DEPTH_LIMIT = 1

三、classmethod方法应用

from_crawler() --> __init__()

Scrapy基础01的更多相关文章

javascript基础01
javascript基础01 Javascript能做些什么? 给予页面灵魂,让页面可以动起来,包括动态的数据,动态的标签,动态的样式等等. 如实现到轮播图.拖拽.放大镜等,而动态的数据就好比不像没有 ...
Androd核心基础01
Androd核心基础01包含的主要内容如下 Android版本简介 Android体系结构 JVM和DVM的区别常见adb命令操作 Android工程目录结构点击事件的四种形式电话拨号器Demo ...
java基础学习05(面向对象基础01)
面向对象基础01 1.理解面向对象的概念 2.掌握类与对象的概念3.掌握类的封装性4.掌握类构造方法的使用实现的目标 1.类与对象的关系.定义.使用 2.对象的创建格式,可以创建多个对象3.对象的内 ...
Linux基础01 学会使用命令帮助
Linux基础01 学会使用命令帮助概述在linux终端,面对命令不知道怎么用,或不记得命令的拼写及参数时,我们需要求助于系统的帮助文档:linux系统内置的帮助文档很详细,通常能解决我们的问题, ...
可满足性模块理论(SMT)基础 - 01 - 自动机和斯皮尔伯格算术
可满足性模块理论(SMT)基础 - 01 - 自动机和斯皮尔伯格算术前言如果,我们只给出一个数学问题的(比如一道数独题)约束条件,是否有程序可以自动求出一个解? 可满足性模理论(SMT - Sat ...
LibreOJ 2003. 「SDOI2017」新生舞会基础01分数规划最大权匹配
#2003. 「SDOI2017」新生舞会内存限制:256 MiB时间限制:1500 ms标准输入输出题目类型:传统评测方式:文本比较上传者: 匿名提交提交记录统计讨论测试数据题目描述 ...
java基础 01
java基础01 1. /** * JDK: (Java Development ToolKit) java开发工具包.JDK是整个java的核心! * 包括了java运行环境 JRE(Java Ru ...
0.Python 爬虫之Scrapy入门实践指南（Scrapy基础知识）
目录 0.0.Scrapy基础 0.1.Scrapy 框架图 0.2.Scrapy主要包括了以下组件: 0.3.Scrapy简单示例如下: 0.4.Scrapy运行流程如下: 0.5.还有什么? 0. ...
081 01 Android 零基础入门 02 Java面向对象 01 Java面向对象基础 01 初识面向对象 06 new关键字
081 01 Android 零基础入门 02 Java面向对象 01 Java面向对象基础 01 初识面向对象 06 new关键字本文知识点:new关键字说明:因为时间紧张,本人写博客过程中只是 ...

随机推荐

ssh-keygen适用场景与rsync使用id_rsa技巧
ssh-keygen工具可以实现免密码登录服务器可参考之前的blog:http://www.cnblogs.com/Mrhuangrui/p/4565333.html写的比较粗糙原理说明使用ssh- ...
linux中$#,$0,$1,$2,$@,$*,$$,$?的含义
$# 是传给脚本的参数个数$0 是脚本本身的文件名$1 是脚本后接的第一个参数$2 是脚本后接的第二个参数$@ 是传给脚本的所有参数列表,"$1" "$2" & ...
【BZOJ2940】条纹（博弈论）
[BZOJ2940]条纹(博弈论) 题面 BZOJ 神TM权限题. 题解我们把题目看成取石子的话,题目就变成了这样: 有一堆$m$个石头,每次可以取走$c,z,n$个,每次取完之后可以把当前 ...
[luogu1198][bzoj1012][JSOI2008]最大数【线段树+分块】
题目描述区间查询最大值,结尾插入,强制在线. 分析线段树可以做,但是练了一下分块,发现自己打错了两个地方,一个是分块的地方把/打成了%,还有是分块的时候标号要-1. 其他也没什么要多讲的. 代码 ...
CentOS下Denyhosts的安装和使用
安装默认yum就可以进行安装 yum install denyhosts* -y 配置配置文件路径: /etc/denyhosts.conf ; YUM安装时其实已经配置好了大部分,我们自己稍作改 ...
gevent多协程运用
#导包 import gevent #猴子补丁 from gevent import monkey monkey.patch_all() from d8_db import ConnectMysql ...
Vue+Django2.0 restframework打造前后端分离的生鲜电商项目（3）
1.drf前期准备 1.django-rest-framework官方文档 https://www.django-rest-framework.org/ #直接百度找到的djangorestframe ...
详解清除浮动的多种方式（clearfix）
说明本文适合知道HTML 与 CSS基础知识的读者,或者想要了解清除浮动背后原理的读者! 1.什么是浮动首先我们需要知道定位元素在页面中的位置就是定位,解决问题之前我们先来了解下几种定位方式 : ...
自定义QMenu
参考: http://blog.csdn.net/qq1623803207/article/details/77449884 http://blog.sina.com.cn/s/blog_a6fb6c ...
Java多线程-详细版
基本概念解释并发:一个处理器处理多个任务,这些任务对于处理器来说是交替运行的,每个时间点只有一个任务在进行. 并行:多个处理器处理多个任务,这些任务是同时运行的.每个时间点有多个任务同时进行. 进程 ...

Scrapy基础01

一、Scarpy简介

二、爬取chouti.com新闻示例

三、classmethod方法应用

Scrapy基础01的更多相关文章

随机推荐

热门专题