QuotesBot

This is a Scrapy project to scrape quotes from famous people from http://quotes.toscrape.com (github repo).

This project is only meant for educational purposes.

任务：

爬取该网站的名人名言、作者、作者信息（名字，生日、描述）以及名言标签，并保存

import scrapy

import re

class AuthorSpider(scrapy.Spider):

    name = "author"

    start_urls = ["http://quotes.toscrape.com/"]

    def parse(self, response):

        author_page_links = response.css('.author + a')

        yield from response.follow_all(author_page_links, self.parse_author)

        next_page_links = response.css('li.next a')

        yield from response.follow_all(next_page_links, self.parse)

    def parse_author(self, response):

        def extract_with_css(query):

            return response.css(query).get(default="").strip()

        yield {

            "name": extract_with_css("h3.author-title::text"),

            "birthdate": extract_with_css(".author-born-date::text"),

            "bio": extract_with_css(".author-description::text"),

        }

保存：

scrapy crawl spidername -o test.csv

项目练习：

Extracted data

This project extracts quotes, combined with the respective author names and tags. The extracted data looks like this sample:

{

    'author': 'Douglas Adams',

    'text': '“I may not have gone where I intended to go, but I think I ...”',

    'tags': ['life', 'navigation']

}

Spiders

This project contains two spiders and you can list them using the list command:

$ scrapy list

toscrape-css

toscrape-xpath

Both spiders extract the same data from the same website, but toscrape-css employs CSS selectors, while toscrape-xpath employs XPath expressions.

You can learn more about the spiders by going through the Scrapy Tutorial.

Running the spiders

You can run a spider using the scrapy crawl command, such as:

scrapy crawl toscrape-css

If you want to save the scraped data to a file, you can pass the -o option:

scrapy crawl toscrape-css -o quotes.json

项目代码：

class QuotesbotSpider(scrapy.Spider):

    name = "quotesbot"

    start_urls = ["http://quotes.toscrape.com"]

    def parse(self, response, **kwargs):

        for quote in response.css('div.quote'):

            yield {

                "author":quote.css(".author::text").get(),

                "text":quote.css(".text::text").get(),

                "tags":quote.css(".tags meta::attr(content)").get(),

            }

        next_page_link = response.css("li.next a")

        if next_page_link is not None:

            yield from response.follow_all(next_page_link, callbac

结果：

Scrapy 项目：QuotesBot的更多相关文章

亲测——pycharm下运行第一个scrapy项目 ©seven_clear
最近在学习scrapy,就想着用pycharm调试,但不知道怎么弄,从网上搜了很多方法,这里总结一个我试成功了的. 首先当然是安装scrapy,安装教程什么的网上一大堆,这里推荐一个详细的:http: ...
scrapy（一）建立一个scrapy项目
本项目实现了获取stack overflow的问题,语言使用python,框架scrapy框架,选取mongoDB作为持久化数据库,redis做为数据缓存项目源码可以参考我的github:https ...
Python Scrapy项目创建（基础普及篇）
在使用Scrapy开发爬虫时,通常需要创建一个Scrapy项目.通过如下命令即可创建 Scrapy 项目: scrapy startproject ZhipinSpider 在上面命令中,scrapy ...
pycharm创建scrapy项目教程及遇到的坑
最近学习scrapy爬虫框架,在使用pycharm安装scrapy类库及创建scrapy项目时花费了好长的时间,遇到各种坑,根据网上的各种教程,花费了一晚上的时间,终于成功,其中也踩了一些坑,现在整理 ...
【Python3爬虫】第一个Scrapy项目
Python版本:3.5 IDE:Pycharm 今天跟着网上的教程做了第一个Scrapy项目,遇到了很多问题,花了很多时间终于解决了== 一.Scrapy终端(scrapy shell) Sc ...
eclipse创建scrapy项目
1. 您必须创建一个新的Scrapy项目. 进入您打算存储代码的目录中(比如否F:/demo),运行下列命令: scrapy startproject tutorial 2.在eclipse中创建一个 ...
python爬虫scrapy项目详解（关注、持续更新）
python爬虫scrapy项目(一) 爬取目标:腾讯招聘网站(起始url:https://hr.tencent.com/position.php?keywords=&tid=0&st ...
Scrapy项目创建以及目录详情
Scrapy项目创建已经目录详情一.新建项目(scrapy startproject) 在开始爬取之前,必须创建一个新的Scrapy项目.进入自定义的项目目录中,运行下列命令: PS C:\scra ...
Scrapy项目结构分析和工作流程
新建的空Scrapy项目: spiders目录: 负责存放继承自scrapy的爬虫类.里面主要是用于分析response并提取返回的item或者是下一个URL信息,每个Spider负责处理特定的网站或 ...
爬虫系列2：scrapy项目入门案例分析
本文从一个基础案例入手,较为详细的分析了scrapy项目的建设过程(在官方文档的基础上做了调整).主要内容如下: 0.准备工作 1.scrapy项目结构 2.编写spider 3.编写item.py ...

随机推荐

java中的IO处理和使用，API详细介绍（一）
写在前面:本文章基本覆盖了java IO的全部内容,java新IO没有涉及,因为我想和这个分开,以突出那个的重要性,新IO哪一篇文章还没有开始写,估计很快就能和大家见面.照旧,文章依旧以例子为主,因为 ...
linux基础进阶命令详解（输出重定向(2>&1,1>&2,&>file)、输入重定向、管道符、通配符、三种引号、软连接、硬链接、根“/”、绝对路径vs相对路径）
本章命令(共9个): 1 2 3 4 5 6 7 8 9 输出重定向输入重定向管道符通配符三种引号软连接硬链接根"/" 绝对路径vs相对路径 1.输出重定向作用:一 ...
遇到的一个bug
/// <summary> /// 检测玩家是否在机器人的球形碰撞体内,这个碰撞体是机器人的侦测范围,玩家在内部会进行视野检测和声音检测 /// </summary> priv ...
使用 noexcept 我们需要知道什么？
noexcept 关键字 noexcept 是什么? noexcept 是自 C++11 引入的新特性,指定函数是否可能会引发异常,以下是 noexcept 的标准语法: noexcept-expre ...
Testing Round #16 (Unrated)
比赛链接:https://codeforces.com/contest/1351 A - A+B (Trial Problem) #include <bits/stdc++.h> usin ...
Codeforces Round #673 (Div. 2) C. k-Amazing Numbers (DP,思维)
题意:有一组数,分别用长度从\([1,n]\)的区间去取子数组,要求取到的所有子数组中必须有共同的数,如果满足条件数组共同的数中最小的数,否则输出\(-1\). 题解:我们先从后面确定每两个相同数之间 ...
Codeforces Round #669 (Div. 2) A. Ahahahahahahahaha (构造)
题意:有一个长度为偶数只含\(0\)和\(1\)的序列,你可以移除最多\(\frac{n}{2}\)个位置的元素,使得操作后奇数位置的元素和等于偶数位置的元素和,求新序列. 题解:统计\(0\)和\( ...
二叉树增删改查 && 程序实现
二叉排序树定义一棵空树,或者是具有下列性质的二叉树:(1)若左子树不空,则左子树上所有结点的值均小于它的根结点的值:(2)若右子树不空,则右子树上所有结点的值均大于它的根结点的值:(3)左.右子树也 ...
1.搭建NFS环境,用于存储数据
作者微信:tangy8080 电子邮箱:914661180@qq.com 更新时间:2019-06-12 14:59:50 星期三欢迎您订阅和分享我的订阅号,订阅号内会不定期分享一些我自己学习过程 ...
【非原创】codeforces 1029F Multicolored Markers 【贪心+构造】
题目:戳这里题意:给a个红色小方块和b个蓝色小方块,求其能组成的周长最小的矩形,要求红色或蓝色方块至少有一个也是矩形. 思路来源:戳这里解题思路:遍历大矩形可能满足的所有周长,维护最小值即可.需要 ...