示例说明：

本示例主要是PyQuery解析返回的response页面数据。response.doc解析页面数据是pyspider的主要用法，应该熟练掌握基本使用方法。其他返回类型示例见后续文章。

pyspider爬取的内容通过回调的参数response返回，response有多种解析方式。
1、response.json用于解析json数据
2、response.doc返回的是PyQuery对象
3、response.etree返回的是lxml对象
4、response.text返回的是unicode文本
5、response.content返回的是字节码

使用方法：

PyQuery可以采用CSS选择器作为参数对网页进行解析。关于PyQuery更多用法，可以参考PyQuery complete API。

PyQuery官方参考手册: https://pythonhosted.org/pyquery/

response.doc('.ml.mlt.mtw.cl > li').items()

response.doc('.pti > .pdbt > .authi > em > span').attr('title')

CSS选择器：更多详情请参考CSS Selector Reference

选择器	示例	示例说明
.class	.intro	Selects all elements with class=”intro”
#id	#firstname	Selects the element with id=”firstname”
element	p	Selects all <p> elements
element,element	div, p	Selects all <div> elements and all <p> elements
element element	div p	Selects all <p> elements inside <div> elements
element>element	div > p	Selects all <p> elements where the parent is a <div> element
[attribute]	[target]	Selects all elements with a target attribute
[attribute=value]	[target=_blank]	Selects all elements with target=”_blank”
[attribute^=value]	a[href^=”https”]	Selects every <a> element whose href attribute value begins with “https”
[attribute$=value]	a[href$=”.pdf”]	Selects every <a> element whose href attribute value ends with “.pdf”
[attribute*=value]	a[href*=”w3schools”]	Selects every <a> element whose href attribute value contains the substring “w3schools”
:checked	input:checked	Selects every checked <input> element

示例代码：

1、PyQuery基本语法应用

#!/usr/bin/env python

# -*- encoding: utf-8 -*-

# Created on 2016-10-09 15:16:04

# Project: douban_rent

from pyspider.libs.base_handler import *

groups = {'上海租房': 'https://www.douban.com/group/shanghaizufang/discussion?start=',

          '上海租房@长宁租房/徐汇/静安租房': 'https://www.douban.com/group/zufan/discussion?start=',

          '上海短租日租小组 ': 'https://www.douban.com/group/chuzu/discussion?start=',

          '上海短租': 'https://www.douban.com/group/275832/discussion?start=',

          }

class Handler(BaseHandler):

    crawl_config = {

    }

    @every(minutes=24 * 60)

    def on_start(self):

        for i in groups:

            url = groups[i] + '0'

            self.crawl(url, callback=self.index_page, validate_cert=False)

    @config(age=10 * 24 * 60 * 60)

    def index_page(self, response):

        for each in response.doc('.olt .title a').items():

            self.crawl(each.attr.href, callback=self.detail_page)

        #next page

        for each in response.doc('.next a').items():

            self.crawl(each.attr.href, callback=self.index_page)

    @config(priority=2)

    def detail_page(self, response):

        count_not=0

        notlist=[]

        for i in response.doc('.bg-img-green a').items():

            if i.text() <> response.doc('.from a').text():

                count_not +=1

                notlist.append(i.text())

        for i in notlist: print i

        return {

            "url": response.url,

            "title": response.doc('title').text(),

            "author": response.doc('.from a').text(),

            "time": response.doc('.color-green').text(),

            "content": response.doc('#link-report p').text(),

            "回应数": len([x.text() for x in response.doc('h4').items()])   ,

            # "最后回帖时间": [x for x in response.doc('.pubtime').items()][-1].text(),

            "非lz回帖数": count_not,

            "非lz回帖人数": len(set(notlist)),

            # "主演": [x.text() for x in response.doc('.actor a').items()],

        }

2、PyQuery基本语法应用

#!/usr/bin/env python

# -*- encoding: utf-8 -*-

# Created on 2015-10-09 09:00:10

# Project: learn_pyspider_db

from pyspider.libs.base_handler import *

import re

class Handler(BaseHandler):

    crawl_config = {

    }

    @every(minutes=24 * 60)

    def on_start(self):

        self.crawl('http://movie.douban.com/tag/', callback=self.index_page)

    @config(age=10 * 24 * 60 * 60)

    def index_page(self, response):

        for each in response.doc('.tagCol a').items():

            self.crawl(each.attr.href, callback=self.list_tag_page)

    @config(age=10 * 24 * 60 * 60)

    def list_tag_page(self, response):

        all = response.doc('.more-links')

        self.crawl(all.attr.href, callback=self.list_page)

    @config(age=10*24*60*60)

    def list_page(self, response):

        for each in response.doc('.title').items():

            self.crawl(each.attr.href, callback=self.detail_page)

        # craw the next page

        next_page = response.doc('.next > a')

        self.crawl(next_page.attr.href, callback=self.list_page)

    @config(priority=2)

    def detail_page(self, response):

        director = ''

        for item in response.doc('#info > span > .pl'):

            if item.text == u'导演':

                next_item = item.getnext()

                director = next_item.getchildren()[0].text

                break

        return {

            "url": response.url,

            "title": response.doc('title').text(),

            "director": director,

        }

3、PyQuery复杂选择及翻页处理

#!/usr/bin/env python

# -*- encoding: utf-8 -*-

# Created on 2016-10-12 06:39:31

# Project: qx_zj_poi2

import re

from pyspider.libs.base_handler import *

class Handler(BaseHandler):

    crawl_config = {

    }

    @every(minutes=24 * 60)

    def on_start(self):

        self.crawl('http://www.poi86.com/poi/amap/province/330000.html', callback=self.city_page)

    @config(age=24 * 60 * 60)

    def city_page(self, response):

        for each in response.doc('body > div:nth-child(2) > div > div.panel-body > ul > li > a').items():

            self.crawl(each.attr.href, callback=self.district_page)

    @config(age=24 * 60 * 60)

    def district_page(self, response):

        for each in response.doc('body > div:nth-child(2) > div > div.panel-body > ul > li > a').items():

            self.crawl(each.attr.href, callback=self.poi_idx_page)

    @config(age=24 * 60 * 60)

    def poi_idx_page(self, response):

        for each in response.doc('td > a').items():

            self.crawl(each.attr.href, callback=self.poi_dtl_page)

        # 翻页

        for each in response.doc('body > div:nth-child(2) > div > div.panel-body > div > ul > li:nth-child(13) > a').items():

            self.crawl(each.attr.href, callback=self.poi_idx_page)

    @config(priority=100)

    def poi_dtl_page(self, response):

        return {

            "url": response.url,

            "id": re.findall('\d+',response.url)[1],

            "name": response.doc('body > div:nth-child(2) > div:nth-child(2) > div.panel-heading > h1').text(),

            "province": response.doc('body > div:nth-child(2) > div:nth-child(2) > div.panel-body > ul > li:nth-child(1) > a').text(),

            "city": response.doc('body > div:nth-child(2) > div:nth-child(2) > div.panel-body > ul > li:nth-child(2) > a').text(),

            "district": response.doc('body > div:nth-child(2) > div:nth-child(2) > div.panel-body > ul > li:nth-child(3) > a').text(),

            "addr": response.doc('body > div:nth-child(2) > div:nth-child(2) > div.panel-body > ul > li:nth-child(4)').text(),

            "tel": response.doc('body > div:nth-child(2) > div:nth-child(2) > div.panel-body > ul > li:nth-child(5)').text(),

            "type": response.doc('body > div:nth-child(2) > div:nth-child(2) > div.panel-body > ul > li:nth-child(6)').text(),

            "dd_map": response.doc('body > div:nth-child(2) > div:nth-child(2) > div.panel-body > ul > li:nth-child(7)').text(),

            "hx_map": response.doc('body > div:nth-child(2) > div:nth-child(2) > div.panel-body > ul > li:nth-child(8)').text(),

            "bd_map": response.doc('body > div:nth-child(2) > div:nth-child(2) > div.panel-body > ul > li:nth-child(9)').text(),

        }

pyspider用PyQuery解析页面数据的更多相关文章

pyspider示例代码三：用PyQuery解析页面数据
本系列文章主要记录和讲解pyspider的示例代码,希望能抛砖引玉.pyspider示例代码官方网站是http://demo.pyspider.org/.上面的示例代码太多,无从下手.因此本人找出一些 ...
python爬虫解析页面数据的三种方式
re模块 re.S表示匹配单行 re.M表示匹配多行使用re模块提取图片url,下载所有糗事百科中的图片普通版 import requests import re import os if not ...
pyspider示例代码：解析JSON数据
pyspider示例代码官方网站是http://demo.pyspider.org/.上面的示例代码太多,无从下手.因此本人找出一下比较经典的示例进行简单讲解,希望对新手有一些帮助. 示例说明: py ...
pyspider示例代码二：解析JSON数据
本系列文章主要记录和讲解pyspider的示例代码,希望能抛砖引玉.pyspider示例代码官方网站是http://demo.pyspider.org/.上面的示例代码太多,无从下手.因此本人找出一下 ...
python爬虫的页面数据解析和提取/xpath/bs4/jsonpath/正则(2)
上半部分内容链接 : https://www.cnblogs.com/lowmanisbusy/p/9069330.html 四.json和jsonpath的使用 JSON(JavaScript Ob ...
python爬虫的页面数据解析和提取/xpath/bs4/jsonpath/正则(1)
一.数据类型及解析方式一般来讲对我们而言,需要抓取的是某个网站或者某个应用的内容,提取有用的价值.内容一般分为两部分,非结构化的数据和结构化的数据. 非结构化数据:先有数据,再有结构, 结构化数 ...
JS解析Json 数据并跳转到一个新页面，取消A 标签跳转
JS解析Json 数据并跳转到一个新页面,代码如下 $.getJSON("http://api.cn.abb.com/common/api/staff/employee/" + o ...
python爬虫使用xpath解析页面和提取数据
XPath解析页面和提取数据一.简介关注公众号"轻松学编程"了解更多. XPath即为XML路径语言,它是一种用来确定XML(标准通用标记语言的子集)文档中某部分位置的语言.X ...
pytho爬虫使用bs4 解析页面和提取数据
页面解析和数据提取关注公众号"轻松学编程"了解更多. 一般来讲对我们而言,需要抓取的是某个网站或者某个应用的内容,提取有用的价值.内容一般分为两部分,非结构化的数据和结构化的 ...

随机推荐

官方：MySQL 5.7 并行复制实现原理与调优 | InsideMySQL（转载）
MySQL 5.7并行复制时代众所周知,MySQL的复制延迟是一直被诟病的问题之一,然而在Inside君之前的两篇博客中(1,2)中都已经提到了MySQL 5.7版本已经支持“真正”的并行复制功能, ...
html调bug
F12-->Sources-->相应文件-->找有波浪线
python学习之控制语句
#if statement number=int(input("please input a number")); if number<10 : print("is ...
【JavsScript高级程序设计】学习笔记[1]
1. 在HTML中使用javascript 在使用script嵌入脚本时,脚本内容不能包含'</script>'字符串(会被认为是结束的标志).可以通过转义字符解决('\') 默认scri ...
高级C/C++编译技术之读书笔记（一）之编译/链接
最近有幸阅读了<高级C/C++编译技术>深受启发,该书深入浅出地讲解了构建过程(编译.链接)中的各种细节,从多个角度展示了程序与库文件或代码的集成方法,提出了面向代码复用和系统集成的软件架 ...
Django之mysql表单操作
在Django之ORM模型中总结过django下mysql表的创建操作,接下来总结mysql表记录操作,包括表记录的增.删.改.查. 1. 添加表记录 class UserInfo(models.Mo ...
hibernate的级联（hibernate注解的CascadeType属性）
[自己项目遇到的问题]: 新增删除都可以实现 ,就是修改的时候无法同步更新设计三个类: 问题类scask 正文内容类text类查看数+回复数+讨论数的runinfo类 [正文类和查看数 ...
nmon的使用
Linux性能评测工具之一:nmon篇分类: 敏捷实践2010-06-08 11:27 7458人阅读评论(0) 收藏举报工具linuxfilesystemsaixx86excel 目录( ...
JavaScript define
1. AMD的由来前端技术虽然在不断发展之中,却一直没有质的飞跃.除了已有的各大著名框架,比如Dojo,jQuery,ExtJs等等,很多公司也都有着自己的前端开发框架.这些框架的使用效率以及开发质 ...
百度分享和bshare
社会法社交分享组件bshare http://www.bshare.cn/ 百度share也不错