scrapy中Selector的使用
scrapy的Selector选择器其实也可以用来解析,今天主要总结下css和xpath的用法,其实我个人最喜欢用css
以慕课网嵩天老师教程中的一个网页为例,python123.io/ws/demo.html
解析是提取信息的一种手段,主要提取的信息包括:标签节点、属性、文本,下面从这三个方面来分别说明
一、提取标签节点
response = ”<html><head><title>This is a python demo page</title></head>
<body>
<p class="title"><b>The demo python introduces several python courses.</b></p>
<p class="course">Python is a wonderful general-purpose programming language. You can learn Python from novice to professional by tracking the following courses:
<a href="http://www.icourse163.org/course/BIT-268001" class="py1" id="link1">Basic Python</a> and <a href="http://www.icourse163.org/course/BIT-1001870001" class="py2" id="link2">Advanced Python</a>.</p>
</body></html>”
上面这个就是网页的html信息了,比如我要提取<p>标签
使用css选择器
selector = Selector(text=response)
p = selector.css('p').extract()
print(p)
#['<p class="title"><b>The demo python introduces several python courses.</b></p>', '<p class="course">Python is a wonderful general-purpose programming language. You can learn Python from novice to professional by tracking the following courses:\r\n<a href="http://www.icourse163.org/course/BIT-268001" class="py1" id="link1">Basic Python</a> and <a href="http://www.icourse163.org/course/BIT-1001870001" class="py2" id="link2">Advanced Python</a>.</p>']
这样就得到了所有p节点的信息,得到的是一个列表信息,如果只想得到第一个,实际上可以使用extract_first()方法,而不是使用extract()方法
对于简单的节点查找,这样就够了,但是如果同样的节点很多,而且我要查找的节点不在第一个,这样处理就不行。解决的方法是添加限制条件,添加class、id等等限制信息
比如我想提取class=course的p节点信息,使用p[class='course'],当然,如果有其他的属性,也可以用其他属性作为限定
selecor = Selector(text=result)
response = selecor.css('p[class="course"]').extract_first()
print(response) #<p class="course">Python is a wonderful general-purpose programming language. You can learn Python from novice to professional by tracking the following courses:<a href="http://www.icourse163.org/course/BIT-268001" class="py1" id="link1">Basic Python</a> and <a href="http://www.icourse163.org/course/BIT-1001870001" class="py2" id="link2">Advanced Python</a>.</p>
使用xpath
使用xpath大体思路也是一样的,只不过语法有点不同
使用xpath实现上述第一个例子
selecor = Selector(text=result)
response = selecor.xpath('//p').extract_first()
print(response)
使用xpath实现上述第二个例子
selecor = Selector(text=result)
response = selecor.xpath('//p[@class="course"]').extract_first()
print(response)
细心点的可能会发现xpath选取标签节点,就比css多了个//和@,//代表从当前节点进行选择,@后面接的是属性
二、提取属性
有时候我们需要提取属性值,比如src、href
response = ”<html><head><title>This is a python demo page</title></head>
<body>
<p class="title"><b>The demo python introduces several python courses.</b></p>
<p class="course">Python is a wonderful general-purpose programming language. You can learn Python from novice to professional by tracking the following courses:
<a href="http://www.icourse163.org/course/BIT-268001" class="py1" id="link1">Basic Python</a> and <a href="http://www.icourse163.org/course/BIT-1001870001" class="py2" id="link2">Advanced Python</a>.</p>
</body></html>”
还是这段例子,为了方便观看,我拷过来
比如我现在要提取第一个a标签的href
使用css
直接在标签后面加上::attr(href),attr代表提取的是属性,括号内的href代表我要提取的是哪种属性
selecor = Selector(text=result)
response = selecor.css('a::attr(href)').extract_first()
print(response)
#http://www.icourse163.org/course/BIT-268001
如果要提取特性的a标签的href属性,比如第二个a标签的href,同样可以使用限制条件
selecor = Selector(text=result)
response = selecor.css('a[class="py2"]::attr(href)').extract_first()
print(response)
#http://www.icourse163.org/course/BIT-1001870001
使用xpath
实现上面第一个例子
selecor = Selector(text=result)
response = selecor.xpath('//a/@href').extract_first()
print(response)
实现上面第二个例子
selecor = Selector(text=result)
response = selecor.xpath('//a[@class="py2"]/@href').extract_first()
print(response)
三、提取文本信息
response = ”<html><head><title>This is a python demo page</title></head>
<body>
<p class="title"><b>The demo python introduces several python courses.</b></p>
<p class="course">Python is a wonderful general-purpose programming language. You can learn Python from novice to professional by tracking the following courses:
<a href="http://www.icourse163.org/course/BIT-268001" class="py1" id="link1">Basic Python</a> and <a href="http://www.icourse163.org/course/BIT-1001870001" class="py2" id="link2">Advanced Python</a>.</p>
</body></html>”
提取第一个a标签的文本
使用css选择器
只需要在标签后面加上::text,至于怎么选择标签参照上面
selecor = Selector(text=result)
response = selecor.css('a::text').extract_first()
print(response)
#Basic Python
选择特定标签的文本,比如第二个a标签文本,同样是加一个限制条件就好
selecor = Selector(text=result)
response = selecor.css('a[class="py2"]::text').extract_first()
print(response)
#Advanced Python
使用xpath来实现
首先是第一个例子,使用//a选择到a节点,再/text()选择到文本信息
selecor = Selector(text=result)
response = selecor.xpath('//a/text()').extract_first()
print(response)
实现第二个例子,添加xpath限制条件的时候前面一定不要忘记加@,而且text后面要加()
selecor = Selector(text=result)
response = selecor.xpath('//a[@class="py2"]/text()').extract_first()
print(response)
最后总结下:对于提取而言,xpath多了/和@符号,即使在添加限制条件时,xpath也需要在限制的属性前加@,所以这也是我喜欢css的原因,因为我懒。
scrapy中Selector的使用的更多相关文章
- 18.scrapy中selector的用法
Selector是一个独立的模块. Selector主要是与scrapy结合使用的. 开启Scrapy shell: 1.打开命令行cmd 2.scrapy shell http://doc.scra ...
- 论Scrapy中的数据持久化
引入 Scrapy的数据持久化,主要包括存储到数据库.文件以及内置数据存储. 那我们今天就来讲讲如何把Scrapy中的数据存储到数据库和文件当中. 终端指令存储 保证爬虫文件的parse方法中有可迭代 ...
- 使用scrapy中xpath选择器的一个坑点
情景如下: 一个网页下有一个ul,这个ur下有125个li标签,每个li标签下有我们想要的 url 字段(每个 url 是唯一的)和 price 字段,我们现在要访问每个li下的url并在生成的请求中 ...
- scrapy框架Selector提取数据
从页面中提取数据的核心技术是HTTP文本解析,在python中常用的模块处理: BeautifulSoup 非常流行的解析库,API简单,但解析的速度慢. lxml 是一套使用c语言编写的xml解析 ...
- scrapy 中用selector来提取数据的用法
一. 基本概念 1. Selector是一个可独立使用的模块,我们可以用Selector类来构建一个选择器对象,然后调用它的相关方法如xpaht(), css()等来提取数据,如下 from sc ...
- scrapy中选择器用法
一.Selector选择器介绍 python从网页中提取数据常用以下两种方法: lxml:基于ElementTree的XML解析库(也可以解析HTML),不是python的标准库 BeautifulS ...
- Scrapy中get和extract_first的区别
在scrapy中,从xpath中取得selector对象后,需要取出需要的数据. 使用get以及getall获取的是带标签的数据 比如 <p>这是一段文字</p> 如果用get ...
- python的scrapy框架的使用 和xpath的使用 && scrapy中request和response的函数参数 && parse()函数运行机制
这篇博客主要是讲一下scrapy框架的使用,对于糗事百科爬取数据并未去专门处理 最后爬取的数据保存为json格式 一.先说一下pyharm怎么去看一些函数在源码中的代码实现 按着ctrl然后点击函数就 ...
- Scrapy中使用Django的Model访问数据库
Scrapy中使用Django的Model进行数据库访问 当已存在Django项目的时候,直接引入Django的Model来使用比较简单 # 使用以下语句添加Django项目的目录到path impo ...
随机推荐
- 【题解】HNOI2008GT考试
这题好难啊……完全不懂矩阵加速递推的我TAT 这道题目要求我们求出不含不吉利数字的字符串总数,那么我们有dp方程 : dp[i][j](长度为 i 的字符串,最长与不吉利数字前缀相同的后缀长度为 j ...
- cdh版本的zookeeper安装以及配置(伪分布式模式)
需要的软件包:zookeeper-3.4.5-cdh5.3.6.tar.gz 1.将软件包上传到Linux系统指定目录下: /opt/softwares/cdh 2.解压到指定的目录:/opt/mo ...
- [Leetcode] Path Sum II路径和
Given a binary tree and a sum, find all root-to-leaf paths where each path's sum equals the given su ...
- 【COGS 1873】 [国家集训队2011]happiness(吴确) 最小割
这是一种最小割模型,就是对称三角,中间双向边,我们必须满足其最小割就是满足题目条件的互斥关系的最小舍弃,在这道题里面我们S表示文T表示理,中间一排点是每个人,每个人向两边连其选文或者选理的价值,中间每 ...
- mybatis学习(七)——resultType解析
resultType是sql映射文件中定义返回值类型,返回值有基本类型,对象类型,List类型,Map类型等.现总结一下再解释 总结: resultType: 1.基本类型 :resultType= ...
- yum命令Header V3 RSA/SHA1 Signature, key ID c105b9de: NOKEY
yum命令Header V3 RSA/SHA1 Signature, key ID c105b9de: NOKEY 博客分类: linux 三种解决方案 我采取第三种方案解决的 第一种: linu ...
- init_connect基本用法
服务器为每个连接的客户端执行的字符串.字符串由一个或多个SQL语句组成.要想指定多个语句,用分号间隔开.例如,每个客户端开始时默认启用autocommit模式.没有全局服务器变量可以规定autocom ...
- 7月20号day12总结
今天学习过程和小结 先进行了复习,主要 1,hive导入数据的方式有 本地导入 load data [local] inpath 'hdfs-dir' into table tablename; s ...
- Ubuntu1604 install netease-cloud music
Two issue: 1. There is no voice on my computer, and the system was mute and cannot unmute. eric@E641 ...
- Python爬虫学习笔记之爬今日头条的街拍图片
代码: import requests import os from hashlib import md5 from urllib.parse import urlencode from multip ...