python爬虫学习之使用XPath解析开奖网站

实例需求：运用python语言爬取http://kaijiang.zhcw.com/zhcw/html/ssq/list_1.html这个开奖网站所有的信息，并且保存为txt文件。

实例环境：python3.7
　　　　　 BeautifulSoup库、XPath(需手动安装)
　　　　　 urllib库(内置的python库，无需手动安装)

实例网站：

　　第一步，点击链接http://kaijiang.zhcw.com/zhcw/html/ssq/list_1.html进入网站，查看网站基本信息，注意一共要爬取118页数据。

　　第二步，查看网页源代码，熟悉网页结构，标签等信息。

实例代码：

#encoding=utf-8

#pip install lxml

from bs4 import BeautifulSoup

import urllib.request

from lxml import etree

class GetDoubleColorBallNumber(object):

    def __init__(self):

        self.urls = []

        self.getUrls()

        self.items = self.spider(self.urls)

        self.pipelines(self.items)

    def getUrls(self):

        URL = r'http://kaijiang.zhcw.com/zhcw/html/ssq/list.html'

        htmlContent = self.getResponseContent(URL)

        soup = BeautifulSoup(htmlContent, 'html.parser')

        tag = soup.find_all('p')[-1]

        pages = tag.strong.get_text()

        pages = '3'

        for i in range(2, int(pages)+1):

            url = r'http://kaijiang.zhcw.com/zhcw/html/ssq/list_' + str(i) + '.html'

            self.urls.append(url)

    #3、	网络模块（NETWORK）

    def getResponseContent(self, url):

        try:

            response = urllib.request.urlopen(url)

        except urllib.request.URLError as e:

            raise e

        else:

            return response.read().decode("utf-8")

    #3、爬虫模块（Spider）

    def spider(self,urls):

        items = []

        for url in urls:

            try:

                html = self.getResponseContent(url)

                xpath_tree = etree.HTML(html)

                trTags = xpath_tree.xpath('//tr[not(@*)]')   # 匹配所有tr下没有任何属性的节点

                for tag in trTags:

                    # if tag.xpath('../html'):

                    #     print("找到了html标签")

                    # if tag.xpath('/td/em'):

                    #     print("****************")

                    #如果存在em子孙节点

                    if tag.xpath('./td/em'):

                        item = {}

                        item['date'] = tag.xpath('./td[1]/text()')[0]

                        item['order'] = tag.xpath('./td[2]/text()')[0]

                        item['red1'] = tag.xpath('./td[3]/em[1]/text()')[0]

                        item['red2'] = tag.xpath('./td[3]/em[2]/text()')[0]

                        item['red3'] = tag.xpath('./td[3]/em[3]/text()')[0]

                        item['red4'] = tag.xpath('./td[3]/em[4]/text()')[0]

                        item['red5'] = tag.xpath('./td[3]/em[5]/text()')[0]

                        item['red6'] = tag.xpath('./td[3]/em[6]/text()')[0]

                        item['blue'] = tag.xpath('./td[3]/em[7]/text()')[0]

                        item['money'] = tag.xpath('./td[4]/strong/text()')[0]

                        item['first'] = tag.xpath('./td[5]/strong/text()')[0]

                        item['second'] = tag.xpath('./td[6]/strong/text()')[0]

                        items.append(item)

            except Exception as e:

                print(str(e))

                raise e

        return items

    def pipelines(self,items):

        fileName = u'双色球.txt'

        with open(fileName, 'w') as fp:

            for item in items:

                fp.write('%s %s \t %s %s %s %s %s %s  %s \t %s \t %s %s \n'

                      %(item['date'],item['order'],item['red1'],item['red2'],item['red3'],item['red4'],item['red5'],item['red6'],item['blue'],item['money'],item['first'],item['second']))

if __name__ == '__main__':

    GDCBN = GetDoubleColorBallNumber()

实例结果：

python爬虫学习之使用XPath解析开奖网站的更多相关文章

Python爬虫学习之使用beautifulsoup爬取招聘网站信息
菜鸟一只,也是在尝试并学习和摸索爬虫相关知识. 1.首先分析要爬取页面结构.可以看到一列搜索的结果,现在需要得到每一个链接,然后才能爬取对应页面. 关键代码思路如下: html = getHtml(& ...
python爬虫学习之使用BeautifulSoup库爬取开奖网站信息-模块化
实例需求:运用python语言爬取http://kaijiang.zhcw.com/zhcw/html/ssq/list_1.html这个开奖网站所有的信息,并且保存为txt文件和excel文件. 实 ...
python爬虫学习笔记（一）——环境配置（windows系统）
在进行python爬虫学习前,需要进行如下准备工作: python3+pip官方配置 1.Anaconda(推荐,包括python和相关库) [推荐地址:清华镜像] https://mirrors ...
Python爬虫学习：三、爬虫的基本操作流程
本文是博主原创随笔,转载时请注明出处Maple2cat|Python爬虫学习:三.爬虫的基本操作与流程一般我们使用Python爬虫都是希望实现一套完整的功能,如下: 1.爬虫目标数据.信息: 2.将 ...
Python爬虫教程-22-lxml-etree和xpath配合使用
Python爬虫教程-22-lxml-etree和xpath配合使用 lxml:python 的HTML/XML的解析器官网文档:https://lxml.de/ 使用前,需要安装安 lxml 包 ...
小白学 Python 爬虫（21）：解析库 Beautiful Soup（上）
小白学 Python 爬虫(21):解析库 Beautiful Soup(上) 人生苦短,我用 Python 前文传送门: 小白学 Python 爬虫(1):开篇小白学 Python 爬虫(2):前 ...
小白学 Python 爬虫（22）：解析库 Beautiful Soup（下）
人生苦短,我用 Python 前文传送门: 小白学 Python 爬虫(1):开篇小白学 Python 爬虫(2):前置准备(一)基本类库的安装小白学 Python 爬虫(3):前置准备(二)Li ...
小白学 Python 爬虫（23）：解析库 pyquery 入门
人生苦短,我用 Python 前文传送门: 小白学 Python 爬虫(1):开篇小白学 Python 爬虫(2):前置准备(一)基本类库的安装小白学 Python 爬虫(3):前置准备(二)Li ...
python爬虫学习01--电子书爬取
python爬虫学习01--电子书爬取 1.获取网页信息 import requests #导入requests库 ''' 获取网页信息 ''' if __name__ == '__main__': ...

随机推荐

Smartforms
Include text Populate indicator in program perform get_text using '0002' ls_detail-vbeln"Header ...
Mapbox Studio Classic 闪退问题解决方案
之前安装过Mapbox Studio Classic 0.38,好久没有用了,今天用的时候发现不停的闪退,经过一番折腾,发现删除 %USERPROFILE%\.mapbox-studio 目录下所有文 ...
Git多账号配置，同一电脑多个ssh-key的管理
为什么有这种需求? 在我们开发过程中,可能会遇到使用同一台机器,既要向公司git服务器提交代码,也要向gitlib或者gitee等 git仓库提交代码,2个仓库设置的用户名信息,不一样,此时需要用到多 ...
MVC htmlAttributes and additionalViewData
@Html.TextBoxFor(m => m.UserName, new { title = "ABC" }) // 输出结果为 <input data-val=&q ...
《C++实践之路.pdf》源码
> 源码下载方法 < >> 打开微信 >> 扫描下方二维码 >> 关注林哥私房菜 >> 输入对应编号获取百度网盘提取密码全书源码[已更新完 ...
用turtle库实现汉诺塔问题~~~~~
汉诺塔问题问题描述和背景: 汉诺塔是学习"递归"的经典入门案例,该案例来源于真实故事.‪‬‪‬‪‬‪‬‪‬‮‬‪‬‫‬‪‬‪‬‪‬‪‬‪‬‮‬‪‬‭‬‪‬‪‬‪‬‪‬‪‬‮‬‪‬ ...
主要的Ajax框架都有什么?
* Dojo(dojotoolkit.org):* Prototype和Scriptaculous (www.prototypejs.org和script.aculo.us):* Direct We ...
linux中tomcat startup.sh出现commond not found
问题: 前些天,再Linux提交更新代码启动tomcat时报commond not found 过程: 查了下百度,http://code2care.org/2015/-bash:-startup.s ...
Filter笔记
1.Filter [1] Filter简介 > Filter翻译为中文是过滤器的意思. > Filter是JavaWeb的三大web组件之一:Servlet.Filter.Listener ...
es6数组
将两类对象转为真正的数组 Array.from方法用于将两类对象转为真正的数组:类似数组的对象(array-like object)和可遍历(iterable)的对象(包括ES6新增的数据结构Set和 ...

python爬虫学习之使用XPath解析开奖网站

python爬虫学习之使用XPath解析开奖网站的更多相关文章

随机推荐

热门专题