Tencent 招聘信息网站

创建项目

scrapy startproject Tencent

创建爬虫

scrapy  genspider -t crawl tencent

1. 起始url start_url = 'https://hr.tencent.com/position.php'

在起始页面，需要获取该也页面上的每个职位的详情页的url,同时需要提取下一页的url地址，做同样的操作。

因此起始页url地址的提取，分为两类：

　　1. 每个职位详情页的url地址的提取

　　2. 下一页url地址的提取，并且得到的页面做的操作和起始页的操作一样。

url地址的提取

1. 提取详情页url，详情页的url地址如下：

提取规则详情页的规则：

rules = (

        # 提取详情页的url地址  ，详情页url地址对应的响应，需要进行数据提取，所有需要有回调函数，用来解析数据

        Rule(LinkExtractor(restrict_xpaths=("//table[@class='tablelist']//td[@class='l square']")), callback='parse_item')

    )

提取下一页的htmlj所在的位置：

2 获取下一页的url 规则：

rules = (

        # 提取详情页的url地址

        # Rule(LinkExtractor(allow=r'position_detail.php?id=\d+\&keywords=&tid=0&lid=0'), callback='parse_item'), # 这个表达式有错，这里不用正则

        Rule(LinkExtractor(restrict_xpaths=("//table[@class='tablelist']//td[@class='l square']")), callback='parse_item'),

        # 翻页

        Rule(LinkExtractor(restrict_xpaths=("//a[@id='next']")), follow=True),

    )

获取详情页数据

1.详情数据提取(爬虫逻辑)

1.获取标题

xpath:

item['title'] = response.xpath('//td[@id="sharetitle"]/text()').extract_first()

2. 获取工作地点，职位，招聘人数

xpath:

 item['addr'] = response.xpath('//tr[@class="c bottomline"]/td[1]//text()').extract()[1]

 item['position'] = response.xpath('//tr[@class="c bottomline"]/td[2]//text()').extract()[1]

 item['num'] = response.xpath('//tr[@class="c bottomline"]/td[3]//text()').extract()[1]

3.工作要求抓取

xpath:

item['skill'] =response.xpath('//ul[@class="squareli"]/li/text()').extract()

爬虫的代码：

# -*- coding: utf-8 -*-

import scrapy

from scrapy.linkextractors import LinkExtractor

from scrapy.spiders import CrawlSpider, Rule

from ..items import TencentItem

class TencentSpider(CrawlSpider):

    name = 'tencent'

    allowed_domains = ['hr.tencent.com']

    start_urls = ['https://hr.tencent.com/position.php']

    rules = (

        # 提取详情页的url地址

        # Rule(LinkExtractor(allow=r'position_detail.php?id=\d+\&keywords=&tid=0&lid=0'), callback='parse_item'), # 这个表达式有错

        Rule(LinkExtractor(restrict_xpaths=("//table[@class='tablelist']//td[@class='l square']")), callback='parse_item'),

        # 翻页

        Rule(LinkExtractor(restrict_xpaths=("//a[@id='next']")), follow=True),

    )

    def parse_item(self, response):

        item = TencentItem()

        item['title'] = response.xpath('//td[@id="sharetitle"]/text()').extract_first()

        item['addr'] = response.xpath('//tr[@class="c bottomline"]/td[1]//text()').extract()[0]

        item['position'] = response.xpath('//tr[@class="c bottomline"]/td[2]//text()').extract()[0]

        item['num'] = response.xpath('//tr[@class="c bottomline"]/td[3]//text()').extract()[0]

        item['skill'] =response.xpath('//ul[@class="squareli"]/li/text()').extract()

        print(dict(item))

        return item

tencent.py

2. 数据存储

1.settings.py 配置文件，配置如下信息

ROBOTSTXT_OBEY = False

USER_AGENT = 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_13_2) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/63.0.3239.132 Safari/537.36'

ITEM_PIPELINES = {

   'jd.pipelines.TencentPipeline': 300,

}

2. items.py 中：

import scrapy

class TencentItem(scrapy.Item):

    # define the fields for your item here like:

    title = scrapy.Field()

    addr = scrapy.Field()

    position = scrapy.Field()

    num = scrapy.Field()

    skill = scrapy.Field()

3. pipeline.py中：

import  pymongo

class TencentPipeline(object):

    def open_spider(self,spider):

        # 爬虫开启是连接数据库

        client = pymongo.MongoClient()

        collention = client.tencent.ten

        self.client =client

        self.collention = collention

        pass

    def process_item(self, item, spider):

        # 数据保存在mongodb 中

        self.collention.insert(dict(item))

        return item

    def colse_spdier(self,spider):

        # 爬虫结束，关闭数据库

        self.client.close()

启动项目

1.先将MongoDB数据库跑起来。

2.执行爬虫命令：

scrapy  crawl  tencent

3. 执行程序后的效果：

使用scrapy-crawlSpider 爬取tencent 招聘的更多相关文章

Scrapy框架——CrawlSpider爬取某招聘信息网站
CrawlSpider Scrapy框架中分两类爬虫,Spider类和CrawlSpider类. 它是Spider的派生类,Spider类的设计原则是只爬取start_url列表中的网页, 而Craw ...
Python爬虫【实战篇】scrapy 框架爬取某招聘网存入mongodb
创建项目 scrapy startproject zhaoping 创建爬虫 cd zhaoping scrapy genspider hr zhaopingwang.com 目录结构 items.p ...
Python+Scrapy+Crawlspider 爬取数据且存入MySQL数据库
1.Scrapy使用流程 1-1.使用Terminal终端创建工程,输入指令:scrapy startproject ProName 1-2.进入工程目录:cd ProName 1-3.创建爬虫文件( ...
简单的scrapy实战:爬取腾讯招聘北京地区的相关招聘信息
简单的scrapy实战:爬取腾讯招聘北京地区的相关招聘信息简单的scrapy实战:爬取腾讯招聘北京地区的相关招聘信息系统环境:Fedora22(昨天已安装scrapy环境) 爬取的开始URL:ht ...
爬虫07 /scrapy图片爬取、中间件、selenium在scrapy中的应用、CrawlSpider、分布式、增量式
爬虫07 /scrapy图片爬取.中间件.selenium在scrapy中的应用.CrawlSpider.分布式.增量式目录爬虫07 /scrapy图片爬取.中间件.selenium在scrapy ...
scrapy-redis + Bloom Filter分布式爬取tencent社招信息
scrapy-redis + Bloom Filter分布式爬取tencent社招信息什么是scrapy-redis 什么是 Bloom Filter 为什么需要使用scrapy-redis + B ...
scrapy-redis分布式爬取tencent社招信息
scrapy-redis分布式爬取tencent社招信息什么是scrapy-redis 目标任务安装爬虫创建爬虫编写 items.py 编写 spiders/tencent.py 编写 pip ...
python-scrapy爬取某招聘网站(二)
首先要准备python3+scrapy+pycharm 一.首先让我们了解一下网站拉勾网https://www.lagou.com/ 和Boss直聘类似的网址设计方式,与智联招聘不同,它采用普通的页 ...
使用scrapy框架爬取自己的博文（2）
之前写了一篇用scrapy框架爬取自己博文的博客,后来发现对于中文的处理一直有问题- - 显示的时候 [u'python\u4e0b\u722c\u67d0\u4e2a\u7f51\u9875\u76 ...

随机推荐

linux内核中的DMI是什么?
答: 桌面管理接口(Desktop Management Interface).是用来获取硬件信息的,在内核中有一个配置项CONFIG_DMI用来添加此功能到内核中!
Jenkins serving Cake: our recipe for Windows
https://novemberfive.co/blog/windows-jenkins-cake-tutorial/ Where we started, or: why Cake took the ...
【索引失效】什么情况下会引起MySQL索引失效
索引并不是时时都会生效的,比如以下几种情况,将导致索引失效: 1.如果条件中有or,即使其中有条件带索引也不会使用(这也是为什么尽量少用or的原因) 注意:要想使用or,又想让索引生效,只能将or条件 ...
fhqtreap初探
介绍 fhqtreap为利用分裂和合并来满足平衡树的性质,不需要旋转操作的一种平衡树. 并且利用函数式编程可以极大的简化代码量. (题目是抄唐神的来着) 核心操作 (均为按位置分裂合并) struct ...
Thread类的常用方法
String getName() 返回该线程的名称. void setName(String name) 改变线程名称,使之与参数 name 相同. int getPriority() 返回线程的优先 ...
hihoCoder 1233 : Boxes（盒子）
hihoCoder #1233 : Boxes(盒子) 时间限制:1000ms 单点时限:1000ms 内存限制:256MB Description - 题目描述 There is a strange ...
【译】第15节---数据注解-StringLength
原文:http://www.entityframeworktutorial.net/code-first/stringlength-dataannotations-attribute-in-code- ...
oracle 与其他数据库如mysql的区别
想明白一个问题:(1)oracle是以数据库为中心,一个数据库就是一个域(可以看作是一个文件夹的概念),一个数据库可以有多个用户,创建用户是在登陆数据库之后进行的,但是有表空间的概念(2)而mysql ...
C++ 空字符('\0')和空格符(' ')
1.从字符串的长度:-->空字符的长度为0,空格符的长度为1. 2.虽然输出到屏幕是一样的,但是本质的ascii code 是不一样的,他们还是有区别的. #include<iostrea ...
烽火HG220G-U E00L2.03M2000光猫改桥接教程
烽火HG220G-U E00L2.03M2000光猫改桥接教程 P.S. 此教程同样适用于HG221G/HG260G-U/HG261G.(2016.12) 随着北京联通从原有的ONU升级到HGU之后, ...

使用scrapy-crawlSpider 爬取tencent 招聘

Tencent 招聘信息网站

url地址的提取

1. 提取详情页url，详情页的url地址如下：

2 获取下一页的url 规则：

获取详情页数据

1.详情数据提取(爬虫逻辑)

2. 数据存储

1.settings.py 配置文件，配置如下信息

2. items.py 中：

3. pipeline.py中：

启动项目

使用scrapy-crawlSpider 爬取tencent 招聘的更多相关文章

随机推荐

热门专题