python3 scrapy 爬取腾讯招聘

安装scrapy不再赘述，

在控制台中输入scrapy startproject tencent 创建爬虫项目名字为 tencent

接着cd tencent

用pycharm打开tencent项目

构建item文件

# -*- coding: utf-8 -*-

# Define here the models for your scraped items

#

# See documentation in:

# http://doc.scrapy.org/en/latest/topics/items.html

import scrapy

class TencentItem(scrapy.Item):

# define the fields for your item here like:

# name = scrapy.Field()

#职位名

positionname = scrapy.Field()

#详细链接

positionLink = scrapy.Field()

#职位类别

positionType = scrapy.Field()

#招聘人数

peopleNum = scrapy.Field()

#工作地点

workLocation = scrapy.Field()

#发布时间

publishTime = scrapy.Field()

　　接着在spiders文件夹中新建tencentPostition.py文件代码如下注释写的很清楚

# -*- coding: utf-8 -*-

import scrapy

from tencent.items import TencentItem

class TencentpostitionSpider(scrapy.Spider):

#爬虫名

name = 'tencent'

#爬虫域

allowed_domains = ['tencent.com']

#设置URL

url = 'http://hr.tencent.com/position.php?&start='

#设置页码

offset = 0

#默认url

start_urls = [url+str(offset)]

def parse(self, response):

#xpath匹配规则

for each in response.xpath("//tr[@class='even'] | //tr[@class='odd']"):

item = TencentItem()

# 职位名

item["positionname"] = each.xpath("./td[1]/a/text()").extract()[0]

# 详细链接

item["positionLink"] = each.xpath("./td[1]/a/@href").extract()[0]

# 职位类别

try:

item["positionType"] = each.xpath("./td[2]/text()").extract()[0]

except:

item["positionType"] = '空'

# 招聘人数

item["peopleNum"] = each.xpath("./td[3]/text()").extract()[0]

# 工作地点

item["workLocation"] = each.xpath("./td[4]/text()").extract()[0]

# 发布时间

item["publishTime"] = each.xpath("./td[5]/text()").extract()[0]

#把数据交给管道文件

yield item

#设置新URL页码

if(self.offset<2620):

self.offset += 10

#把请求交给控制器

yield scrapy.Request(self.url+str(self.offset),callback=self.parse)

　　接着配置管道文件pipelines.py代码如下

# -*- coding: utf-8 -*-

# Define your item pipelines here

#

# Don't forget to add your pipeline to the ITEM_PIPELINES setting

# See: http://doc.scrapy.org/en/latest/topics/item-pipeline.html

import json

class TencentPipeline(object):

def __init__(self):

#在初始化方法中打开文件

self.fileName = open("tencent.json","wb")

def process_item(self, item, spider):

#把数据转换为字典再转换成json

text = json.dumps(dict(item),ensure_ascii=False)+"\n"

#写到文件中编码设置为utf-8

self.fileName.write(text.encode("utf-8"))

#返回item

return item

def close_spider(self,spider):

#关闭时关闭文件

self.fileName.close()

　　接下来需要配置settings.py文件

不遵循ROBOTS规则

1	`ROBOTSTXT_OBEY` `=` `False`

1 2	`#下载延迟` `DOWNLOAD_DELAY` `=` `3`

#设置请求头

DEFAULT_REQUEST_HEADERS = {

'User-Agent':'Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/63.0.3239.84 Safari/537.36',

'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8',

}

#交给哪个管道文件处理文件夹.管道文件名.类名

ITEM_PIPELINES = {

'tencent.pipelines.TencentPipeline': 300,

}

　接下来再控制台中输入　

scrapy crawl tencent

即可爬取

源码地址

https://github.com/ingxx/scrapy_to_tencent　

python3 scrapy 爬取腾讯招聘的更多相关文章

简单的scrapy实战:爬取腾讯招聘北京地区的相关招聘信息
简单的scrapy实战:爬取腾讯招聘北京地区的相关招聘信息简单的scrapy实战:爬取腾讯招聘北京地区的相关招聘信息系统环境:Fedora22(昨天已安装scrapy环境) 爬取的开始URL:ht ...
利用scrapy爬取腾讯的招聘信息
利用scrapy框架抓取腾讯的招聘信息,爬取地址为:https://hr.tencent.com/position.php 抓取字段包括:招聘岗位,人数,工作地点,发布时间,及具体的工作要求和工作任务 ...
『Scrapy』爬取腾讯招聘网站
分析爬取对象初始网址, http://hr.tencent.com/position.php?@start=0&start=0#a (可选)由于含有多页数据,我们可以查看一下这些网址有什么相 ...
scrapy 第一个案例（爬取腾讯招聘职位信息）
import scrapy import json class TzcSpider(scrapy.Spider): # spider的名字,唯一 name = 'tzc' # 起始地址 start_u ...
python之scrapy爬取某集团招聘信息以及招聘详情
1.定义爬取的字段items.py # -*- coding: utf-8 -*- # Define here the models for your scraped items # # See do ...
Python 爬取腾讯招聘职位详情 2019/12/4有效
我爬取的是Python相关职位,先po上代码,(PS:本人小白,这是跟着B站教学视频学习后,老师留的作业,因为腾讯招聘的网站变动比较大,老师的代码已经无法运行,所以po上),一些想法和过程在后面. f ...
scrapy 爬取智联招聘
准备工作 1. scrapy startproject Jobs 2. cd Jobs 3. scrapy genspider ZhaopinSpider www.zhaopin.com 4. scr ...
利用Crawlspider爬取腾讯招聘数据(全站，深度)
需求: 使用crawlSpider(全站)进行数据爬取 - 首页: 岗位名称,岗位类别 - 详情页:岗位职责 - 持久化存储代码: 爬虫文件: from scrapy.linkextractors ...
python爬虫爬取腾讯招聘信息（静态爬虫）
环境: windows7,python3.4 代码:(亲测可正常执行) import requests from bs4 import BeautifulSoup from math import c ...

随机推荐

C简介与环境配置
C 语言是一种通用的高级语言,最初是由丹尼斯·里奇在贝尔实验室为开发 UNIX 操作系统而设计的.C 语言最开始是于 1972 年在 DEC PDP-11 计算机上被首次实现. 在 1978 年,布莱 ...
How can I list all foreign keys referencing a given table in SQL Server?
How can I list all foreign keys referencing a given table in SQL Server? how to check if columns in ...
【错误解决】SVN常见错误及解决方式
1.Error while creating module:org.apache.subversion.javahl.ClientException:Authorization failed svn: ...
java中年月日的加减法，年月的加减法使用
本文为博主原创,未经允许不得转载: java计算两个年月日之间相差的天数: public static int daysBetween(String smdate,String bdate) thro ...
UVa 1663 净化器
https://vjudge.net/problem/UVA-1663 题意: 给m个长度为n的模板串,每个模板串包含字符0,1和最多一个星号"*",其中星号可以匹配0或1.例如, ...
【Python】【正则】
1. 正则表达式基础 1.1. 简单介绍正则表达式并不是Python的一部分.正则表达式是用于处理字符串的强大工具,拥有自己独特的语法以及一个独立的处理引擎,效率上可能不如str自带的方法,但功能十 ...
python PIL 图像处理库简介(一)
1. Introduction PIL(Python Image Library)是python的第三方图像处理库,但是由于其强大的功能与众多的使用人数,几乎已经被认为是python官方图像处 ...
Codeforces Round #289 (Div. 2, ACM ICPC Rules) E. Pretty Song 算贡献+前缀和
E. Pretty Song time limit per test 1 second memory limit per test 256 megabytes input standard input ...
c语言数组合并
#include<stdio.h> int main() { int m,n,i,j,k; printf("Enter no. of elements in array1:\n& ...
HDU 6124 Euler theorem
Euler theorem 思路:找规律 a 余数个数 1 0 1 2 2 0 2 ...

python3 scrapy 爬取腾讯招聘

python3 scrapy 爬取腾讯招聘的更多相关文章

随机推荐

热门专题