scrapy入门实战-爬取代理网站

入门scrapy。

学习了有这几点

1.如何使用scrapy框架对网站进行爬虫；

2.如何对网页源代码使用xpath进行解析；

3.如何书写spider爬虫文件，对源代码进行解析；

4.学会使用scrapy的基础命令，创建项目，使用模板生成一个爬虫文件spider；

5,通过配置settings.py反爬虫。如设置user-agent；

设定目标:爬取网络代理www.xicidaili.com网站。

使用scrapy startproject 项目名称

scrapy startproject xicidailiSpider

项目名称应该如何命名呢：建议是需要爬虫的域名+Spider.举个例子：比如要爬取www.zhihu.com,那么项目名称可以写成zhihuSpider。

2. 目录中spiders放置的是爬虫文件，然后middlewares.py是中间件，有下载器的中间件，有爬虫文件的中间件。pipelines.py是管道文件，是对spider爬虫文件解析数据的处理。settings.py是设置相关属性，是否遵守爬虫的robotstxt协议，设置User-Agent等。

3.可以使用scrapy提供的模板，命令如下：

scrapy genspider 爬虫名字需要爬虫的网络域名

举例子：

我们需要爬取的www.xicidaili.com

那么可以使用

scarpy genspider xicidaili xicidaili.com

命令完成后，最终的目录如下：

建立后项目后，需要对提取的网页进行分析

经常使用的有三种解析模式：

1.正则表达式

2 xpath response.xpath("表达式")

3 css response.css("表达式")

XPath的语法是w3c的教程。http://www.w3school.com.cn/xpath/xpath_syntax.asp

需要安装一个xpath helper插件在浏览器中，可以帮助验证书写的xpath是否正确。

xpath语法需要多实践，看确实不容易记住。

xicidaili.py

# -*- coding: utf-8 -*-

import scrapy

# 继承scrapy,Spider类

class XicidailiSpider(scrapy.Spider):

    name = 'xicidaili'

    allowed_domains = ['xicidaili.com']

    start_urls = ['https://www.xicidaili.com/nn/',

                  "https://www.xicidaili.com/nt/",

                  "https://www.xicidaili.com/wn/,"

                  "https://www.xicidaili.com/wt/"]

    # 解析响应数据，提取数据和网址等。

    def parse(self, response):

        selectors = response.xpath('//tr')

        for selector in selectors:

            ip = selector.xpath("./td[2]/text()").get()

            port = selector.xpath("./td[3]/text()").get()     #.代表当前节点下

            country = selector.xpath("./td[4]/a/text()").get()   # get()和extract_first() 功能相同，getall()获取多个

            # print(ip,port,country)

            Items={

                "ip":ip,

                "port":port,

                "country":country

            }

            yield  Items

        """

        # 翻页操作

        # 获取下一页的标签

        next_page = response.xpath("//a[@class='next_page']/@href").get()

        # 判断next_page是否有值，也就是是否到了最后一页

        if next_page:

            # 拼接网页url---response.urljoin

            next_url = response.urljoin(next_page)

            # 判断最后一页是否

            yield  scrapy.Request(next_url,callback=self.parse)   # 回调函数不要加括号

    """

# -*- coding: utf-8 -*-

# settings.py设置

# Scrapy settings for xicidailiSpider project

#

# For simplicity, this file contains only settings considered important or

# commonly used. You can find more settings consulting the documentation:

#

#     https://doc.scrapy.org/en/latest/topics/settings.html

#     https://doc.scrapy.org/en/latest/topics/downloader-middleware.html

#     https://doc.scrapy.org/en/latest/topics/spider-middleware.html

BOT_NAME = 'xicidailiSpider'

SPIDER_MODULES = ['xicidailiSpider.spiders']

NEWSPIDER_MODULE = 'xicidailiSpider.spiders'

# 设置到处文件的字符编码

FEED_EXPORT_ENCODING ="UTF8"

# Crawl responsibly by identifying yourself (and your website) on the user-agent

#USER_AGENT = 'xicidailiSpider (+http://www.yourdomain.com)'

# Obey robots.txt rules

# 是否准售robots.txt协议，不遵守

ROBOTSTXT_OBEY = False

# Configure maximum concurrent requests performed by Scrapy (default: 16)

#CONCURRENT_REQUESTS = 32

# Configure a delay for requests for the same website (default: 0)

# See https://doc.scrapy.org/en/latest/topics/settings.html#download-delay

# See also autothrottle settings and docs

#DOWNLOAD_DELAY = 3

# The download delay setting will honor only one of:

#CONCURRENT_REQUESTS_PER_DOMAIN = 16

#CONCURRENT_REQUESTS_PER_IP = 16

# Disable cookies (enabled by default)

#COOKIES_ENABLED = False

# Disable Telnet Console (enabled by default)

#TELNETCONSOLE_ENABLED = False

# Override the default request headers:

DEFAULT_REQUEST_HEADERS = {

   'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8',

   'Accept-Language': 'en',

    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) \

    AppleWebKit/537.36 (KHTML, like Gecko) Chrome/75.0.3770.100 Safari/537.36'

}

# Enable or disable spider middlewares

# See https://doc.scrapy.org/en/latest/topics/spider-middleware.html

#SPIDER_MIDDLEWARES = {

#    'xicidailiSpider.middlewares.XicidailispiderSpiderMiddleware': 543,

#}

# Enable or disable downloader middlewares

# See https://doc.scrapy.org/en/latest/topics/downloader-middleware.html

#DOWNLOADER_MIDDLEWARES = {

#    'xicidailiSpider.middlewares.XicidailispiderDownloaderMiddleware': 543,

#}

# Enable or disable extensions

# See https://doc.scrapy.org/en/latest/topics/extensions.html

#EXTENSIONS = {

#    'scrapy.extensions.telnet.TelnetConsole': None,

#}

# Configure item pipelines

# See https://doc.scrapy.org/en/latest/topics/item-pipeline.html

#ITEM_PIPELINES = {

#    'xicidailiSpider.pipelines.XicidailispiderPipeline': 300,

#}

# Enable and configure the AutoThrottle extension (disabled by default)

# See https://doc.scrapy.org/en/latest/topics/autothrottle.html

#AUTOTHROTTLE_ENABLED = True

# The initial download delay

#AUTOTHROTTLE_START_DELAY = 5

# The maximum download delay to be set in case of high latencies

#AUTOTHROTTLE_MAX_DELAY = 60

# The average number of requests Scrapy should be sending in parallel to

# each remote server

#AUTOTHROTTLE_TARGET_CONCURRENCY = 1.0

# Enable showing throttling stats for every response received:

#AUTOTHROTTLE_DEBUG = False

# Enable and configure HTTP caching (disabled by default)

# See https://doc.scrapy.org/en/latest/topics/downloader-middleware.html#httpcache-middleware-settings

#HTTPCACHE_ENABLED = True

#HTTPCACHE_EXPIRATION_SECS = 0

#HTTPCACHE_DIR = 'httpcache'

#HTTPCACHE_IGNORE_HTTP_CODES = []

#HTTPCACHE_STORAGE = 'scrapy.extensions.httpcache.FilesystemCacheStorage'

　运行

scrapy crawl xicidai 项目名，这个必须唯一。

如果需要输出文件，

scarpy crawl xicidaili --output ip.json 或者ip.csv　

scrapy入门实战-爬取代理网站的更多相关文章

scrapy框架来爬取壁纸网站并将图片下载到本地文件中
首先需要确定要爬取的内容,所以第一步就应该是要确定要爬的字段: 首先去items中确定要爬的内容 class MeizhuoItem(scrapy.Item): # define the fields ...
Scrapy爬虫实战-爬取体彩排列5历史数据
网站地址:http://www.17500.cn/p5/all.php 1.新建爬虫项目 scrapy startproject pfive 2.在spiders目录下新建爬虫 scrapy gens ...
scrapy爬虫框架爬取招聘网站
目录结构 BossFace.py文件中代码: # -*- coding: utf-8 -*-import scrapyfrom ..items import BossfaceItemimport js ...
实战爬取某网站图片-Python
直接上代码 1 #!/usr/bin/python 2 # -*- coding: UTF-8 -*- 3 from bs4 import BeautifulSoup 4 import request ...
简单的scrapy实战:爬取腾讯招聘北京地区的相关招聘信息
简单的scrapy实战:爬取腾讯招聘北京地区的相关招聘信息简单的scrapy实战:爬取腾讯招聘北京地区的相关招聘信息系统环境:Fedora22(昨天已安装scrapy环境) 爬取的开始URL:ht ...
python爬虫-基础入门-爬取整个网站《3》
python爬虫-基础入门-爬取整个网站<3> 描述: 前两章粗略的讲述了python2.python3爬取整个网站,这章节简单的记录一下python2.python3的区别 python ...
python爬虫-基础入门-爬取整个网站《2》
python爬虫-基础入门-爬取整个网站<2> 描述: 开场白已在<python爬虫-基础入门-爬取整个网站<1>>中描述过了,这里不在描述,只附上 python3 ...
python爬虫-基础入门-爬取整个网站《1》
python爬虫-基础入门-爬取整个网站<1> 描述: 使用环境:python2.7.15 ,开发工具:pycharm,现爬取一个网站页面(http://www.baidu.com)所有数 ...
Python 网络爬虫 002 (入门) 爬取一个网站之前，要了解的知识
网站站点的背景调研 1. 检查 robots.txt 网站都会定义robots.txt 文件,这个文件就是给网络爬虫来了解爬取该网站时存在哪些限制.当然了,这个限制仅仅只是一个建议,你可以遵守,也 ...

随机推荐

SQL学习记录：定义(一)
--1.在这里@temp是一个表变量,只有一个批处理中有效,declare @temp table; --2. 如果前面加#就是临时表,可以在tempDB中查看到,它会在最后一个使用它的用户退出后才失 ...
C++ win32 dll 引用外部CLR，加载托管程序集异常-Error 10 error LNK2019: unresolved external symbol _CLRCreateInstancet
异常: Error 10 error LNK2019: unresolved external symbol _CLRCreateInstance@12 referenced in function ...
【翻译】Knowledge-Aware Natural Language Understanding（摘要及目录）
翻译Pradeep Dasigi的一篇长文 Knowledge-Aware Natural Language Understanding 基于知识感知的自然语言理解摘要 Natural Langua ...
QTP场景恢复函数
public Function RecoveryFunction1(Object, Method, Arguments, retVal) Dim FileName ,TimeNow, ResPath ...
QTP read or write XML file
'strNodePath = "/soapenv:Envelope/soapenv:Body/getProductsResponse/transaction/queryProducts/qu ...
“希希敬敬对”队软件工程第九次作业-beta冲刺第五次随笔
“希希敬敬对”队软件工程第九次作业-beta冲刺第五次随笔队名: “希希敬敬对” 龙江腾(队长) 201810775001 杨希 201810812008 何敬 ...
run （简单DP）
链接:https://www.nowcoder.com/acm/contest/140/A 来源:牛客网题目描述 White Cloud is exercising in the playgroun ...
强烈推荐一款功能强大的Tomcat 管理监控工具
专注于Java领域优质技术号,欢迎关注原创: 侯树成 Tomcat那些事儿启动 Tomcat完毕 ,有些时候总会打开浏览器 http://localhost:8080/ 去验证你的Tomcat是否 ...
js 页面跳转新窗口打开
页面跳转:Window.showModalDialog(url,width,height); 弹出一个html文档的模式对话框Parent.window.document.location.href ...
kubernetes容器集群管理启动一个测试示例
创建nginx 创建3个nginx副本 [root@master bin]# kubectl run nginx --image=nginx --replicas=3 kubectl run --ge ...

scrapy入门实战-爬取代理网站

scrapy入门实战-爬取代理网站的更多相关文章

随机推荐

热门专题