web爬虫，requests请求

requests请求，就是用yhthon的requests模块模拟浏览器请求，返回html源码

模拟浏览器请求有两种，一种是不需要用户登录或者验证的请求，一种是需要用户登录或者验证的请求

一、不需要用户登录或者验证的请求

这种比较简单，直接利用requests模块发一个请求即可拿到html源码

#!/usr/bin/env python

# -*- coding:utf8 -*-

import requests     #导入模拟浏览器请求模块

http =requests.get(url="http://www.iqiyi.com/")     #发送http请求

http.encoding = "utf-8"                             #http请求编码

neir = http.text                                    #获取http字符串代码

print(neir)

得到html源码

<!DOCTYPE html>

<html>

<head>

<title>抽屉新热榜-聚合每日热门、搞笑、有趣资讯</title>

        <meta charset="utf-8" />

        <meta name="keywords" content="抽屉新热榜,资讯,段子,图片,公众场合不宜,科技,新闻,节操,搞笑" />

        <meta name="description" content="

            抽屉新热榜，汇聚每日搞笑段子、热门图片、有趣新闻。它将微博、门户、社区、bbs、社交网站等海量内容聚合在一起，通过用户推荐生成最热榜单。看抽屉新热榜，每日热门、有趣资讯尽收眼底。

            " />

        <meta name="robots" content="index,follow" />

        <meta name="GOOGLEBOT" content="index,follow" />

        <meta name="Author" content="搞笑" />

        <meta http-equiv="X-UA-Compatible" content="IE=EmulateIE8">

        <link type="image/x-icon" href="/images/chouti.ico" rel="icon"/>

        <link type="image/x-icon" href="/images/chouti.ico" rel="Shortcut Icon"/>

        <link type="image/x-icon" href="/images/chouti.ico" rel="bookmark"/>

    <link type="application/opensearchdescription+xml"

          href="opensearch.xml" title="抽屉新热榜" rel="search" />

二、需要用户登录或者验证的请求

获取这种页面时，我们首先要了解整个登录过程，一般登录过程是，当用户第一次访问时，会自动在浏览器生成cookie文件，当用户输入登录信息后会携带着生成的cookie文件，如果登录信息正确会给这个cookie

授权，授权后以后访问需要登录的页面时携带授权后cookie即可

1、首先访问一下首页，然后查看是否有自动生成cookie

#!/usr/bin/env python

# -*- coding:utf8 -*-

import requests     #导入模拟浏览器请求模块

### 1、在没登录之前访问一下首页，获取cookie

i1 = requests.get(

    url="http://dig.chouti.com/",

    headers={'Referer': 'http://dig.chouti.com/'}

)

i1.encoding = "utf-8"                               #http请求编码

i1_cookie = i1.cookies.get_dict()

print(i1_cookie)                                    #返回获取到的cookie

#返回：{'JSESSIONID': 'aaaTztKP-KaGLbX-T6R0v', 'gpsd': 'c227f059746c839a28ab136060fe6ebe', 'route': 'f8b4f4a95eeeb2efcff5fd5e417b8319'}

可以看到生成了cookie，说明如果登陆信息正确，后台会给这里的cookie授权，以后访问需要登录的页面携带授权后的cookie即可

2、让程序自动去登录授权cookie

首先我们用浏览器访问登录页面，随便乱输入一下登录密码和账号，获取登录页面url，和登录所需要的字段

携带cookie登录授权

#!/usr/bin/env python

# -*- coding:utf8 -*-

import requests     #导入模拟浏览器请求模块

### 1、在没登录之前访问一下首页，获取cookie

i1 = requests.get(

    url="http://dig.chouti.com/",

    headers={'Referer':'http://dig.chouti.com/'}

)

i1.encoding = "utf-8"                               #http请求编码

i1_cookie = i1.cookies.get_dict()

print(i1_cookie)                                    #返回获取到的cookie

#返回：{'JSESSIONID': 'aaaTztKP-KaGLbX-T6R0v', 'gpsd': 'c227f059746c839a28ab136060fe6ebe', 'route': 'f8b4f4a95eeeb2efcff5fd5e417b8319'}

### 2、用户登陆，携带上一次的cookie，后台对cookie中的随机字符进行授权

i2 = requests.post(

    url="http://dig.chouti.com/login",              #登录url

    data={                                          #登录字段

        'phone': "8615284816568",

        'password': "279819",

        'oneMonth': ""

    },

    headers={'Referer':'http://dig.chouti.com/'},

    cookies=i1_cookie                               #携带cookie

)

i2.encoding = "utf-8"

dluxxi = i2.text

print(dluxxi)                                       #查看登录后服务器的响应

#返回：{"result":{"code":"9999", "message":"", "data":{"complateReg":"0","destJid":"cdu_50072007463"}}}  登录成功

3、登录成功后，说明后台已经给cookie授权，这样我们访问需要登录的页面时，携带这个cookie即可，比如获取个人中心

#!/usr/bin/env python

# -*- coding:utf8 -*-

import requests     #导入模拟浏览器请求模块

### 1、在没登录之前访问一下首页，获取cookie

i1 = requests.get(

    url="http://dig.chouti.com/",

    headers={'Referer':'http://dig.chouti.com/'}

)

i1.encoding = "utf-8"                               #http请求编码

i1_cookie = i1.cookies.get_dict()

print(i1_cookie)                                    #返回获取到的cookie

#返回：{'JSESSIONID': 'aaaTztKP-KaGLbX-T6R0v', 'gpsd': 'c227f059746c839a28ab136060fe6ebe', 'route': 'f8b4f4a95eeeb2efcff5fd5e417b8319'}

### 2、用户登陆，携带上一次的cookie，后台对cookie中的随机字符进行授权

i2 = requests.post(

    url="http://dig.chouti.com/login",              #登录url

    data={                                          #登录字段

        'phone': "8615284816568",

        'password': "279819",

        'oneMonth': ""

    },

    headers={'Referer':'http://dig.chouti.com/'},

    cookies=i1_cookie                               #携带cookie

)

i2.encoding = "utf-8"

dluxxi = i2.text

print(dluxxi)                                       #查看登录后服务器的响应

#返回：{"result":{"code":"9999", "message":"", "data":{"complateReg":"0","destJid":"cdu_50072007463"}}}  登录成功

### 3、访问需要登录才能查看的页面，携带着授权后的cookie访问

shouquan_cookie = i1_cookie

i3 = requests.get(

    url="http://dig.chouti.com/user/link/saved/1",

    headers={'Referer':'http://dig.chouti.com/'},

    cookies=shouquan_cookie                        #携带着授权后的cookie访问

)

i3.encoding = "utf-8"

print(i3.text)                                     #查看需要登录才能查看的页面

获取需要登录页面的html源码成功

全部代码

get()方法，发送get请求
encoding属性，设置请求编码
cookies.get_dict()获取cookies
post()发送post请求
text获取服务器响应信息

#!/usr/bin/env python

# -*- coding:utf8 -*-

import requests     #导入模拟浏览器请求模块

### 1、在没登录之前访问一下首页，获取cookie

i1 = requests.get(

    url="http://dig.chouti.com/",

    headers={'Referer':'http://dig.chouti.com/'}

)

i1.encoding = "utf-8"                               #http请求编码

i1_cookie = i1.cookies.get_dict()

print(i1_cookie)                                    #返回获取到的cookie

#返回：{'JSESSIONID': 'aaaTztKP-KaGLbX-T6R0v', 'gpsd': 'c227f059746c839a28ab136060fe6ebe', 'route': 'f8b4f4a95eeeb2efcff5fd5e417b8319'}

### 2、用户登陆，携带上一次的cookie，后台对cookie中的随机字符进行授权

i2 = requests.post(

    url="http://dig.chouti.com/login",              #登录url

    data={                                          #登录字段

        'phone': "8615284816568",

        'password': "279819",

        'oneMonth': ""

    },

    headers={'Referer':'http://dig.chouti.com/'},

    cookies=i1_cookie                               #携带cookie

)

i2.encoding = "utf-8"

dluxxi = i2.text

print(dluxxi)                                       #查看登录后服务器的响应

#返回：{"result":{"code":"9999", "message":"", "data":{"complateReg":"0","destJid":"cdu_50072007463"}}}  登录成功

### 3、访问需要登录才能查看的页面，携带着授权后的cookie访问

shouquan_cookie = i1_cookie

i3 = requests.get(

    url="http://dig.chouti.com/user/link/saved/1",

    headers={'Referer':'http://dig.chouti.com/'},

    cookies=shouquan_cookie                        #携带着授权后的cookie访问

)

i3.encoding = "utf-8"

print(i3.text)                                     #查看需要登录才能查看的页面

注意：如果登录需要验证码，那就需要做图像处理，根据验证码图片，识别出验证码，将验证码写入登录字段

web爬虫，requests请求的更多相关文章

Python爬虫requests请求库
requests:pip install request 安装实例: import requestsurl = 'http://www.baidu.com'response = requests. ...
第三百二十二节，web爬虫，requests请求
第三百二十二节,web爬虫,requests请求 requests请求,就是用yhthon的requests模块模拟浏览器请求,返回html源码模拟浏览器请求有两种,一种是不需要用户登录或者验证的请 ...
一 web爬虫，requests请求
requests请求,就是用python的requests模块模拟浏览器请求,返回html源码模拟浏览器请求有两种,一种是不需要用户登录或者验证的请求,一种是需要用户登录或者验证的请求一.不需要用 ...
1、web爬虫，requests请求
requests请求,就是用python的requests模块模拟浏览器请求,返回html源码模拟浏览器请求有两种,一种是不需要用户登录或者验证的请求,一种是需要用户登录或者验证的请求一.不需要用 ...
爬虫 Http请求,urllib2获取数据,第三方库requests获取数据,BeautifulSoup处理数据,使用Chrome浏览器开发者工具显示检查网页源代码,json模块的dumps，loads，dump，load方法介绍
爬虫 Http请求,urllib2获取数据,第三方库requests获取数据,BeautifulSoup处理数据,使用Chrome浏览器开发者工具显示检查网页源代码,json模块的dumps,load ...
第三百二十七节，web爬虫讲解2—urllib库爬虫—基础使用—超时设置—自动模拟http请求
第三百二十七节,web爬虫讲解2—urllib库爬虫利用python系统自带的urllib库写简单爬虫 urlopen()获取一个URL的html源码read()读出html源码内容decode(& ...
解决爬虫浏览器中General显示 Status Code:304 NOT MODIFIED，而在requests请求时出现403被拦截的情况。
在此,非常感谢 “完美风暴4” 的无私共享经验的精神在Python爬虫爬取网站时,莫名遇到浏览器中General显示 Status Code: 304 NOT MODIFIED 而在req ...
爬虫（一）—— 请求库（一）requests请求库
目录 requests请求库爬虫:爬取.解析.存储一.请求二.响应三.简单爬虫四.requests高级用法五.session方法(建议使用) 六.selenium模块 requests请求 ...
爬虫 http原理,梨视频,github登陆实例,requests请求参数小总结
回顾:http协议基于请求响应的方式,请求:请求首行请求头{'keys':vales} 请求体 :响应:响应首行,响应头{'keys':'vales'},响应体. import socket soc ...

随机推荐

erlang并发编程（二）
补充-------erlang并发编程 Pid =spawn(fun()-> do_sth() end). 进程监视: Ref = monitor(process, Pid)靠抛异常来终结进程 ...
nodejs-POST数据处理
GET数据:容量小 32K 数据在URL中 POST数据:数据量大 1G 分段传输数据另外发处理方法: const http=require("http"); http.cre ...
nodejs03-GET数据处理
数据请求:--- 前台:form ajax jsonp 后台:一样请求方式: 1.GET 数据在URL中 2.POST 数据在请求体中请求数据组成: 头--header:url,头信息身子--c ...
BluePrism初尝2
在接近三周的自学中,初步体验到了RPA的甜头. 在对BP这个工具慢慢的深入接触中,从0 到1的探索式学习,从最开始的一个个的小功能模块的用途,每一个的属性的功能,到现在自己能初步尝试组织一些简单的流程 ...
hive lock命令的使用
1.hive锁表命令 hive> lock table t1 exclusive;锁表后不能对表进行操作 2.hive表解锁: hive> unlock table t1; 3.查看被锁的 ...
java富文本编辑器KindEditor
在页面写一个编辑框: <textarea name="content" class="form-control" id="content&quo ...
L342 Air Pollution Is Doing More Than Just Slowly Killing Us
Air Pollution Is Doing More Than Just Slowly Killing Us In the future, the authorities might need to ...
UITableView section 圆角阴影
在UITableView实现图片上面的效果,百度一下看了别人的实现方案有下面2种: 1.UITableView section里面嵌套UITableView然后在上面实现圆角和阴影, 弊端代码超 ...
border-radius,box-shadow兼容性解决办法
css3 border-radius不支持IE8/IE7的四种解决方法标签: cssborder-radius兼容性时间:2016-07-18 css3 border-radius用于设置HT ...
excel2013 打开为灰色空白左下角显示就绪要把文件拖进去才能打开！
最近电脑excel2013 打开总是为灰色空白左下角显示就绪要把文件拖进去或者在此再打开一个才能打开! 在网上搜了一下,我是使用下面这个方法解决的: 步骤一:请您在“开始”菜单的搜索框中输入“re ...

web爬虫，requests请求

web爬虫，requests请求的更多相关文章

随机推荐

热门专题