python爬虫入门（1）----- requests

介绍

requests是python实现的简单易用的HTTP库，使用起来比urllib简洁很多

基本使用

requests.get("http://www.baidu.com")

requests.post("http://www.baidu.com")

requests.put("http://www.baidu.com")

requests.delete("http://www.baidu.com")

requests.request("get", "http://www.baidu.com")

get

def get(url, params=None, **kwargs):

        r"""Sends a GET request.

        :param url: URL for the new :class:`Request` object.

        :param params: (optional) Dictionary, list of tuples or bytes to send

            in the body of the :class:`Request`.

        :param \*\*kwargs: Optional arguments that ``request`` takes.

        :return: :class:`Response <Response>` object

        :rtype: requests.Response

        """

        kwargs.setdefault('allow_redirects', True)

        return request('get', url, params=params, **kwargs)

下面凡科微传单获取模板的接口为例子

 import requests

    param = {

    "cmd": "getTemplate"，

    "scrollIndex": 0

    }

    header = {

    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/70.0.3538.77 Safari/537.36"

    }//通过ua识别是否是爬虫

    rep = requests.get("https://cd.fkw.com/ajax/flyerhome.jsp", params=param, headers=header)

    rep.encoding = 'utf8'

    print(rep.text)

post

def post(url, data=None, json=None, **kwargs):

        r"""Sends a POST request.

        :param url: URL for the new :class:`Request` object.

        :param data: (optional) Dictionary, list of tuples, bytes, or file-like

            object to send in the body of the :class:`Request`.

        :param json: (optional) json data to send in the body of the :class:`Request`.

        :param \*\*kwargs: Optional arguments that ``request`` takes.

        :return: :class:`Response <Response>` object

        :rtype: requests.Response

        """

        return request('post', url, data=data, json=json, **kwargs)

一样以凡科微传单接口为例

 import requests

    data = {

    "cmd": "getTemplate"，

    "scrollIndex": 0

    }

    header = {

    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/70.0.3538.77 Safari/537.36"

    }

    rep = requests.post("https://cd.fkw.com/ajax/flyerhome.jsp", data=data, headers=header)

    rep.encoding = 'utf8'

    print(rep.text)

会话对象

在上面操作中request不会持有cookie对象导致每次请求都是新的会话，requests库提供了session的解决方案，下面以凡科登录和登录状态下获取模板为例

import requests

    import _md5

    import json

    import re

    s = requests.session()

    md5 = _md5.md5()

    md5.update("pwd".encode("utf8"))

    pwd = md5.hexdigest()

    data = {

    "cmd": "loginCorpNew",

    "cacct": "username",

    "pwd": pwd

    }

    header = {

    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/70.0.3538.77 Safari/537.36"

    }

    rep = s.post("https://i.fkw.com/ajax/login_h.jsp?dogSrc=3", data=data, headers=header)

    login = json.loads(rep.text)

    tokenStr = login.get("_TOKEN")

    print(tokenStr)

    pattern = "value='(.+)'"

    matcher = re.search(pattern, rep.text)

    if matcher:

        token = matcher.group(1)

        print(token)

        param = {

        "cmd": "getTemplate",

        "_TOKEN": token,

        "scrollIndex": 0

        }

        rep = s.get("https://i.cd.fkw.com/ajax/flyerTemplate_h.jsp", params=param, headers=header)

        print(rep.text)

参考文献

https://cuiqingcai.com/2556.html
http://docs.python-requests.org/en/master/api/

python爬虫入门（1）----- requests的更多相关文章

Python 爬虫入门（requests）
相信最开始接触Python爬虫学习的同学最初大多使用的是urllib,urllib2.在那之后接触到了第三方库requests,requests完全能满足各种http功能,真的是好用爆了 :D 他们是 ...
Python爬虫入门——使用requests爬取python岗位招聘数据
爬虫目的使用requests库和BeautifulSoup4库来爬取拉勾网Python相关岗位数据爬虫工具使用Requests库发送http请求,然后用BeautifulSoup库解析HTML文 ...
Python爬虫入门（二）之Requests库
Python爬虫入门(二)之Requests库我是照着小白教程做的,所以该篇是更小白教程hhhhhhhh 一.Requests库的简介 Requests 唯一的一个非转基因的 Python HTTP ...
python爬虫入门-开发环境与小例子
python爬虫入门开发环境 ubuntu 16.04 sublime pycharm requests库 requests库安装: sudo pip install requests 第一个例子 ...
Python 爬虫入门(二)——爬取妹子图
Python 爬虫入门听说你写代码没动力?本文就给你动力,爬取妹子图.如果这也没动力那就没救了. GitHub 地址: https://github.com/injetlee/Python/blob ...
1.Python爬虫入门一之综述
要学习Python爬虫,我们要学习的共有以下几点: Python基础知识 Python中urllib和urllib2库的用法 Python正则表达式 Python爬虫框架Scrapy Python爬虫 ...
Python 爬虫入门之爬取妹子图
Python 爬虫入门之爬取妹子图来源:李英杰链接: https://segmentfault.com/a/1190000015798452 听说你写代码没动力?本文就给你动力,爬取妹子图.如果 ...
Python爬虫入门一之综述
大家好哈,最近博主在学习Python,学习期间也遇到一些问题,获得了一些经验,在此将自己的学习系统地整理下来,如果大家有兴趣学习爬虫的话,可以将这些文章作为参考,也欢迎大家一共分享学习经验. Pyth ...
Python爬虫入门教程 48-100 使用mitmdump抓取手机惠农APP-手机APP爬虫部分
1. 爬取前的分析 mitmdump是mitmproxy的命令行接口,比Fiddler.Charles等工具方便的地方是它可以对接Python脚本. 有了它我们可以不用手动截获和分析HTTP请求和响应 ...
Python爬虫入门教程 43-100 百思不得姐APP数据-手机APP爬虫部分
1. Python爬虫入门教程爬取背景 2019年1月10日深夜,打开了百思不得姐APP,想了一下是否可以爬呢?不自觉的安装到了夜神模拟器里面.这个APP还是比较有名和有意思的. 下面是百思不得姐的 ...

随机推荐

junit配合catubuter统计单元测试的代码覆盖率
1.视频参考孔浩老师ant视频笔记对应的build-junit.xml脚步如下所示: <?xml version="1.0" encoding="UTF-8&qu ...
Redis高级特性介绍以及实例分析
Redis基础类型回顾转自:http://www.jianshu.com/p/af7043e6c8f9 String Redis中最基本,也是最简单的数据类型.注意,VALUE既可以是简单的Stri ...
服务消费者（Ribbon）
上一篇文章,简单概述了服务注册与发现,在微服务架构中,业务都会被拆分成一个独立的服务,服务之间的通讯是基于http restful的,Ribbon可以很好地控制HTTP和TCP客户端的行为,Sprin ...
js/ts/tsx读取excel表格中的日期格式转换
const formatDate = (timestamp: number) => { const time = new Date((timestamp - 1) * 24 * 3600000 ...
基于C#实现DXF文件读取显示
工控领域的制图软件仍然以AutoCAD为主,很多时候我们希望上位机软件可以读取CAD的图纸文件,从而控制设备按照绘制的路线进行运行,今天给大家分享的是如何使用C#读取DXF文件并进行显示. 公众号:[ ...
如何配置webpack让浏览器自动补全前缀
一.postcss-loader有什么用? PostCSS 本身是一个功能比较单一的工具.它提供了一种方式用 JavaScript 代码来处理 CSS.它负责把 CSS 代码解析成抽象语法树结构(Ab ...
css条纹背景样式、及方格斜纹背景的实现
一.横向条纹如下代码: background: linear-gradient(#fb3 %, #58a %) 上面代码表示整个图片的上部分20%和下部分20%是对应的纯色,只有中间的部分是渐变色.如 ...
django 后端分页
分页处理脚本: # -*- coding: utf-8 -*- # @Time : 2019-01-22 10:41 # @Author : 小贰 # @FileName: page.py # @fu ...
使用Python编写的对拍程序
简介支持数据生成程序模式, 只要有RE或者WA的数据点, 就会停止支持数据文件模式, 使用通配符指定输入文件, 将会对拍所有文件结束后将会打印统计信息第一次在某目录执行,将会通过交互方式获取配 ...
C++的基本输入输出
参考:http://www.runoob.com/cplusplus/cpp-basic-input-output.html