Scrapy学习-10-Request&Response对象

请求URL流程

Scarpy使用请求和响应对象来抓取网站

通常情况下，请求对象会在spider中生成，并在系统中传递，直到到达downloader，它执行请求并返回一个响应对象，该对象返回发送请求的spider。

请求和响应类都有子类，它们添加了基类中不需要的功能。

Request对象

"""

This module implements the Request class which is used to represent HTTP

requests in Scrapy.

See documentation in docs/topics/request-response.rst

"""

import six

from w3lib.url import safe_url_string

from scrapy.http.headers import Headers

from scrapy.utils.python import to_bytes

from scrapy.utils.trackref import object_ref

from scrapy.utils.url import escape_ajax

from scrapy.http.common import obsolete_setter

class Request(object_ref):

    def __init__(self, url, callback=None, method='GET', headers=None, body=None,

                 cookies=None, meta=None, encoding='utf-8', priority=0,

                 dont_filter=False, errback=None, flags=None):

        self._encoding = encoding  # this one has to be set first

        self.method = str(method).upper()

        self._set_url(url)

        self._set_body(body)

        assert isinstance(priority, int), "Request priority not an integer: %r" % priority

        self.priority = priority

        if callback is not None and not callable(callback):

            raise TypeError('callback must be a callable, got %s' % type(callback).__name__)

        if errback is not None and not callable(errback):

            raise TypeError('errback must be a callable, got %s' % type(errback).__name__)

        assert callback or not errback, "Cannot use errback without a callback"

        self.callback = callback

        self.errback = errback

        self.cookies = cookies or {}

        self.headers = Headers(headers or {}, encoding=encoding)

        self.dont_filter = dont_filter

        self._meta = dict(meta) if meta else None

        self.flags = [] if flags is None else list(flags)

    @property

    def meta(self):

        if self._meta is None:

            self._meta = {}

        return self._meta

    def _get_url(self):

        return self._url

    def _set_url(self, url):

        if not isinstance(url, six.string_types):

            raise TypeError('Request url must be str or unicode, got %s:' % type(url).__name__)

        s = safe_url_string(url, self.encoding)

        self._url = escape_ajax(s)

        if ':' not in self._url:

            raise ValueError('Missing scheme in request url: %s' % self._url)

    url = property(_get_url, obsolete_setter(_set_url, 'url'))

    def _get_body(self):

        return self._body

    def _set_body(self, body):

        if body is None:

            self._body = b''

        else:

            self._body = to_bytes(body, self.encoding)

    body = property(_get_body, obsolete_setter(_set_body, 'body'))

    @property

    def encoding(self):

        return self._encoding

    def __str__(self):

        return "<%s %s>" % (self.method, self.url)

    __repr__ = __str__

    def copy(self):

        """Return a copy of this Request"""

        return self.replace()

    def replace(self, *args, **kwargs):

        """Create a new Request with the same attributes except for those

        given new values.

        """

        for x in ['url', 'method', 'headers', 'body', 'cookies', 'meta',

                  'encoding', 'priority', 'dont_filter', 'callback', 'errback']:

            kwargs.setdefault(x, getattr(self, x))

        cls = kwargs.pop('cls', self.__class__)

        return cls(*args, **kwargs)

部分参数解析

url (string) – the URL of this request

callback (callable) – the function that will be called with the response of this request (once its downloaded) as its first parameter. For more information see Passing additional data to callback functions below. If a Request doesn’t specify a callback, the spider’s parse() method will be used. Note that if exceptions are raised during processing, errback is called instead.

method (string) – the HTTP method of this request. Defaults to 'GET'.

meta (dict) – the initial values for the Request.meta attribute. If given, the dict passed in this parameter will be shallow copied.

body (str or unicode) – the request body. If a unicode is passed, then it’s encoded to str using the encoding passed (which defaults to utf-8). If body is not given, an empty string is stored. Regardless of the type of this argument, the final value stored will be a str (never unicode or None).

headers (dict) – the headers of this request. The dict values can be strings (for single valued headers) or lists (for multi-valued headers). If None is passed as value, the HTTP header will not be sent at all.

cookies (dict or list) –

the request cookies. These can be sent in two forms.

1.Using a dict:

    request_with_cookies = Request(url="http://www.example.com",

                               cookies={'currency': 'USD', 'country': 'UY'})

2. Using a list of dicts

    request_with_cookies =     Request(url="http://www.example.com",

                               cookies=[{'name': 'currency',

                                        'value': 'USD',

                                        'domain': 'example.com',

                                        'path': '/currency'}])

Response对象

"""

This module implements the Response class which is used to represent HTTP

responses in Scrapy.

See documentation in docs/topics/request-response.rst

"""

from six.moves.urllib.parse import urljoin

from scrapy.http.request import Request

from scrapy.http.headers import Headers

from scrapy.link import Link

from scrapy.utils.trackref import object_ref

from scrapy.http.common import obsolete_setter

from scrapy.exceptions import NotSupported

class Response(object_ref):

    def __init__(self, url, status=200, headers=None, body=b'', flags=None, request=None):

        self.headers = Headers(headers or {})

        self.status = int(status)

        self._set_body(body)

        self._set_url(url)

        self.request = request

        self.flags = [] if flags is None else list(flags)

    @property

    def meta(self):

        try:

            return self.request.meta

        except AttributeError:

            raise AttributeError(

                "Response.meta not available, this response "

                "is not tied to any request"

            )

    def _get_url(self):

        return self._url

    def _set_url(self, url):

        if isinstance(url, str):

            self._url = url

        else:

            raise TypeError('%s url must be str, got %s:' % (type(self).__name__,

                type(url).__name__))

    url = property(_get_url, obsolete_setter(_set_url, 'url'))

    def _get_body(self):

        return self._body

    def _set_body(self, body):

        if body is None:

            self._body = b''

        elif not isinstance(body, bytes):

            raise TypeError(

                "Response body must be bytes. "

                "If you want to pass unicode body use TextResponse "

                "or HtmlResponse.")

        else:

            self._body = body

    body = property(_get_body, obsolete_setter(_set_body, 'body'))

    def __str__(self):

        return "<%d %s>" % (self.status, self.url)

    __repr__ = __str__

    def copy(self):

        """Return a copy of this Response"""

        return self.replace()

    def replace(self, *args, **kwargs):

        """Create a new Response with the same attributes except for those

        given new values.

        """

        for x in ['url', 'status', 'headers', 'body', 'request', 'flags']:

            kwargs.setdefault(x, getattr(self, x))

        cls = kwargs.pop('cls', self.__class__)

        return cls(*args, **kwargs)

    def urljoin(self, url):

        """Join this Response's url with a possible relative url to form an

        absolute interpretation of the latter."""

        return urljoin(self.url, url)

    @property

    def text(self):

        """For subclasses of TextResponse, this will return the body

        as text (unicode object in Python 2 and str in Python 3)

        """

        raise AttributeError("Response content isn't text")

    def css(self, *a, **kw):

        """Shortcut method implemented only by responses whose content

        is text (subclasses of TextResponse).

        """

        raise NotSupported("Response content isn't text")

    def xpath(self, *a, **kw):

        """Shortcut method implemented only by responses whose content

        is text (subclasses of TextResponse).

        """

        raise NotSupported("Response content isn't text")

    def follow(self, url, callback=None, method='GET', headers=None, body=None,

               cookies=None, meta=None, encoding='utf-8', priority=0,

               dont_filter=False, errback=None):

        # type: (...) -> Request

        """

        Return a :class:`~.Request` instance to follow a link ``url``.

        It accepts the same arguments as ``Request.__init__`` method,

        but ``url`` can be a relative URL or a ``scrapy.link.Link`` object,

        not only an absolute URL.

        :class:`~.TextResponse` provides a :meth:`~.TextResponse.follow`

        method which supports selectors in addition to absolute/relative URLs

        and Link objects.

        """

        if isinstance(url, Link):

            url = url.url

        url = self.urljoin(url)

        return Request(url, callback,

                       method=method,

                       headers=headers,

                       body=body,

                       cookies=cookies,

                       meta=meta,

                       encoding=encoding,

                       priority=priority,

                       dont_filter=dont_filter,

                       errback=errback)

参考官方文档 https://doc.scrapy.org

Scrapy学习-10-Request&Response对象的更多相关文章

Servlet的学习之Request请求对象（3）
本篇接上一篇,将Servlet中的HttpServletRequest对象获取RequestDispatcher对象后能进行的[转发]forward功能和[包含]include功能介绍完. 首先来看R ...
Servlet的学习之Request请求对象（2）
在上一篇<Servlet的学习(十)>中介绍了HttpServletRequest请求对象的一些常用方法,而从这篇起开始介绍和学习HttpServletRequest的常用功能. 使用Ht ...
Servlet的学习之Request请求对象（1）
在本篇中开始对Servlet中的HttpServletRequest请求对象进行学习,请求对象同响应对象一样,我们可以根据该对象中的方法获取例如请求行,请求头和请求实体数据的方法. 在本篇中先对Htt ...
Java-Spring-获取Request,Response对象
转载自:https://www.cnblogs.com/bjlhx/p/6639542.html 第一种.参数 @RequestMapping("/test") @Response ...
request与response对象.
request与response对象. 1. request代表请求对象 response代表的响应对象. 学习它们我们可以操作http请求与响应. 2.request,response体系结构. 在 ...
request与response对象详述
request与response对象. 1. request代表请求对象 response代表的响应对象. 学习它们我们可以操作http请求与响应. 2.request,response体系结构. 在 ...
java中获取request与response对象的方法
Java 获取Request,Response对象方法第一种.参数 @RequestMapping("/test") @ResponseBody public void sa ...
SpringMvc4中获取request、response对象的方法
springMVC4中获取request和response对象有以下两种简单易用的方法: 1.在control层获取在control层中获取HttpServletRequest和HttpServle ...
Scrapy 中 Request 对象和 Response 对象的各参数及属性介绍
Request 对象 Request构造器方法的参数列表: Request(url [, callback=None, method='GET', headers=None, body=None,co ...

随机推荐

产生式模型（生成式模型）与判别式模型<转载>
转自http://dongzipnf.blog.sohu.com/189983746.html 产生式模型与判别式模型产生式模型(Generative Model)与判别式模型(Discrimiti ...
C#编写高并发数据库控制
往往大数据量,高并发时, 瓶颈都在数据库上, 好多人都说用数据库的复制,发布, 读写分离等技术, 但主从数据库之间同步时间有延迟.代码的作用在于保证在上端缓存服务失效(一般来说概率比较低)时,形成倒瓶 ...
BCB:内存泄漏检查工具CodeGuard
一.为什么写这篇东西自己在使用BCB5写一些程序时需要检查很多东西,例如内存泄漏.资源是否有释放等等,在使用了很多工具后,发觉BCB5本身自带的工具―CodeGuard,非常不错,使用也挺方便的,但 ...
Caused by: java.lang.ClassNotFoundException: java.com.bj186.ssm.controller.UserController
在搭建SpringMVC的时候,遇到的这个问题真的很奇葩, 找不到UserController这个类这明明不就在工程目录下吗? 经过了一番艰苦卓绝的斗争, 才发现原来是包导少了之前导入的包是: & ...
ant design table td 文字显示过长添加省略号、ant 文字过长时添加tootip提示
方法1: overflow: hidden; text-overflow: ellipsis; display: -webkit-box; -webkit-line-clamp: 2; -webkit ...
ViewController的lifecycle和autolayout
docker的网络(进阶)
overlay网络 overlay网络驱动程序会在多个docker守护程序(即多个主机上的docker守护程序)之间创建分布式网络.该网络(overlays)位于特定于主机的网络之上,允许连接到它的容 ...
文件操作-mkdir
Linux mkdir命令主要用来创建目录,也可以直接创建多层目录,本文就为大家介绍下 Linux mkdir命令 . 转载自https://www.linuxdaxue.com/linux-com ...
PyCharm 社区版创建Django项目的一个方法
PyCharm 社区版创建项目无法选择Django等项目,只能选择Python项目. 你在进行练习的时候为了方便,可以用过期了的PyCharm专业版在可用的30分钟内创建社区版本不支持的项目,再用Py ...
LIN总线协议
汽车电子类的IC有的采用LIN协议来烧录内部NVM,如英飞凌的TLE8880N和博世的CR665D. LIN总线帧格式如下,一个LIN信息帧有同步间隔.同步域.标示符域(PID域).数据域.校验码域. ...

Scrapy学习-10-Request&Response对象

Scrapy学习-10-Request&Response对象的更多相关文章

随机推荐

热门专题