一个网页抓取的类支持get+post+cookie存储

前段时间提取了一个工具类，分享给大家：

<?php

class httpconnector {

    private $curl;

    private $cookie;

    private $kv;

    function __construct(){

        $this->kv = new SaeKV();

        $this->kv->init();

        if($data=$this->kv->get("cookie"))

          $this->cookie=$data;

    }

    public function get($url) {

        $this->curl = curl_init();

        curl_setopt($this->curl, CURLOPT_URL, $url);

        curl_setopt($this->curl, CURLOPT_HEADER, 1);

        curl_setopt($this->curl, CURLOPT_USERAGENT, "Mozilla/4.0 (compatible; MSIE 5.01; Windows NT 5.0)");

        curl_setopt($this->curl, CURLOPT_COOKIE, $this->cookie);

        curl_setopt($this->curl, CURLOPT_RETURNTRANSFER, 1);

        $data = curl_exec($this->curl);

        curl_close($this->curl);

        preg_match_all("/Set-Cookie:(.*?);/", $data, $match, PREG_SET_ORDER);

        foreach ($match as $r) {

            if ($this->cookie != '') {

                $this->cookie = $this->cookie . ';';

            }

            if (isset($r[1])) {

                $this->cookie .= trim(str_replace("\r\n", "", $r[1]));

            }

        }

        $this->kv->set("cookie",$this->cookie);

        return $data;

    }

    public function post($url, $params) {

        $this->curl = curl_init();

        curl_setopt($this->curl, CURLOPT_URL, $url);

        curl_setopt($this->curl, CURLOPT_HEADER, 1);

        curl_setopt($this->curl, CURLOPT_COOKIE, $this->cookie);

        curl_setopt($this->curl, CURLOPT_POST, 1);

        curl_setopt($this->curl, CURLOPT_USERAGENT, "Mozilla/4.0 (compatible; MSIE 5.01; Windows NT 5.0)");

        curl_setopt($this->curl, CURLOPT_POSTFIELDS, $params);

        curl_setopt($this->curl, CURLOPT_RETURNTRANSFER, 1);

        $data = curl_exec($this->curl);

        curl_close($this->curl);

        preg_match_all("/Set-Cookie:(.*?);/", $data, $match, PREG_SET_ORDER);

        foreach ($match as $r) {

            if ($this->cookie != '') {

                $this->cookie = $this->cookie . ';';

            }

            if (isset($r[1])) {

                $this->cookie .= trim(str_replace("\r\n", "", $r[1]));

            }

        }

        $this->kv->set("cookie",$this->cookie);

        return $data;

    }

}

?>

一个网页抓取的类支持get+post+cookie存储的更多相关文章

分享一个c#t的网页抓取类
using System; using System.Collections.Generic; using System.Web; using System.Text; using System.Ne ...
Java实现网页抓取的一个Demo
这个小案例的话我是存放在我的github 上. 下面给出链接自己可以去看下,也可以直接下载源码.有具体的说明 <Java网页抓取>
基于Casperjs的网页抓取技术【抓取豆瓣信息网络爬虫实战示例】
CasperJS is a navigation scripting & testing utility for the PhantomJS (WebKit) and SlimerJS (Ge ...
Python开发爬虫之动态网页抓取篇：爬取博客评论数据——通过Selenium模拟浏览器抓取
区别于上篇动态网页抓取,这里介绍另一种方法,即使用浏览器渲染引擎.直接用浏览器在显示网页时解析 HTML.应用 CSS 样式并执行 JavaScript 的语句. 这个方法在爬虫过程中会打开一个浏览器 ...
Python爬虫之三种网页抓取方法性能比较
下面我们将介绍三种抓取网页数据的方法,首先是正则表达式,然后是流行的 BeautifulSoup 模块,最后是强大的 lxml 模块. 1. 正则表达式如果你对正则表达式还不熟悉,或是需要一些提 ...
Python之HTML的解析（网页抓取一）
http://blog.csdn.net/my2010sam/article/details/14526223 --------------------- 对html的解析是网页抓取的基础,分析抓取的 ...
java网页抓取
网页抓取就是,我们想要从别人的网站上得到我们想要的,也算是窃取了,有的网站就对这个网页抓取就做了限制,比如百度直接进入正题 //要抓取的网页地址 String urlStr = "http ...
网页抓取：PHP实现网页爬虫方式小结
来源:http://www.ido321.com/1158.html 抓取某一个网页中的内容,需要对DOM树进行解析,找到指定节点后,再抓取我们需要的内容,过程有点繁琐.LZ总结了几种常用的.易于实现 ...
Python实现简单的网页抓取
现在开源的网页抓取程序有很多,各种语言应有尽有. 这里分享一下Python从零开始的网页抓取过程第一步:安装Python 点击下载适合的版本https://www.python.org/ 我这里选择 ...

随机推荐

利用Browser Link提高前端开发的生产力
(此文章同时发表在本人微信公众号"dotNET每日精华文章",欢迎右边二维码来关注.) 题记:Browser Link是VS 2013开始引入的一个强大功能,让前端代码(比如Ang ...
阿里云（ECS）Centos服务器LNMP环境搭建
阿里云( ECS ) Centos7 服务器 LNMP 环境搭建前言第一次接触阿里云是大四的时候,当时在校外公司做兼职,关于智能家居项目的,话说当时俺就只有一个月左右的 php 后台开发经验(还是 ...
mathematica练习程序（获得股票数据）
从去年的11月开始,中国的股市就一直大涨,不知道这次能持续多长时间. 为了获得股票数据,我用matlab试了网上的一些方法,总是失败,所以就改用mathematica,一行代码就可以了. DateLi ...
封装自己的printf函数
#include <stdio.h> #include <stdarg.h> //方式一 #define DBG_PRINT (printf("%s:%u %s:%s ...
loj 1167(二分+最大流）
题目链接:http://acm.hust.edu.cn/vjudge/problem/viewProblem.action?id=26881 思路:我们可以二分最大危险度,然后建图,由于每个休息点只能 ...
sql 提取数字、字母、汉字
--提取数字 IF OBJECT_ID('DBO.GET_NUMBER2') IS NOT NULL DROP FUNCTION DBO.GET_NUMBER2 GO )) ) AS BEGIN BE ...
How to choose the number of topics/partitions in a Kafka cluster?
This is a common question asked by many Kafka users. The goal of this post is to explain a few impor ...
模仿QQ左滑删除
需求: 1.左滑删除 2.向左滑动距离超过一半的时候让它自动滑开,向右滑动超过一半的时候自动隐藏 3.一次只允许滑开一个item 还有,根本不需要自定义view来实现,谨防入坑布局: <?xm ...
Hibernate的持久化类状态
Hibernate的持久化类状态持久化类:就是一个实体类与数据库表建立了映射. Hibernate为了方便管理持久化类,将持久化类分成了三种状态. 瞬时态 transient (临时态):持久化 ...
让一个div在不同的显示器中永远居中
<!DOCTYPE html> <html> <head lang="en"> <meta charset="UTF-8&quo ...

一个网页抓取的类支持get+post+cookie存储

一个网页抓取的类支持get+post+cookie存储的更多相关文章

随机推荐

热门专题