爬取知名社区技术文章_pipelines

获取字段的存储处理和获取普通的路径

#!/usr/bin/python3

# -*- coding: utf-8 -*-

import pymysql

import gevent

import pymysql

from gevent import monkey

from scrapy.pipelines.images import ImagesPipeline

import pymysql.cursors

class JobboleImagerPipeline(ImagesPipeline):

    """

    获得图片下载路径

    """

    def item_completed(self, results, item, info):

        if 'img_url' in item:

            for key, value in results:

                # print(key)

                img_path = value['path']

                # print(value['path'])

                item['img_path'] = img_path

        return item

# class SqlSave(object):

#     """常规同步方式存入数据库"""

#     def __init__(self):

#         SQL_DBA = {

#             'host': 'localhost',

#             'db': 'jobole',

#             'user': 'root',

#             'password': 'password',

#             'use_unicode': True,

#             'charset': 'utf8'

#         }

#         self.conn = pymysql.connect(**SQL_DBA)

#         self.cursor = self.conn.cursor()

#

#     def process_item(self, item, spider):

#         sql = self.get_sql(item)

#         print(sql)

#         self.cursor.execute(sql)

#         self.conn.commit()

#

#         return item

#

#     def get_sql(self, item):

#         sql = """insert into article(cont_id, cont_url, title, publish_time, cont, img_url, img_path, like_num, collection_num, comment_num) value ('%s','%s','%s','%s','%s','%s','%s', %d, %d, %d)

#         """ % (item['cont_id'], item['cont_url'],item['title'],item['publish_time'],item['cont'],item['img_url'][0],item['img_path'],item['link_num'],item['collection_num'],item['comment_num'],)

#         return sql

class SqlSave(object):

    """

    协程方式向数据库插入数据

    """

    def __init__(self):

        # 初始数据库连接和参数，SQL_DBA可写在setting中，通过 获取在settings.py中设置的SQL_DBA字典

        # @classmethod

        # def from_settings(cls, settings):

        #     sql_dba = settings[SQL_DBA]

        #     return cls(cls，sql_dba)           需要__init__中新添个参数接收这个值

        SQL_DBA = {

            'host': 'localhost',

            'db': 'jobole',

            'user': 'root',

            'password': 'password',

            'use_unicode': True,

            'charset': 'utf8'

        }

        self.conn = pymysql.connect(**SQL_DBA)

        self.cursor = self.conn.cursor()

    def process_item(self, item, spider):

        sql = self.__get_sql(item)

        # 协程方式对数据库插入操作

        gevent.joinall([

            gevent.spawn(self.__go_sql, self.cursor, self.conn, sql, item),

        ])

        return item

    def __go_sql(self, cursor, conn, sql, item):

        try:

            # 数据库插入操作

            cursor.execute(sql,

                           (item['cont_id'], item['cont_url'], item['title'], item['publish_time'],

                            item['cont'], item['img_url'][0], item['img_path'], item['link_num'],

                            item['collection_num'], item['comment_num']))

            conn.commit()

        except Exception as e:

            print(e)

    def __get_sql(self, item):

        # 生成sql语句

        sql = """insert into

                  article(cont_id, cont_url, title, publish_time,

                  cont, img_url, img_path, like_num,

                  collection_num, comment_num)

                value

                  (%s,%s,%s,%s,%s,%s,%s,%s,%s,%s)"""

        return sql

爬取知名社区技术文章_pipelines_4的更多相关文章

爬取知名社区技术文章_items_2
item中定义获取的字段和原始数据进行处理并合法化数据 #!/usr/bin/python3 # -*- coding: utf-8 -*- import scrapy import hashlib ...
爬取知名社区技术文章_setting_5
# -*- coding: utf-8 -*- # Scrapy settings for JobBole project # # For simplicity, this file contains ...
爬取知名社区技术文章_article_3
爬虫主逻辑处理,获取字段,获取主url和子url #!/usr/bin/python3 # -*- coding: utf-8 -*- import scrapy from scrapy.http i ...
第4章 scrapy爬取知名技术文章网站(2)
4-8~9 编写spider爬取jobbole的所有文章 # -*- coding: utf-8 -*- import re import scrapy import datetime from sc ...
爬取博主所有文章并保存到本地（.txt版）--python3.6
闲话: 一位前辈告诉我大学期间要好好维护自己的博客,在博客园发布很好,但是自己最好也保留一个备份. 正好最近在学习python,刚刚从py2转到py3,还有点不是很习惯,正想着多练习,于是萌生了这个想 ...
爬虫实战——Scrapy爬取伯乐在线所有文章
Scrapy简单介绍及爬取伯乐在线所有文章一.简说安装相关环境及依赖包 1.安装Python(2或3都行,我这里用的是3) 2.虚拟环境搭建: 依赖包:virtualenv,virtualenvwr ...
Node爬取简书首页文章
Node爬取简书首页文章博主刚学node,打算写个爬虫练练手,这次的爬虫目标是简书的首页文章流程分析使用superagent发送http请求到服务端,获取HTML文本用cheerio解析获得的 ...
使用Python爬取微信公众号文章并保存为PDF文件(解决图片不显示的问题)
前言第一次写博客,主要内容是爬取微信公众号的文章,将文章以PDF格式保存在本地. 爬取微信公众号文章(使用wechatsogou) 1.安装 pip install wechatsogou --up ...
Python3.6+Scrapy爬取知名技术文章网站
爬取分析伯乐在线已经提供了所有文章的接口,还有下一页的接口,所有我们可以直接爬取一页,再翻页爬. 环境搭建 Windows下安装Python: http://www.cnblogs.com/0bug ...

随机推荐

Java 非线程安全的HashMap如何在多线程中使用
Java 非线程安全的HashMap如何在多线程中使用 HashMap 是非线程安全的.在多线程条件下,容易导致死循环,具体表现为CPU使用率100%.因此多线程环境下保证 HashMap 的线程安全 ...
【转】NO.2、Appium之IOS第一个demo
接第一篇:Appium之iOS环境搭建 http://blog.csdn.net/clean_water/article/details/52946191 这个实例继承了unittest,重写了它的s ...
PE解析器的编写（四）——数据目录表的解析
在PE结构中最重要的就是区块表和数据目录表,上节已经说明了如何解析区块表,下面就是数据目录表,在数据目录表中一般只关心导入表,导出表和资源这几个部分,但是资源实在是太复杂了,而且在一般的病毒木马中也不 ...
CSS背景-background
复合属性-background 如果同时设置了background-color和background-image时,背景颜色会被图片覆盖. background-image: 用作背景的图片,back ...
JavaScript（六）函数
函数的声明方式 function name () {} 函数声明 var name = function(){} 函数表达式所有函数都有返回值未return 的函数返回值是 unde ...
Django-Views模块详解
http请求中产生的两个核心对象 http请求: HttpRequest http响应: HttpResponse 所在位置 django.http httpRequest属性: HttpReques ...
DAY5-小别-2018-1-15
有两天没有写了,前天考完试出去浪了,惭愧自己没有学习:昨天,启程回家看完了循环内容的视频,晚上十点半火车到站,没抽出时间写了,还看了<黑客帝国>,有点小感触,人工智能的时代即将到来,我们该 ...
有具体名称的匿名函数var bar = function foo(){}
http://kangax.github.io/nfe/ 命名的函数表达式函数表达式实际上可以经常看到.Web开发中的一个常见模式是基于某种特性测试来"分叉"函数定义,从而获得最 ...
ISE14.7安装教程（转）
ISE14.7可在百度云中下载链接:http://pan.baidu.com/s/1boQKyzd密码:a0m2 原文链接:http://blog.chinaaet.com/crazybird/p/3 ...
LVS集群ipvsadm命令和调度算法（6）
一.ipvsadm命令参考为了更好的让大家理解这份命令手册,将手册里面用到的几个术语先简单的介绍一下: 术语解释: 1.virtual-service-address:是指虚拟服务器的ip地址2.r ...

爬取知名社区技术文章_pipelines_4

爬取知名社区技术文章_pipelines_4的更多相关文章

随机推荐

热门专题