Python爬虫之路——简单网页抓图升级版（添加多线程支持）

转载自我的博客:http://www.mylonly.com/archives/1418.html

经过两个晚上的奋斗。将上一篇文章介绍的爬虫略微改进了下（Python爬虫之路——简单网页抓图），主要是将获取图片链接任务和下载图片任务用线程分开来处理了，并且这次的爬虫不只能够爬第一页的图片链接的，整个http://desk.zol.com.cn/meinv/以下的图片都会被爬到，并且提供了多种分辨率图片的文件下载，详细设置方法代码凝视里面有介绍。

这次的代码仍然有点不足，Ctrl-C无法终止程序，应该是线程无法响应主程序的终止消息导致的，（最好放在后台跑程序）还有线程的分配还能够优化的更好一点。兴许会陆续改进.

#coding: utf-8 #############################################################

# File Name: main.py

# Author: mylonly

# mail: mylonly@gmail.com

# Created Time: Wed 11 Jun 2014 08:22:12 PM CST

#########################################################################

#!/usr/bin/python

import re,urllib2,HTMLParser,threading,Queue,time

#各图集入口链接

htmlDoorList = []

#包括图片的Hmtl链接

htmlUrlList = []

#图片Url链接Queue

imageUrlList = Queue.Queue(0)

#捕获图片数量

imageGetCount = 0

#已下载图片数量

imageDownloadCount = 0

#每一个图集的起始地址。用于推断终止

nextHtmlUrl = ''

#本地保存路径

localSavePath = '/data/1920x1080/'

#假设你想下你须要的分辨率的，请改动replace_str,有例如以下分辨率可供选择1920x1200。1980x1920,1680x1050,1600x900,1440x900,1366x768,1280x1024,1024x768,1280x800

replace_str = '1920x1080'

replaced_str = '960x600'

#内页分析处理类

class ImageHtmlParser(HTMLParser.HTMLParser):

	def __init__(self):

		self.nextUrl = ''

		HTMLParser.HTMLParser.__init__(self)

	def handle_starttag(self,tag,attrs):

		global imageUrlList

		if(tag == 'img' and len(attrs) > 2 ):

			if(attrs[0] == ('id','bigImg')):

				url = attrs[1][1]

				url = url.replace(replaced_str,replace_str)

				imageUrlList.put(url)

				global imageGetCount

				imageGetCount = imageGetCount + 1

				print url

		elif(tag == 'a' and len(attrs) == 4):

			if(attrs[0] == ('id','pageNext') and attrs[1] == ('class','next')):

				global nextHtmlUrl

				nextHtmlUrl = attrs[2][1];

#首页分析类

class IndexHtmlParser(HTMLParser.HTMLParser):

	def __init__(self):

		self.urlList = []

		self.index = 0

		self.nextUrl = ''

		self.tagList = ['li','a']

		self.classList = ['photo-list-padding','pic']

		HTMLParser.HTMLParser.__init__(self)

	def handle_starttag(self,tag,attrs):

		if(tag == self.tagList[self.index]):

			for attr in attrs:

				if (attr[1] == self.classList[self.index]):

					if(self.index == 0):

						#第一层找到了

						self.index = 1

					else:

						#第二层找到了

						self.index = 0

						print attrs[1][1]

						self.urlList.append(attrs[1][1])

						break

		elif(tag == 'a'):

			for attr in attrs:

				if (attr[0] == 'id' and attr[1] == 'pageNext'):

					self.nextUrl = attrs[1][1]

					print 'nextUrl:',self.nextUrl

					break

#首页Hmtl解析器

indexParser = IndexHtmlParser()

#内页Html解析器

imageParser = ImageHtmlParser()

#依据首页得到全部入口链接

print '開始扫描首页...'

host = 'http://desk.zol.com.cn'

indexUrl = '/meinv/'

while (indexUrl != ''):

	print '正在抓取网页:',host+indexUrl

	request = urllib2.Request(host+indexUrl)

	try:

		m = urllib2.urlopen(request)

		con = m.read()

		indexParser.feed(con)

		if (indexUrl == indexParser.nextUrl):

			break

		else:

			indexUrl = indexParser.nextUrl

	except urllib2.URLError,e:

		print e.reason

print '首页扫描完毕，全部图集链接已获得：'

htmlDoorList = indexParser.urlList

#依据入口链接得到全部图片的url

class getImageUrl(threading.Thread):

	def __init__(self):

		threading.Thread.__init__(self)

	def run(self):

		for door in htmlDoorList:

			print '開始获取图片地址,入口地址为:',door

			global nextHtmlUrl

			nextHtmlUrl = ''

			while(door != ''):

				print '開始从网页%s获取图片...'% (host+door)

				if(nextHtmlUrl != ''):

					request = urllib2.Request(host+nextHtmlUrl)

				else:

					request = urllib2.Request(host+door)

				try:

					m = urllib2.urlopen(request)

					con = m.read()

					imageParser.feed(con)

					print '下一个页面地址为:',nextHtmlUrl

					if(door == nextHtmlUrl):

						break

				except urllib2.URLError,e:

					print e.reason

		print '全部图片地址均已获得:',imageUrlList

class getImage(threading.Thread):

	def __init__(self):

		threading.Thread.__init__(self)

	def run(self):

		global imageUrlList

		print '開始下载图片...'

		while(True):

			print '眼下捕获图片数量:',imageGetCount

			print '已下载图片数量:',imageDownloadCount

			image = imageUrlList.get()

			print '下载文件路径:',image

			try:

				cont = urllib2.urlopen(image).read()

				patter = '[0-9]*\.jpg';

				match = re.search(patter,image);

				if match:

					print '正在下载文件：',match.group()

					filename = localSavePath+match.group()

					f = open(filename,'wb')

					f.write(cont)

					f.close()

					global imageDownloadCount

					imageDownloadCount = imageDownloadCount + 1

				else:

					print 'no match'

				if(imageUrlList.empty()):

					break

			except urllib2.URLError,e:

				print e.reason

		print '文件全部下载完毕...'

get = getImageUrl()

get.start()

print '获取图片链接线程启动:'

time.sleep(2)

download = getImage()

download.start()

print '下载图片链接线程启动:'

Python爬虫之路——简单网页抓图升级版（添加多线程支持）的更多相关文章

Python爬虫之路——简单的网页抓图
转载自我自己的博客:http://www.mylonly.com/archives/1401.html 用Python的urllib2库和HTMLParser库写了一个简单的抓图脚本.主要抓的是htt ...
Python 爬虫修养-处理动态网页
Python 爬虫修养-处理动态网页本文转自:i春秋社区 0x01 前言在进行爬虫开发的过程中,我们会遇到很多的棘手的问题,当然对于普通的问题比如 UA 等修改的问题,我们并不在讨论范围,既然要将 ...
python爬虫之路——无头浏览器初识及简单例子
from selenium import webdriver url='https://www.jianshu.com/p/a64529b4ccf3' def get_info(url): inclu ...
Python爬虫学习之获取网页源码
偶然的机会,在知乎上看到一个有关爬虫的话题<利用爬虫技术能做到哪些很酷很有趣很有用的事情?>,因为强烈的好奇心和觉得会写爬虫是一件高大上的事情,所以就对爬虫产生了兴趣. 关于网络爬虫的定义 ...
Python爬虫实战：将网页转换为pdf电子书
写爬虫似乎没有比用 Python 更合适了,Python 社区提供的爬虫工具多得让你眼花缭乱,各种拿来就可以直接用的 library 分分钟就可以写出一个爬虫出来,今天就琢磨着写一个爬虫,将廖雪峰的 ...
【python爬虫】一个简单的爬取百家号文章的小爬虫
需求用"老龄智能"在百度百家号中搜索文章,爬取文章内容和相关信息. 观察网页红色框框的地方可以选择资讯来源,我这里选择的是百家号,因为百家号聚合了来自多个平台的新闻报道.首先看 ...
python爬虫之路——初识爬虫三大库，requests,lxml,beautiful.
三大库:requests,lxml,beautifulSoup. Request库作用:请求网站获取网页数据. get()的基本使用方法 #导入库 import requests #向网站发送请求,获 ...
python爬虫之路——初识基本页面构造原理
通过chrome浏览器的使用简单介绍网页构成 360浏览器使用右键审查元素,Chrome浏览器使用右键检查,都可查看网页代码. 网页代码有两部分:HTML文件和CSS样式.其中有<script& ...
python 爬虫（爬取网页的img并下载）
from urllib.request import urlopen # 引用第三方库 import requests #引用requests/用于访问网站(没安装需要安装) from pyquery ...

随机推荐

Python虚拟机函数机制之无参调用（一）
PyFunctionObject对象在Python中,任何一个东西都是对象,函数也不例外.函数这种抽象机制,是通过一个Python对象——PyFunctionObject来实现的 typedef s ...
线程中更新ui方法汇总
一.为何写作此文你是不是经常看到很多书籍中说:不能在子线程中操作ui,不然会报错.你是不是也遇到了如下的疑惑(见下面的代码): @Override protected void onCreate ...
LoadRunner11使用方法以及注意点收集
一:安装loadrunner http://jingyan.baidu.com/article/f7ff0bfc1cc82c2e26bb13b7.html http://www.cnblogs.com ...
深度学习：Sigmoid函数与损失函数求导
1.sigmoid函数 sigmoid函数,也就是s型曲线函数,如下: 函数: 导数: 上面是我们常见的形式,虽然知道这样的形式,也知道计算流程,不够感觉并不太直观,下面来分析一下. 1.1 ...
换一种思维看待PHP VS Node.js
php和javascript都是非常流行的编程语言,刚刚开始一个服务于服务端,一个服务于前端,长久以来,它们都能够和睦相处,直到有一天,一个叫做node.js的JavaScript运行环境诞生后,再加 ...
springboot集成shiro——登陆记住我
在shiro配置类中增加两个方法: com.resthour.config.shrio.ShiroConfiguration /** * cookie管理对象 * @return */ @Bean p ...
[LOJ#6002]「网络流 24 题」最小路径覆盖
[LOJ#6002]「网络流 24 题」最小路径覆盖试题描述给定有向图 G=(V,E).设 P 是 G 的一个简单路(顶点不相交)的集合.如果 V 中每个顶点恰好在 P 的一条路上,则称 P 是 ...
[luoguP2770] 航空路线问题（最小费用最大流）
传送门模型求最长两条不相交路径,用最大费用最大流解决. 实现为了限制经过次数,将每个点i拆成xi,yi. 1.从xi向yi连一条容量为1,费用为1的有向边(1<i<N), 2.从x1 ...
java中文乱码问题解决
1 处理乱码方式: 1 连接数据库的时候 jdbc.properties:jdbc:mysql://localhost:3306/myproject?useUnicode=true&chara ...
java Logger 的使用与配置
原文来自:http://blog.csdn.net/nash603/article/details/6749914 Logger所对应的属性文件在安装jdk目录下的jre/lib/logging.pr ...

Python爬虫之路——简单网页抓图升级版（添加多线程支持）

Python爬虫之路——简单网页抓图升级版（添加多线程支持）的更多相关文章

随机推荐

热门专题