python爬虫学习(7) —— 爬取你的AC代码

上一篇文章中，我们介绍了python爬虫利器——requests，并且拿HDU做了小测试。

这篇文章，我们来爬取一下自己AC的代码。

1 确定ac代码对应的页面

如下图所示，我们一般情况可以通过该顺序找到我们曾经AC过的代码

登陆hdu -> 点击自己的信息 -> 点击Last accepted submissions -> 在Code Len 处选择一个代码 -> 看到你AC的代码

我们可以看到，所有AC代码的页面都是

http://acm.hdu.edu.cn/viewcode.php?rid= + RunID

而这个RunID，正好在表格的最前面：

很自然我们可以想到，用正则表达式进行匹配。

2. 处理换页问题

很显然，，如果你AC的代码多了，必然会存在换页问题

不过我们可以在源代码中找到换页对应的URL，我们直接跳转，直到找不到为止。

3. 代码处理问题

html中有一些转义字符，使得我们不能直接将代码保存下来。

这时候我们需要用到 HTMLParser

code = html_parser.unescape(down_code)

4. 具体实现

#coding=utf-8

import re, HTMLParser, requests, getpass, os

# 初始化会话对象 以及 cookies

s = requests.session()

cookies = dict(cookies_are='working')

# 一些基础的url

host_url = 'http://acm.hdu.edu.cn'

post_url = 'http://acm.hdu.edu.cn/userloginex.php?action=login'

status_url = 'http://acm.hdu.edu.cn/status.php?user='

codebase_url = 'http://acm.hdu.edu.cn/viewcode.php?rid='

# 正则表达式的匹配串

runid_pat = re.compile(r'<tr.*?align=center ><td height=22px>(.*?)</td>.*?</tr>',re.S)

code_pat = re.compile(r'<textarea id=usercode style="display:none;text-align:left;">(.+?)</textarea>',re.S)

lan_pat = re.compile(r'Language : (.*?)&nbsp;&nbsp',re.S)

problem_pat = re.compile(r'Problem : <a href=.*?target=_blank>(.*?) .*?</a>',re.S)

nextpage_pat = re.compile(r'Prev Page</a><a style="margin-right:20px" href="(.*?)">Next Page ></a>',re.S)

# 代码保存目录

if not os.path.exists('./ac_code'):

	os.mkdir(r'./ac_code')

base_path = r'./ac_code/'

# 登陆

def login(usr,psw):

	data = {'username':usr,'userpass':psw,'login':'Sign In'}

	r = s.post(post_url,data=data,cookies=cookies)

# 代码语言判断

def lan_judge(language):

	if language == 'G++':

		suffix = '.cpp'

	elif language == 'GCC':

		suffix = '.c'

	elif language == 'C++':

		suffix = '.cpp'

	elif language == 'C':

		suffix = '.c'

	elif language == 'Pascal':

		suffix = '.pas'

	elif language == 'Java':

		suffix = '.java'

	else:

		suffix = '.cpp'

	return suffix

if __name__ == '__main__':

	usr = raw_input('input your username:')

	psw = getpass.getpass('input your password:')

	login(usr,psw)

	# 用于处理html中的转义字符

	html_parser = HTMLParser.HTMLParser()

	# 遍历每一页，并下载其代码

	status_url = status_url  + usr + '&status=5'

	status_html = s.get(status_url,cookies=cookies).text

	flag = True

	print "Just go!"

	while( flag ):

		runid_list = runid_pat.findall(status_html)

		for id in runid_list:

			code_url = codebase_url + id

			down_html = s.get(code_url,cookies=cookies).text

			down_code = code_pat.search(down_html).group(1)

			language = lan_pat.search(down_html).group(1)

			problemid = problem_pat.search(down_html).group(1)

			suffix = lan_judge(language)

			code = html_parser.unescape(down_code).encode('utf-8')

			code = code.replace('\r\n','\n')

			open( base_path + 'hdu' + problemid + '__' + id + suffix,"wb").write(code)

		nexturl = nextpage_pat.search(status_html)

		if nexturl == None:

			flag = False

		else:

			status_url = host_url + nexturl.group(1)

			status_html = s.get(status_url,cookies=cookies).text

	print "all of your ac codes were saved!"

5. 效果截图

ubuntu下测试：

windows下测试：

6. 写在后面

额。。很久很久很久没有刷题了。。。233333333，其实我是想告诉你一个事实

这里就可以下载AC的代码，，，哈哈哈。so，这个爬虫仅仅用来练习就好。

python爬虫学习(7) —— 爬取你的AC代码的更多相关文章

Python爬虫学习(二) ——————爬取前程无忧招聘信息并写入excel
作为一名Pythoner,相信大家对Python的就业前景或多或少会有一些关注.索性我们就写一个爬虫去获取一些我们需要的信息,今天我们要爬取的是前程无忧!说干就干!进入到前程无忧的官网,输入关键字&q ...
python爬虫学习之爬取全国各省市县级城市邮政编码
实例需求:运用python语言在http://www.ip138.com/post/网站爬取全国各个省市县级城市的邮政编码,并且保存在excel文件中实例环境:python3.7 requests库 ...
【转载】教你分分钟学会用python爬虫框架Scrapy爬取心目中的女神
原文:教你分分钟学会用python爬虫框架Scrapy爬取心目中的女神本博文将带领你从入门到精通爬虫框架Scrapy,最终具备爬取任何网页的数据的能力.本文以校花网为例进行爬取,校花网:http:/ ...
Python爬虫实例：爬取B站《工作细胞》短评——异步加载信息的爬取
很多网页的信息都是通过异步加载的,本文就举例讨论下此类网页的抓取. <工作细胞>最近比较火,bilibili 上目前的短评已经有17000多条. 先看分析下页面右边 li 标签中的就是短 ...
Python爬虫实例：爬取猫眼电影——破解字体反爬
字体反爬字体反爬也就是自定义字体反爬,通过调用自定义的字体文件来渲染网页中的文字,而网页中的文字不再是文字,而是相应的字体编码,通过复制或者简单的采集是无法采集到编码后的文字内容的. 现在貌似不少网 ...
Python爬虫实例：爬取豆瓣Top250
入门第一个爬虫一般都是爬这个,实在是太简单.用了 requests 和 bs4 库. 1.检查网页元素,提取所需要的信息并保存.这个用 bs4 就可以,前面的文章中已经有详细的用法阐述. 2.找到下一 ...
python爬虫-基础入门-爬取整个网站《3》
python爬虫-基础入门-爬取整个网站<3> 描述: 前两章粗略的讲述了python2.python3爬取整个网站,这章节简单的记录一下python2.python3的区别 python ...
python爬虫-基础入门-爬取整个网站《2》
python爬虫-基础入门-爬取整个网站<2> 描述: 开场白已在<python爬虫-基础入门-爬取整个网站<1>>中描述过了,这里不在描述,只附上 python3 ...
python爬虫-基础入门-爬取整个网站《1》
python爬虫-基础入门-爬取整个网站<1> 描述: 使用环境:python2.7.15 ,开发工具:pycharm,现爬取一个网站页面(http://www.baidu.com)所有数 ...

随机推荐

在 CSS 预编译器之后：PostCSS
提到css预编译器(css preprocessor),你可能想到Sass.Less以及Stylus.而本文要介绍的PostCSS,正是一个这样的工具:css预编译器可以做到的事,它同样可以做到. “ ...
构建ASP.NET MVC4+EF5+EasyUI+Unity2.x注入的后台管理系统（9）-MVC与EasyUI结合增删改查
系列目录文章于2016-12-17日重写在第八讲中,我们已经做到了怎么样分页.这一讲主要讲增删改查.第六讲的代码已经给出,里面包含了增删改,大家可以下载下来看下. 这讲主要是,制作漂亮的工具栏,虽 ...
NPOI导出Excel
using System;using System.Collections.Generic;using System.Linq;using System.Text;#region NPOIusing ...
提升用户体验的最佳免费 jQuery 表单插件
网页表单是一个老生常谈的话题.出于这样或那样的目的,一些示例中都会包括用户注册,电子商务结算,用户设置甚至联系人表格.而输入栏是非常容易用现代的CSS3技术来应用样式.但是到底什么决定整体用户体验? ...
php内核分析（二）－ZTS和zend_try
这里阅读的php版本为PHP-7.1.0 RC3,阅读代码的平台为linux ZTS 我们会看到文章中有很多地方是: #ifdef ZTS # define CG(v) ZEND_TSRMG(comp ...
unsafe
今天无意中发现C#这种完全面向对象的高级语言中也可以用不安全的指针类型,即要用到unsafe关键字.在公共语言运行库 (CLR) 中,不安全代码是指无法验证的代码.C# 中的不安全代码不一定是危险的, ...
联想 Thinkpad X230 SLIC 2.1 Marker
等了好久,终于等到了 X230 的 SLIC 2.1 的 Marker !特发帖备份... 基本情况笔记本:Lenovo X230(i5+8G+500G) 操作系统:Windows 7 Pro x6 ...
.NET 版本区别，以及与 Windows 的关系
老是记不住各 Windows 版本中的 .NET 版本号,下面汇总一下: .NET Framework各版本汇总以及之间的关系 Mailbag: What version of the .NET Fr ...
C#写文本日志帮助类(支持多线程)改进版(不适用于ASP.NET程序)
由于iis的自动回收机制,不适用于ASP.NET程序代码: using System; using System.Collections.Concurrent; using System.Confi ...
来玩Play框架02 响应
作者:Vamei 出处:http://www.cnblogs.com/vamei 欢迎转载,也请保留这段声明.谢谢! 我上一章总结了Play框架的基本使用.这一章里,我将修改和增加响应. HTTP协议 ...