Python 爬基金数据

爬科学基金共享服务网中基金数据

#coding=utf-8

import json

import requests

from lxml import etree

from HTMLParser import HTMLParser

from pymongo import MongoClient

data = {'pageSize':10,'currentPage':1,'fundingProject.projectNo':'','fundingProject.name':'','fundingProject.person':'','fundingProject.org':'',

'fundingProject.applyCode':'','fundingProject.grantCode':'','fundingProject.subGrantCode':'','fundingProject.helpGrantCode':'','fundingProject.keyword':'',

'fundingProject.statYear':'','checkCode':'%E8%AF%B7%E8%BE%93%E5%85%A5%E9%AA%8C%E8%AF%81%E7%A0%81'}

url = 'http://npd.nsfc.gov.cn/fundingProjectSearchAction!search.action'

headers = {'Accept':'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,image/apng,*/*;q=0.8',

'Accept-Encoding':'gzip, deflate',

'Accept-Language':'zh-CN,zh;q=0.9',

'Cache-Control':'max-age=0',

'Connection':'keep-alive',

'Content-Length':'',

'Content-Type':'application/x-www-form-urlencoded',

'Cookie':'JSESSIONID=8BD27CE37366ED8022B42BFC68FF82D4',

'Host':'npd.nsfc.gov.cn',

'Origin':'http://npd.nsfc.gov.cn',

'Referer':'http://npd.nsfc.gov.cn/fundingProjectSearchAction!search.action',

'Upgrade-Insecure-Requests':'',

'User-Agent':'Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/63.0.3239.132 Safari/537.36'}

def main():

    client = MongoClient('localhost', 27017)

    db = client.ScienceFund

    db.authenticate("","")

    collection=db.science_fund

    for i in range(1, 43184):

        print i

        data['currentPage'] = i

        result = requests.post(url, data = data, headers = headers)

        html = result.text

        tree = etree.HTML(html)

        table = tree.xpath("//dl[@class='time_dl']")

        for item in table:

            content = etree.tostring(item, method='html')

            content =  HTMLParser().unescape(content)

            # print content

            bson = jiexi(content)

            collection.insert(bson)

def jiexi(content):

    # 标题

    title1 = content.find('">', 20)

    title2 = content.find('</')

    title = content[title1+2:title2]

    # print title

    # 批准号

    standard_no1 = content.find(u'批准号', title2)

    standard_no2 = content.find('</dd>', standard_no1)

    standard_no = content[standard_no1+4:standard_no2].strip()

    # print standard_no

    # 项目类别

    standard_type1 = content.find(u'项目类别', standard_no2)

    standard_type2 = content.find('</dd>', standard_type1)

    standard_type = content[standard_type1+5:standard_type2].strip()

    # print standard_type

    # 依托单位

    supporting_institution1 = content.find(u'依托单位', standard_type2)

    supporting_institution2= content.find('</dd>', supporting_institution1)

    supporting_institution = content[supporting_institution1+5:supporting_institution2].strip()

    # print supporting_institution

    # 项目负责人

    project_principal1 = content.find(u'项目负责人', supporting_institution2)

    project_principal2 = content.find('</dd>', project_principal1)

    project_principal = content[project_principal1+6:project_principal2].strip()

    # print project_principal

    # 资助经费

    funds1 = content.find(u'资助经费', project_principal2)

    funds2 = content.find('</dd>', funds1)

    funds = content[funds1+5:funds2].strip()

    # print funds

    # 批准年度

    year1 = content.find(u'批准年度', funds2)

    year2 = content.find('</dd>', year1)

    year = content[year1+5:year2].strip()

    # print year

    # 关键词

    keywords1 = content.find(u'关键词', year2)

    keywords2 = content.find('</dd>', keywords1)

    keywords = content[keywords1+4:keywords2].strip()

    # print keywords

    dc = {}

    dc['title'] = title

    dc['standard_no'] = standard_no

    dc['standard_type'] = standard_type

    dc['supporting_institution'] = supporting_institution

    dc['project_principal'] = project_principal

    dc['funds'] = funds

    dc['year'] = year

    dc['keywords'] = keywords

    return dc

if __name__ == '__main__':

    main()

Python 爬基金数据的更多相关文章

python爬取数据需要注意的问题
1 爬取https的网站或是接口的时候,如果是不受信用的SSL证书,会报错,需要添加如下代码,如下代码可以保证当前代码块内所有的请求都自动屏蔽ssl证书问题: import ssl # 这个是爬取ht ...
python爬取数据保存到Excel中
# -*- conding:utf-8 -*- # 1.两页的内容 # 2.抓取每页title和URL # 3.根据title创建文件,发送URL请求,提取数据 import requests fro ...
python爬取数据保存入库
import urllib2 import re import MySQLdb class LatestTest: #初始化 def __init__(self): self.url="ht ...
Python 爬起数据时 'gbk' codec can't encode character '\xa0' 的问题
1.被这个问题折腾了一上午终于解决了,再网上看到有用 string.replace(u'\xa0',u' ') 替换成空格的,方法试了没用. 后来发现要在open的时候加utf-8才解决问题. 以 ...
Python 爬取数据入库mysql
# -*- enconding:etf-8 -*- import pymysql import os import time import re serveraddr="localhost& ...
Python 爬取美团酒店信息
事由:近期和朋友聊天,聊到黄山酒店事情,需要了解一下黄山的酒店情况,然后就想着用python 爬一些数据出来,做个参考主要思路:通过查找,基本思路清晰,目标明确,仅仅爬取美团莫一地区的酒店信息,不过 ...
如何使用Python爬取基金数据，并可视化显示
本文的文字及图片来源于网络,仅供学习.交流使用,不具有任何商业用途,版权归原作者所有,如有问题请及时联系我们以作处理以下文章来源于Will的大食堂,作者打饭大叔前言美国疫情越来越严峻,大选也进入 ...
python爬取股票最新数据并用excel绘制树状图
大家好,最近大A的白马股们简直跌妈不认,作为重仓了抱团白马股基金的养鸡少年,每日那是一个以泪洗面啊. 不过从金融界最近一个交易日的大盘云图来看,其实很多中小股还是红色滴,绿的都是白马股们. 以下截图 ...
python爬取网站数据
开学前接了一个任务,内容是从网上爬取特定属性的数据.正好之前学了python,练练手. 编码问题因为涉及到中文,所以必然地涉及到了编码的问题,这一次借这个机会算是彻底搞清楚了. 问题要从文字的编码讲 ...

随机推荐

【C++】类的特殊成员变量+初始化列表
参考资料: 1.黄邦勇帅 2.http://blog.163.com/sunshine_linting/blog/static/448933232011810101848652/ 3.http://w ...
HDU-3221
Brute-force Algorithm Time Limit: 2000/1000 MS (Java/Others) Memory Limit: 32768/32768 K (Java/Ot ...
Linux命令之：tr
1. 用途: tr,translate的简写,主要用于压缩重复字符,删除文件中的控制字符以及进行字符转换操作. 2. 语法: tr [OPTION]... SET1 [SET2] 3. 参数: -s: ...
mysql TIMESTAMPDIFF
在MySQL应用时,经常要使用这两个函数TIMESTAMPDIFF和TIMESTAMPADD. 一,TIMESTAMPDIFF 语法: TIMESTAMPDIFF(interval,datetime_ ...
【xunsearch】笔记
1.添加索引 $ cd /usr/local/xunsearch/sdk/php/ $ util/Indexer.php --rebuild --source=mysql://数据库用户名:数据库密码 ...
HDU 3045 Picnic Cows
$dp$,斜率优化. 设$dp[i]$表示$1$至$i$位置的最小费用,则$dp[i]=min(dp[j]+s[i]-s[j]-(i-j)*x[j+1])$,$dp[n]$为答案. 然后斜率优化就可以 ...
RPD Volume 168 Issue 4 March 2016 评论6
Natural variation of ambient dose rate in the air of Izu-Oshima Island after the Fukushima Daiichi N ...
[CTSC2017]密钥
传送门:http://uoj.ac/problem/297 “无论哪场比赛,都要相信题目是水的” 这不仅是HNOI2018D2T3的教训,也是这题的教训,思维定势真的很可怕. 普及组水题,真是愧对CT ...
【DFS】Codeforces Round #398 (Div. 2) C. Garland
设sum是所有灯泡的亮度之和有两种情况: 一种是存在结点U和V,U是V的祖先,并且U的子树权值和为sum/3*2,且U不是根,且V的子树权值和为sum/3. 另一种是存在结点U和V,他们之间没有祖先 ...
【bzoj1370】【团伙】原来并查集还能这么用？！
(画师当然是武内崇啦) Description 在某城市里住着n个人,任何两个认识的人不是朋友就是敌人,而且满足: 1. 我朋友的朋友是我的朋友: 2. 我敌人的敌人是我的朋友: 所有是朋友的人组成一 ...

Python 爬基金数据

Python 爬基金数据的更多相关文章

随机推荐

热门专题