python crawler

crawl blog website: www.apress.com

# -*- coding: utf-8 -*-

"""

Created on Wed May 10 18:01:41 2017

@author: Raghav Bali

"""

"""

This script crawls apress.com's blog page to:

    + extract list of recent blog post titles and their URLS

    + extract content related to each blog post in plain text

using requests and BeautifulSoup packages

``Execute``

        $ python crawl_bs.py

"""

import requests

from time import sleep

from bs4 import BeautifulSoup

def get_post_mapping(content):

    """This function extracts blog post title and url from response object

    Args:

        content (request.content): String content returned from requests.get

    Returns:

        list: a list of dictionaries with keys title and url

    """

    post_detail_list = []

    post_soup = BeautifulSoup(content,"lxml")

    h3_content = post_soup.find_all("h3")

    for h3 in h3_content:

        post_detail_list.append(

            {'title':h3.a.get_text(),'url':h3.a.attrs.get('href')}

            )

    return post_detail_list

def get_post_content(content):

    """This function extracts blog post content from response object

    Args:

        content (request.content): String content returned from requests.get

    Returns:

        str: blog's content in plain text

    """

    plain_text = ""

    text_soup = BeautifulSoup(content,"lxml")

    para_list = text_soup.find_all("div",

                                   {'class':'cms-richtext'})

    for p in para_list[0]:

        plain_text += p.getText()

    return plain_text

if __name__ =='__main__':

    crawl_url = "http://www.apress.com/in/blog/all-blog-posts"

    post_url_prefix = "http://www.apress.com"

    print("Crawling Apress.com for recent blog posts...\n\n")    

    response = requests.get(crawl_url)

    if response.status_code == 200:

        blog_post_details = get_post_mapping(response.content)

    if blog_post_details:

        print("Blog posts found:{}".format(len(blog_post_details)))

        for post in blog_post_details:

            print("Crawling content for post titled:",post.get('title'))

            post_response = requests.get(post_url_prefix+post.get('url'))

            if post_response.status_code == 200:

                post['content'] = get_post_content(post_response.content)

            print("Waiting for 10 secs before crawling next post...\n\n")

            sleep(10)

        print("Content crawled for all posts")

        # print/write content to file

        for post in blog_post_details:

            print(post)

python crawler的更多相关文章

Python crawler access to web pages the get requests a cookie
Python in the process of accessing the web page,encounter with cookie,so we need to get it. cookie i ...
【python爬虫】根据查询词爬取网站返回结果
最近在做语义方面的问题,需要反义词.就在网上找反义词大全之类的,但是大多不全,没有我想要的.然后就找相关的网站,发现了http://fanyici.xpcha.com/5f7x868lizu.html ...
python脚本工具－ 3 目录遍历
遍历系统中某一目录下的所有文件名 #! /usr/bin/python # coding:utf-8 import os def dirList(path): filelist = os.listdi ...
pyrailgun 0.24 : Python Package Index
pyrailgun 0.24 : Python Package Index pyrailgun 0.24 Download pyrailgun-0.24.zip Fast Crawler For Py ...
[Python]新手写爬虫全过程（转）
今天早上起来,第一件事情就是理一理今天该做的事情,瞬间get到任务,写一个只用python字符串内建函数的爬虫,定义为v1.0,开发中的版本号定义为v0.x.数据存放?这个是一个练手的玩具,就写在tx ...
python编写知乎爬虫实践
爬虫的基本流程网络爬虫的基本工作流程如下: 首先选取一部分精心挑选的种子URL 将种子URL加入任务队列从待抓取URL队列中取出待抓取的URL,解析DNS,并且得到主机的ip,并将URL对应的网页 ...
python爬虫之urllib
#coding=utf-8 #urllib操作类 import time import urllib.request import urllib.parse from urllib.error imp ...
Python实现自动登录/登出校园网网关
学校校园网的网络连接有免费连接和收费连接两种类型,可想而知收费连接浏览体验更佳,比如可以访问更多的网站.之前收费地址只能开通包月服务才可使用,后来居然有了每个月60小时的免费使用收费地址的优惠.但是, ...
python爬虫实践
模拟登陆与文件下载爬取http://moodle.tipdm.com上面的视频并下载模拟登陆由于泰迪杯网站问题,测试之后发现无法用正常的账号密码登陆,这里会使用访客账号登陆. 我们先打开泰迪杯的 ...

随机推荐

windows7下安装msys2
系统: windows 7 首先需要msys2的安装包,可以去官网下载安装包官网地址: http://www.msys2.org/本次下载的是 msys2-x86_64-20190524.exe 注意 ...
物联网学习笔记三：物联网网关协议比较：MQTT 和 Modbus
物联网学习笔记三:物联网网关协议比较:MQTT 和 Modbus 物联网 (IoT) 不只是新技术,还是与旧技术的集成,其关键在于通信.可用的通信方法各不相同,但是,各种不同的协议在将海量“事物”连接 ...
linux安装php nginx mysql
linux装软件方式: systemctl status firewalld.service 查看防火墙systemctl stop firewalld.service systemctl disab ...
Java Web项目搭建过程记录（struts2）
开发工具:eclipse 搭建环境:jdk1.7 tomcat 8.0 基础的java开发环境搭建过程不再赘述,下面从打开eclipse 之后的操作开始第一步: 创建项目,File -> ...
NumPy 之案例(随机漫步)
import numpy as np The numpy.random module supplements(补充) the built-in Python random with functions ...
Zepto——简化版jQuery，移动端首选js库
转载请注明原文地址:https://www.cnblogs.com/ygj0930/p/10826054.html 一:Zepto是什么 Zepto最初是为移动端开发的js库,是jQuery的轻量级替 ...
进程间通信之数据传输--Socket
The client server model Most interprocess communication uses the client server model. These terms re ...
如何测试Web服务.2
-->全文字数:2700,需要占用你几分钟的阅读时间 ,您也可以收藏后,时间充足时再阅读- -->上一节讲了<Web服务基础介绍>,本节介绍可用于测试web服务的开源测试工具. ...
PAT 乙级 1022.D进制的A+B C++/Java
1022 D进制的A+B (20 分) 题目来源输入两个非负 10 进制整数 A 和 B (≤),输出 A+B 的 D (1)进制数. 输入格式: 输入在一行中依次给出 3 个整数 A.B 和 D. ...
BFS算法的优化双向宽度优先搜索
双向宽度优先搜索 (Bidirectional BFS) 算法适用于如下的场景: 无向图所有边的长度都为 1 或者长度都一样同时给出了起点和终点以上 3 个条件都满足的时候,可以使用双向宽度优先 ...

python crawler

python crawler的更多相关文章

随机推荐

热门专题