[Python] Wikipedia Crawler

import time

import urllib

import bs4

import requests

start_url = "https://en.wikipedia.org/wiki/Special:Random"

target_url = "https://en.wikipedia.org/wiki/Philosophy"

def find_first_link(url):

    response = requests.get(url)

    html = response.text

    soup = bs4.BeautifulSoup(html, "html.parser")

    # This div contains the article's body

    # (June 2017 Note: Body nested in two div tags)

    content_div = soup.find(id="mw-content-text").find(class_="mw-parser-output")

    # stores the first link found in the article, if the article contains no

    # links this value will remain None

    article_link = None

    # Find all the direct children of content_div that are paragraphs

    for element in content_div.find_all("p", recursive=False):

        # Find the first anchor tag that's a direct child of a paragraph.

        # It's important to only look at direct children, because other types

        # of link, e.g. footnotes and pronunciation, could come before the

        # first link to an article. Those other link types aren't direct

        # children though, they're in divs of various classes.

        if element.find("a", recursive=False):

            article_link = element.find("a", recursive=False).get('href')

            break

    if not article_link:

        return

    # Build a full url from the relative article_link url

    first_link = urllib.parse.urljoin('https://en.wikipedia.org/', article_link)

    return first_link

def continue_crawl(search_history, target_url, max_steps=25):

    if search_history[-1] == target_url:

        print("We've found the target article!")

        return False

    elif len(search_history) > max_steps:

        print("The search has gone on suspiciously long, aborting search!")

        return False

    elif search_history[-1] in search_history[:-1]:

        print("We've arrived at an article we've already seen, aborting search!")

        return False

    else:

        return True

article_chain = [start_url]

while continue_crawl(article_chain, target_url):

    print(article_chain[-1])

    first_link = find_first_link(article_chain[-1])

    if not first_link:

        print("We've arrived at an article with no links, aborting search!")

        break

    article_chain.append(first_link)

    time.sleep(2) # Slow things down so as to not hammer Wikipedia's servers

[Python] Wikipedia Crawler的更多相关文章

Python Web Crawler
Python版本:3.5.2 pycharm URL Parsing¶ https://docs.python.org/3.5/library/urllib.parse.html?highlight= ...
【Python五篇慢慢弹】快速上手学python
快速上手学python 作者:白宁超 2016年10月4日19:59:39 摘要:python语言俨然不算新技术,七八年前甚至更早已有很多人研习,只是没有现在流行罢了.之所以当下如此盛行,我想肯定是多 ...
python百科
Python 编辑词条添加义项名 B 添加义项 ? Python(英语发音:/ˈpaɪθən/), 是一种面向对象.解释型计算机程序设计语言,由Guido van Rossum于1989年底发明,第 ...
Python in minute
Python 性能优化相关专题: https://www.ibm.com/developerworks/cn/linux/l-cn-python-optim/ Python wikipedi ...
Python网络数据采集7-单元测试与Selenium自动化测试
Python网络数据采集7-单元测试与Selenium自动化测试单元测试 Python中使用内置库unittest可完成单元测试.只要继承unittest.TestCase类,就可以实现下面的功能. ...
深入了解Python
一.Python的风格 Python在设计上坚持了清晰划一的风格,这使得Python成为一门易读.易维护,并且被大量用户所欢迎的.用途广泛的语言. 设计者开发时总的指导思想是,对于一个特定的问题,只要 ...
######【Python】【基础知识】Python的介绍 ######
Python 是一种面向对象.解释型计算机程序设计语言. Python是什么? Python(英国发音:/ˈpaɪθən/ 美国发音:/ˈpaɪθɑːn/), 是一种面向对象的解释型计算机程序设计语言 ...
所有selenium相关的库
通过爬虫获取官方文档库如果想获取相应的库修改对应配置即可代码如下 from urllib.parse import urljoin import requests from lxml im ...
500lines项目简介
"500行或更少" "What I cannot create, I do not understand." -- Richard Feynman <50 ...

随机推荐

Linux-php7安装redis
Linux-php7安装redis 标签(空格分隔): 未分类安装redis服务 1 下载redis cd /usr/local/ 进入安装目录 wget http://download.redis ...
java9新特性-11-String存储结构变更
1. 官方Feature JEP254: Compact Strings 2. 产生背景 Motivation The current implementation of the String cla ...
ubuntu12.04
最近越来越觉得必须用Linux了,于是装了15.04,好不习惯的感觉,思维还是10.10的时代. 尝试做种http://jingyan.baidu.com/article/a681b0dedad55c ...
【转载】eclipse中批量修改Java类文件中引入的package包路径
原博客地址:http://my.oschina.net/leeoo/blog/37852 当复制其他工程中的包到新工程的目录中时,由于包路径不同,出现红叉,下面的类要一个一个修改包路径,类文件太多的话 ...
ivew语法中'${}`的用法
BootStrap--from 表单
1 垂直表单(默认) 2 内联表单 3 水平表单使用 class .sr-only,您可以隐藏内联表单的标签. 垂直或基本表单基本的表单结构是 Bootstrap 自带的,个别的表单控件自动接收一 ...
解决mongodb TypeError: Cannot read property 'XXX' of null 问题
有时候我们的字段里的数据为空而去查询就会报错. 比如以下形式: “cartList”:[] “cartList”:[{}] “cartList”:{} “cartList”:null 在查询的时候加上 ...
java 线程传参方式
第一类:主动向线程传参 public class ThreadTest extends Thread { public ThreadTest() { } /** * 第一种通过构造方法来传递参数 ...
laravel中soapServer支持wsdl的例子
最近在对接客户的CRM系统,获取令牌时,要用DES方式加密解密,由于之前没有搞错这种加密方式,经过请教了"百度"和"谷歌"两个老师后,结合了多篇文档内容后,终于 ...
Redis序列化存储Java集合List等自定义类型
在"Redis学习总结和相关资料"http://blog.csdn.net/fansunion/article/details/49278209 这篇文章中,对Redis做了总体的 ...

[Python] Wikipedia Crawler

[Python] Wikipedia Crawler的更多相关文章

随机推荐

热门专题