[Python] Wikipedia Crawler

import time

import urllib

import bs4

import requests

start_url = "https://en.wikipedia.org/wiki/Special:Random"

target_url = "https://en.wikipedia.org/wiki/Philosophy"

def find_first_link(url):

    response = requests.get(url)

    html = response.text

    soup = bs4.BeautifulSoup(html, "html.parser")

    # This div contains the article's body

    # (June 2017 Note: Body nested in two div tags)

    content_div = soup.find(id="mw-content-text").find(class_="mw-parser-output")

    # stores the first link found in the article, if the article contains no

    # links this value will remain None

    article_link = None

    # Find all the direct children of content_div that are paragraphs

    for element in content_div.find_all("p", recursive=False):

        # Find the first anchor tag that's a direct child of a paragraph.

        # It's important to only look at direct children, because other types

        # of link, e.g. footnotes and pronunciation, could come before the

        # first link to an article. Those other link types aren't direct

        # children though, they're in divs of various classes.

        if element.find("a", recursive=False):

            article_link = element.find("a", recursive=False).get('href')

            break

    if not article_link:

        return

    # Build a full url from the relative article_link url

    first_link = urllib.parse.urljoin('https://en.wikipedia.org/', article_link)

    return first_link

def continue_crawl(search_history, target_url, max_steps=25):

    if search_history[-1] == target_url:

        print("We've found the target article!")

        return False

    elif len(search_history) > max_steps:

        print("The search has gone on suspiciously long, aborting search!")

        return False

    elif search_history[-1] in search_history[:-1]:

        print("We've arrived at an article we've already seen, aborting search!")

        return False

    else:

        return True

article_chain = [start_url]

while continue_crawl(article_chain, target_url):

    print(article_chain[-1])

    first_link = find_first_link(article_chain[-1])

    if not first_link:

        print("We've arrived at an article with no links, aborting search!")

        break

    article_chain.append(first_link)

    time.sleep(2) # Slow things down so as to not hammer Wikipedia's servers

[Python] Wikipedia Crawler的更多相关文章

Python Web Crawler
Python版本:3.5.2 pycharm URL Parsing¶ https://docs.python.org/3.5/library/urllib.parse.html?highlight= ...
【Python五篇慢慢弹】快速上手学python
快速上手学python 作者:白宁超 2016年10月4日19:59:39 摘要:python语言俨然不算新技术,七八年前甚至更早已有很多人研习,只是没有现在流行罢了.之所以当下如此盛行,我想肯定是多 ...
python百科
Python 编辑词条添加义项名 B 添加义项 ? Python(英语发音:/ˈpaɪθən/), 是一种面向对象.解释型计算机程序设计语言,由Guido van Rossum于1989年底发明,第 ...
Python in minute
Python 性能优化相关专题: https://www.ibm.com/developerworks/cn/linux/l-cn-python-optim/ Python wikipedi ...
Python网络数据采集7-单元测试与Selenium自动化测试
Python网络数据采集7-单元测试与Selenium自动化测试单元测试 Python中使用内置库unittest可完成单元测试.只要继承unittest.TestCase类,就可以实现下面的功能. ...
深入了解Python
一.Python的风格 Python在设计上坚持了清晰划一的风格,这使得Python成为一门易读.易维护,并且被大量用户所欢迎的.用途广泛的语言. 设计者开发时总的指导思想是,对于一个特定的问题,只要 ...
######【Python】【基础知识】Python的介绍 ######
Python 是一种面向对象.解释型计算机程序设计语言. Python是什么? Python(英国发音:/ˈpaɪθən/ 美国发音:/ˈpaɪθɑːn/), 是一种面向对象的解释型计算机程序设计语言 ...
所有selenium相关的库
通过爬虫获取官方文档库如果想获取相应的库修改对应配置即可代码如下 from urllib.parse import urljoin import requests from lxml im ...
500lines项目简介
"500行或更少" "What I cannot create, I do not understand." -- Richard Feynman <50 ...

随机推荐

ubuntu升级到14.04后终端显示重叠
系统升级后,发现这个问题非常不爽,问题不大,但有时候找不到解决方法,让人纠结好久.解决方法例如以下: 编辑->配置文件首选项->常规-> monospace 改为ubuntu mon ...
使用iOS原生sqlite3框架对sqlite数据库进行操作
摘要: iOS中sqlite3框架可以很好的对sqlite数据库进行支持,通过面向对象的封装,可以更易于开发者使用. 使用iOS原生sqlite3框架对sqlite数据库进行操作一.引言 sqlit ...
远程登录工具 —— filezilla（FTP vs. SFTP）、xshell、secureCRT
filezilla:是一个免费开源的 FTP 软件,分为客户端版本和服务器版本,具备所有的 FTP 软件功能. 支持的协议:FTP & SFTP(Secure File Transfer Pr ...
django笔记10 cookie整理
感谢武沛齐老师 Alex老师 cookie 没有cookie所有的网站都登录不上客户端浏览器上的一个文件 {'user':'ljc'} {"user":'zpt'} reques ...
关于Fragment的setUserVisibleHint() 方法和onCreateView（）的执行顺序
1:setUserVisibleHint(boolean isVisibleToUser)的方法就很重要,根据方法名来看当前页面是否可见, 里面的boolean值就是判断当前页面是否可见的变量,所以大 ...
ACM-ICPC 2016 Qingdao Preliminary Contest
A I Count Two Three I will show you the most popular board game in the Shanghai Ingress Resistance T ...
jQuery的效果函数
jQuery的效果函数有很多,下面让我们一起看看jQuery的效果函数吧: jQuery的效果函数列表: animate():对被选元素应用“自定义”的动画. clearQueue():对被选元素移除 ...
题解 CF896C 【Willem, Chtholly and Seniorious】
貌似珂朵莉树是目前为止(我学过的)唯一一个可以维护区间x次方和查询的高效数据结构. 但是这玩意有个很大的毛病,就是它的高效建立在数据随机的前提下. 在数据随机的时候assign操作比较多,所以它的复杂 ...
centos7 安装rsyslog
http://blog.csdn.net/u011630575/article/details/50896781 http://blog.chinaunix.net/uid-21142030-id-5 ...
C#文件拖放至窗口的ListView控件获取文件类型
using System; using System.Collections.Generic; using System.ComponentModel; using System.Data; usin ...

[Python] Wikipedia Crawler

[Python] Wikipedia Crawler的更多相关文章

随机推荐

热门专题