零基础Python爬虫实现(百度贴吧)

提示:本学习来自Ehco前辈的文章, 经过实现得出的笔记。

目标

http://tieba.baidu.com/f?kw=linux&ie=utf-8

网站结构

学习目标

由于是第一个实验性质爬虫，我们要做的不多，我们需要做的就是：

1. 从网上爬下特定页码的网页

2. 对于爬下的页面内容进行简单的筛选分析

3. 找到每一篇帖子的 标题、发帖人、日期、楼层、以及跳转链接

4. 将结果保存到文本。

发现规律

&pn=0 ： 首页

&pn=50： 第二页

&pn=100：第三页

&pn=50*n 第n页

50 表示 每一页都有50篇帖子。
这样就能实现翻页操作

附上代码

import requests

import time

from bs4 import BeautifulSoup

def get_html(url):

    try:

        r = requests.get(url, timeout=30)

        r.raise_for_status()

        r.encoding = 'utf-8'

        return r.text

    except:

        return "error"

def get_content(url):

    comments = []

    html = get_html(url)

    soup = BeautifulSoup(html, 'lxml')

    liTags = soup.find_all('li', attrs={'class':' j_thread_list clearfix'})

    for li in liTags:

        comment = {}

        try:

            #标题

            comment['title'] = li.find(

                'a', attrs={'class':'j_th_tit '}).text.strip()

            #链接

            comment['link'] = "http://tieba.baidu.com/" + \

                li.find('a', attrs={'class' : 'j_th_tit'})['href']

            #发帖人

            comment['name'] = li.find(

                'span', attrs = {'class':'tb_icon_author '}

            ).text.strip()

            #发帖时间

            comment['time'] = li.find(

                'span', attrs={'class':'pull-right is_show_create_time'}

            ).text.strip()

            #回复数量

            comment['replyNum'] = li.find(

                'span', attrs={'class':'threadlist_rep_num center_text'}

            ).text.strip()

            comments.append(comment)

        except:

            print("出了点小问题")

    return comments

def Out2File(dict):

    with open('TTBT.txt', 'a+') as f:

        for comment in dict:

            f.write('标题: {} \t 连接: {} \t 发帖人: {} \t 发帖时间: {} \t 回复数量: {} \n'.format(

                comment['title'], comment['link'], comment['name'], comment['time'], comment['replyNum']

            ))

        print("当前页面爬取完成")

def main(base_url, deep):

    url_list = []

    for i in range(0, deep):

        url_list.append(base_url + '&pn' + str(50 * i))

    print("所有的网页已经下载到本地! 开始筛选信息")

    for url in url_list:

        content = get_content(url)

        Out2File(content)

    print("所有的信息都已经保存完毕")

base_url = 'http://tieba.baidu.com/f?kw=linux&ie=utf-8'

deep = 3

if __name__ == '__main__':

    main(base_url, deep)

结果

零基础Python爬虫实现(百度贴吧)的更多相关文章

零基础Python爬虫实现(爬取最新电影排行)
提示:本学习来自Ehco前辈的文章, 经过实现得出的笔记. 目标网站 http://dianying.2345.com/top/ 网站结构要爬的部分,在ul标签下(包括li标签), 大致来说迭代li ...
嵩天老师的零基础Python笔记：https://www.bilibili.com/video/av15123607/?from=search&seid=10211084839195730432#page=25 中的42-45讲 {字典}
#coding=gbk#嵩天老师的零基础Python笔记:https://www.bilibili.com/video/av15123607/?from=search&seid=1021108 ...
嵩天老师的零基础Python笔记：https://www.bilibili.com/video/av13570243/?from=search&seid=15873837810484552531 中的15-23讲
#coding=gbk#嵩天老师的零基础Python笔记:https://www.bilibili.com/video/av13570243/?from=search&seid=1587383 ...
嵩天老师的零基础Python笔记：https://www.bilibili.com/video/av13570243/?from=search&seid=15873837810484552531 中的1-14讲
#coding=gbk#嵩天老师的零基础Python笔记:https://www.bilibili.com/video/av13570243/?from=search&seid=1587383 ...
零基础Python应该怎样学习呢？（附视频教程）
Python应该怎样学习呢? 阶段一:适合自己的学习方式对于零基础的初学者来说,最迷茫的是不知道怎样开始学习?那这里小编建议可以采用视频+书籍的方式进行学习.看视频学习可以让你迅速掌握编程的基础语法 ...
如何用Python爬虫实现百度图片自动下载？
Github:https://github.com/nnngu/LearningNotes 制作爬虫的步骤制作一个爬虫一般分以下几个步骤: 分析需求分析网页源代码,配合开发者工具编写正则表达式或 ...
python爬虫获取百度图片（没有精华，只为娱乐）
python3.7,爬虫技术,获取百度图片资源,msg为查询内容,cnt为查询的页数,大家快点来爬起来.注:现在只能爬取到百度的小图片,以后有大图片的方法,我会陆续发贴. #!/usr/bin/env ...
【学习笔记】第二章 python安全编程基础---python爬虫基础（urllib）
一.爬虫基础 1.爬虫概念网络爬虫(又称为网页蜘蛛),是一种按照一定的规则,自动地抓取万维网信息的程序或脚本.用爬虫最大的好出是批量且自动化得获取和处理信息.对于宏观或微观的情况都可以多一个侧面去了 ...
零基础Python接口测试教程
目录一.Python基础 Python简介.环境搭建及包管理 Python基本语法基本数据类型(6种) 条件/循环文件读写(文本文件) 函数/类模块/包常见算法二.接口测试快速实践简单接 ...

随机推荐

is_readable() 函数检查指定的文件是否可读。
定义和用法 is_readable() 函数判断指定文件名是否可读. 语法 is_readable(file) 参数描述 file 必需.规定要检查的文件. 说明如果由 file 指定的文件或目录 ...
nodejs发送邮件
这里我主要使用的是 nodemailer 这个插件第一步下载依赖 cnpm install nodemailer --save 第二步建立email.js 'use strict'; const ...
JavaScript-switch-case-电话系统
<!DOCTYPE html> <html> <head lang="en"> <meta charset="UTF-8&quo ...
java 基础功能
1.str.length();// 获取整个字符串的长度 public class Test { public static void main(String[] args) { String s = ...
java中垃圾回收机制中的引用计数法和可达性分析法（最详细）
首先,我这是抄写过来的,写得真的很好很好,是我看过关于GC方面讲解最清楚明白的一篇.原文地址是:https://www.zhihu.com/question/21539353
[ Windows BAT Script ] 删除某个目录下的所有某类文件
删除某个目录下的所有某类文件 @echo off for /R %%s in (*.txt) do ( echo %%s del %%s ) pause @echo on
Mvcpager以下各节已定义，但尚未为布局页“~/Views/Shared/_Layout.cshtml”呈现:“Scripts”。
解决办法如下: 1.在_Layout.cshtml布局body内,添加section,Scripts.Render和RenderSection标签示例代码如下: <body class=&quo ...
Centos 6.5初始化配置
安装好centos 6.5 # -*- coding:utf-8 -*- import win32api import time import os from Tkinter import * are ...
html5-垂直定位
*{ padding: 0px; margin: 0px; }#div2{ background: green; padding: 15px; width: 200px; ...
华为手机安装 charles 证书

零基础Python爬虫实现(百度贴吧)

目标

网站结构

学习目标

发现规律

附上代码

结果

零基础Python爬虫实现(百度贴吧)的更多相关文章

随机推荐

热门专题