python BeautifulSoup4解析网页

html = """

<html><head><title>The Dormouse's story</title></head>

<body>

<p class="title" name="dromouse"><b>The Dormouse's story</b></p>

<p class="story">Once upon a time there were three little sisters; and their names were

<a href="http://example.com/elsie" class="sister" id="link1"><!-- Elsie --></a><a href="http://example.com/lacie" class="sister" id="link2">Lacie</a> and

<a href="http://example.com/tillie" class="sister" id="link3">Tillie</a>

and they lived at the bottom of a well.</p>

<p class="story">...</p></body></html>

"""

soup=BS(html,'html.parser')

for i in soup.find_all('a'):

    print('i.text:',i.text)#注释掉的内容就不打印了  str类型

    print('i.string:',i.string)  #注释掉的内容 都会打印出来，NavigableString对象

print('soup.head.contents:',soup.head.contents,type(soup.head.contents))

print('soup.head.children:',soup.head.children,type(soup.head.children))

print('soup.body.contents:',soup.body.contents)#返回一个子元素的列表

print('soup.body.children:',soup.body.children)#返回一个子元素的迭代器

for i in soup.body.children:

    print(i)

print('子孙节点 都显示出来')

for i in soup.body.descendants:

    print(i)

print('soup.body.string:',soup.body.string)

print('soup.body.strings:',soup.body.strings)

print('soup.body.stripped_strings:',soup.body.stripped_strings)  #过滤掉所有空格显示

print('去掉空格的body子元素：')

for i  in soup.body.stripped_strings:

    print(i)

print('soup.a.parent:',soup.a.parent)

print('soup.a.next_sibling:',soup.a.next_sibling)  #注意文本节点、换行\n都可能成为当前节点的上一个或者下一个同级节点

print('soup.a.previous_sibling:',soup.a.previous_sibling)

print('soup.a.next_element:',soup.a.next_element)  #下一个元素 不一定同级

print('soup.a.previous_element:',soup.a.previous_element)

print('打印所有后面的同级节点:\n')

for i in soup.a.next_siblings:

    print(i)

print('soup.a.next_element:',list(soup.a.next_elements)[1])

print('***********find_all*****')

print(soup.find_all('a'))

print('引入正则表达式：')

import re

print(soup.find_all(re.compile(r'^title')))  #正则匹配的是 标签的名字

print('列表的方式匹配：')

print(soup.find_all(['a','b']))

print('函数的方式匹配，类似filter')

def func(tag):

    if tag.has_attr('class') and re.search(r'^a',tag.name):

        return tag

print(soup.find_all(func))

html = """

<html><head><title>The Dormouse's story</title></head>

<body>

<p class="title" name="dromouse"><b>The Dormouse's story</b></p>

<p class="story">Once upon a time there were three little sisters; and their names were

<a href="http://example.com/elsie" class="sister" id="link1"><!-- Elsie --></a><a href="http://example.com/lacie" class="sister" id="link2">Lacie</a> and

<a href="http://example.com/tillie" class="sister" id="link3">Tillie</a>

and they lived at the bottom of a well.</p>

<p class="story">...</p></body></html>

"""

soup=BS(html,'html.parser')

print('按属性值查找:')

print(soup.find_all(id='link1'))

print(soup.find_all('a',id='link1'))

print(soup.find_all(id='link2',href=re.compile(r'laci')))  #返回的都是列表

print(soup.find_all(class_='story')) #注意后面加的下划线

print(soup.find_all(attrs={'class':'sister'}))

print('按元素内容查找text参数：')

print(soup.find_all(text='Tillie'))

print(soup.find_all(text=['Tillie','Lacie']))  #返回的都是元素内容

print(soup.find_all(text=re.compile(r'ormous')))

print('通过内容元素 找到上级元素')

print(soup.find_all(text=re.compile(r'ormous'))[1].parent.parent)

#限制查找数量

print('limit:')

print(soup.find_all('a',limit=2))

print('只在子节点查找：')

print(soup.body.find_all('a',limit=2,recursive=False))  #只查找子节点 recursive循环的、递归的

print(soup.body.find_all(class_='story',recursive=False))

python BeautifulSoup4解析网页的更多相关文章

Python爬虫解析网页的4种方式值得收藏
用Python写爬虫工具在现在是一种司空见惯的事情,每个人都希望能够写一段程序去互联网上扒一点资料下来,用于数据分析或者干点别的事情. 我们知道,爬虫的原理无非是把目标网址的内容下载下来存储到内存 ...
python bs4解析网页时 bs4.FeatureNotFound: Couldn't find a tree builder with the features you requested: lxml. Do you need to inst（转）
Python小白,学习时候用到bs4解析网站,报错 bs4.FeatureNotFound: Couldn't find a tree builder with the features you re ...
使用Python中的urlparse、urllib抓取和解析网页（一）（转）
对搜索引擎.文件索引.文档转换.数据检索.站点备份或迁移等应用程序来说,经常用到对网页(即HTML文件)的解析处理.事实上,通过Python 语言提供的各种模块,我们无需借助Web服务器或者Web浏览 ...
python网络爬虫之解析网页的BeautifulSoup(爬取电影图片)[三]
目录前言一.BeautifulSoup的基本语法二.爬取网页图片扩展学习后记前言本章同样是解析一个网页的结构信息在上章内容中(python网络爬虫之解析网页的正则表达式(爬取4k动漫图 ...
Python中的urlparse、urllib抓取和解析网页（一）
对搜索引擎.文件索引.文档转换.数据检索.站点备份或迁移等应用程序来说,经常用到对网页(即HTML文件)的解析处理.事实上,通过Python 语言提供的各种模块,我们无需借助Web服务器或者Web浏览 ...
python网络爬虫之解析网页的XPath(爬取Path职位信息)[三]
目录前言 XPath的使用方法 XPath爬取数据后言 @(目录) 前言本章同样是解析网页,不过使用的解析技术为XPath. 相对于之前的BeautifulSoup,我感觉还行,也是一个比较常用 ...
python网络爬虫-解析网页（六）
解析网页主要使用到3种方法提取网页中的数据,分别是正则表达式.beautifulsoup和lxml. 使用正则表达式解析网页正则表达式是对字符串操作的逻辑公式 .代替任意字符 . *匹配前0个或多 ...
Python爬虫之解析网页
常用的类库为lxml, BeautifulSoup, re(正则) 以获取豆瓣电影正在热映的电影名为例,url='https://movie.douban.com/cinema/nowplaying/ ...
[技术博客] BeautifulSoup4分析网页
[技术博客] BeautifulSoup4分析网页使用BeautifulSoup4进行网页文本分析前言进行网络爬虫时我们需要从网页源代码中提取自己所需要的信息,分析整理后存入数据库中. 在pyt ...

随机推荐

常用HTML转义字符,html转义符,JavaScript转义符,html转义字符表,HTML语言特殊字符对照表(ISO Latin-1字符集)
HTML字符实体(Character Entities),转义字符串(Escape Sequence) 为什么要用转义字符串? HTML中<,>,&等有特殊含义(<,> ...
ACS712电流传感器应用
1. 原理图其中第7脚输出的是电压值,那么电压值和测量的电流什么关系?看下图,有3个量程,我用的是20A电流的,100mv电压对应1A电流看下图,不同的温度会有影响,不过区别不大最后计算的公式是 ...
Socket测试工具（客户端、服务端）
Socket是什么? SOCKET用于在两个基于TCP/IP协议的应用程序之间相互通信.最早出现在UNIX系统中,是UNIX系统主要的信息传递方式.在WINDOWS系统中,SOCKET称为WINSOC ...
高级UI-CardView
CardView是在Android 5.0推出的新控件,为了兼容之前的版本,将其放在了v7包里面,在现在扁平化设计潮流的驱使下,越来越多的软件使用到了CardView这一控件,那么这篇文章就来看看Ca ...
List<E>
List<E>——列表有序,存储和读取的顺序是一致的由整数索引允许重复 add(int index,E element)——将元素插入指定位置 get(int index)——获取指 ...
RabbitMQ基础命令rabbitmqctl
官网文档 https://www.rabbitmq.com/rabbitmqctl.8.html 一般操作命令后台管理页面都有的,部分没有(应用程序管理,和集群管理). 直接使用命令,必须配置环境变量 ...
计算机网络自顶向下方法第4章网络层:数据平面 (Network layer)
4.1 网络层概述网络层主要功能为转发(将数据从路由器输入接口转移到合适的输出接口)和路由选择(端到端的路径选择),每台路由器都有一张转发表,用最长前缀匹配规则来转发. 4.1.1 转发和路由选择 ...
Shiro身份认证、盐加密
目的: Shiro认证盐加密工具类 Shiro认证 1.导入pom依赖 <dependency> <groupId>org.apache.shiro</groupId& ...
应用编排服务之ELK技术栈示例模板详解
日志对互联网应用的运维尤为重要,它可以帮助我们了解服务的运行状态.了解数据流量来源甚至可以帮助我们分析用户的行为等.当进行故障排查时,我们希望能够快速的进行日志查询和过滤,以便精准的定位并解决问题. ...
idea .gitignore模板
IDEA 创建的项目,需要搞个.gitignore文件,文件内容可以参考插件的. # Created by .ignore support plugin (hsz.mobi) ### JetBrain ...

python BeautifulSoup4解析网页

python BeautifulSoup4解析网页的更多相关文章

随机推荐

热门专题