Python-docx 读取word.docx内容
第一次写博客,也不知道要写点儿什么好,所以就把我在学习Python的过程中遇到的问题记录下来,以便之后查看,本人小白,写的不好,如有错误,还请大家批评指正!
中文编码问题总是让人头疼,想要用Python读取word中的内容,用open()经常报错,上网一搜结果发现了Python有专门读取.docx的模块python_docx(只能读取.docx文件,不能读取.doc文件),用起来很方便。
安装python-docx:
pip install python_docx
(注意:不是pip install docx ! docx也可以安装,但总是报错,缺少exceptions,无法导入)
接下来就可以用Python_docx 来读取word文本了。
代码如下:
import docx
from docx import Document
path = "C:\\Users\\Administrator\\Desktop\\word.docx"
document = Document(path)
for paragraph in document.paragraphs:
print(paragraph.text)
运行即可输出文本。
我尝试用docx读取.doc文本
代码如下:
import os
import docx
for filename in os.listdir(os.getcwd()):
if filename.endswith('.doc'):
print(filename[:-4])
doc = docx.Document(filename[:-4]+".docx")
for para in doc.paragraphs:
print (para.text)
结果报错:docx.opc.exceptions.PackageNotFoundError: Package not found。还是无法识别doc
引用1楼,“改变拓展名并没有改变其编码方式,因此无法读取文本内容,需将doc文件另存为docx文件后再用python-docx读取其内容”
# Document 还有添加标题、分页、段落、图片、章节等方法,说明如下
| add_heading(self, text='', level=1)
| Return a heading paragraph newly added to the end of the document,
| containing *text* and having its paragraph style determined by
| *level*. If *level* is 0, the style is set to `Title`. If *level* is
| 1 (or omitted), `Heading 1` is used. Otherwise the style is set to
| `Heading {level}`. Raises |ValueError| if *level* is outside the
| range 0-9.
|
| add_page_break(self)
| Return a paragraph newly added to the end of the document and
| containing only a page break.
|
| add_paragraph(self, text='', style=None)
| Return a paragraph newly added to the end of the document, populated
| with *text* and having paragraph style *style*. *text* can contain
| tab (``\t``) characters, which are converted to the appropriate XML
| form for a tab. *text* can also include newline (``\n``) or carriage
| return (``\r``) characters, each of which is converted to a line
| break.
|
| add_picture(self, image_path_or_stream, width=None, height=None)
| Return a new picture shape added in its own paragraph at the end of
| the document. The picture contains the image at
| *image_path_or_stream*, scaled based on *width* and *height*. If
| neither width nor height is specified, the picture appears at its
| native size. If only one is specified, it is used to compute
| a scaling factor that is then applied to the unspecified dimension,
| preserving the aspect ratio of the image. The native size of the
| picture is calculated using the dots-per-inch (dpi) value specified
| in the image file, defaulting to 72 dpi if no value is specified, as
| is often the case.
|
| add_section(self, start_type=2)
| Return a |Section| object representing a new section added at the end
| of the document. The optional *start_type* argument must be a member
| of the :ref:`WdSectionStart` enumeration, and defaults to
| ``WD_SECTION.NEW_PAGE`` if not provided.
|
| add_table(self, rows, cols, style=None)
| Add a table having row and column counts of *rows* and *cols*
| respectively and table style of *style*. *style* may be a paragraph
| style object or a paragraph style name. If *style* is |None|, the
| table inherits the default table style of the document.
|
| save(self, path_or_stream)
| Save this document to *path_or_stream*, which can be eit a path to
| a filesystem location (a string) or a file-like object.
docx还有许多其它功能,还正在学习中,详见官方文档:https://python-docx.readthedocs.io/en/latest/user/quickstart.html
Python-docx 读取word.docx内容的更多相关文章
- python读取word表格内容(1)
1.首页介绍下word表格内容,实例如下: 每两个表格后面是一个合并的单元格
- poi读取word的内容
pache POI是Apache软件基金会的开放源码函式库,POI提供API给Java程序对Microsoft Office格式档案读和写的功能. 1.读取word 2003及word 2007需要的 ...
- Python configparser 读取指定节点内容失败
# !/user/bin/python # -*- coding: utf-8 -*- import configparser # 生成一个config文件 config = configparser ...
- java 实现poi方式读取word文件内容
1.下载poi的jar包 下载地址:https://www.apache.org/dyn/closer.lua/poi/release/bin/poi-bin-3.17-20170915.tar.gz ...
- aspose.word 读取word段落内容
注:转载请标明文章原始出处及作者信息 aspose.word 插件下载 链接: http://pan.baidu.com/s/1qXIgOXY 密码: wsj2 使用原因:无需安装office,无兼容 ...
- Python中读取csv文件内容方法
gg 224@126.com 85 男 dd 123@126.com 52 女 fgf 125@126.com 23 女 csv文件内容如上图,首先导入csv包,调用csv中的方法reader()创建 ...
- python读取word中的段落、表、图+++++++++++Doc转换Docx
读取文本.图.表.解压信息 import docx import zipfile import os import shutil '''读取word中的文本''' def gettxt(): file ...
- Python 读取word中表格数据、读取word修改并保存、替换word中词汇、读取word中每段内容,读取一段话中相同样式内容,理解Document中run
from docx import Document path = r'D:\pywork\12' # word信息表所在文件夹 w = Document(path + '/' + 'word信息表.d ...
- 使用poi读取word2007(.docx)中的复杂表格
使用poi读取word2007(.docx)中的复杂表格 最近工作需要做一个读取word(.docx)中的表格,并以html形式输出.经过上网查询,使用了poi. 对于2007及之后的word文档,需 ...
随机推荐
- 容器(docker)内运行Nginx
容器内运行nginx其实很简单,但是一开始还是浪费了我很多时间.这里写下来给大家省点时间. 1.创建nginx文件夹,放置各种配置及日志等. mkdir /docker/nginx docker 文件 ...
- SQLServer 2008 已成功与服务器建立连接,但是在登录前的握手期间发生错误。 (provider: SSL Provider, error: 0 - 等待的操作过时。
在用SQL Server 2008 在连接其他电脑的实例时,一直提示“已成功与服务器建立连接,但是在登录前的握手期间发生错误. (provider: SSL Provider, error: 0 - ...
- leetcode每日刷题计划-简单篇day9
Num 38 报数 Count and Say 题意读起来比较费劲..看懂了题还是不难的 注意最后的长度是sz的长度,开始写错写的len 在下次计算的时候len要更新下 说明 直接让char和int进 ...
- 使用shell命令给文件中每一行的前面、后面添加字符
shell command shell给一个文件中的每一行开头插入字符的方法:awk '{print "xxx"$0}' fileName shell给一个文件中的每一行结尾插入字 ...
- orcal - 分组
执行顺序 from where group by having select order by 多表查询与分组查询的时候,查询结果相当于是一张临时表,所有的分组是在临时表操作 分组统计查询 COUNT ...
- mysql5.0手动升级8.0.15,并链接到navicat
一.卸载老版本的mysql 1.1 在控制面板中删除即可 1.2 将老版本的mysql安装残留文件彻底删除 二.彻底删除mysql-注册表 2.1 开始->运行-> regedit 看看注 ...
- 001_angular4.0框架学习
1. Cannot find module 'angular2-in-memory-web-api' 报这个错误的时候 是没有安装这个包 要手动安装下包 命令: npm i angular-in ...
- 新建DataTable
//创建DataTable DataTable dt = new DataTable("NewDt"); //创建自增长的ID列 DataColumn dc = dt.Column ...
- 查看log日志
本地环境的的log日志 可以直接查看, 对于新手来说怎么查看正式环境下的log日志呢 1, SSH到服务器 2,cd 到logs所在目录 3, tail -f 对应日志名字
- js:基于原生js的上啦下啦刷新功能
链接:https://www.jianshu.com/p/a8392115e6f0演示地址:http://wonghan.cn/iscroll-demo/html:<body> <d ...