英文词频统计的java实现方法
需求概要
1.读取文件,文件内包可含英文字符,及常见标点,空格级换行符。
2.统计英文单词在本文件的出现次数
3.将统计结果排序
4.显示排序结果
分析
1.读取文件可使用BufferedReader类按行读取
2.针对读入行根据分隔符拆分出单词,使用java.util工具提供的Map记录单词和其出现次数的信息,HashMap和TreeMap均可,如果排序结果按字母序可选用TreeMap,本例选择用HashMap。
3.将Map中存储的键值对,存入List中,List中的元素为二维字符串数组,并将单词出现次数作为比对依据,利用list提供的sort方法排序,需要实现Comparator.compare接口方法。
4.循环遍历List中的String[],输出结果。
部分功能实现
输出结果
public void printSortedWordGroupCount(String filename) {
List<String[]> result = getSortedWordGroupCount(filename);
if (result == null) {
System.out.println("no result");
return;
}
for (String[] sa : result) {
System.out.println(sa[1] + ": " + sa[0]);
}
}
统计
public Map<String, Integer> getWordGroupCount(String filename) {
try {
FileReader fr = new FileReader(filename);
BufferedReader br = new BufferedReader(fr);
String content = "";
Map<String, Integer> result = new HashMap<String, Integer>();
while ((content = br.readLine()) != null) {
StringTokenizer st = new StringTokenizer(content, "!&(){}+-= ':;<> /\",");
while (st.hasMoreTokens()) {
String key = st.nextToken();
if (result.containsKey(key))
result.put(key, result.get(key) + 1);
else
result.put(key, 1);
}
}
br.close();
fr.close();
return result;
} catch (FileNotFoundException e) {
System.out.println("failed to open file:" + filename);
e.printStackTrace();
} catch (Exception e) {
System.out.println("some expection occured");
e.printStackTrace();
}
return null;
}
排序
public List<String[]> getSortedWordGroupCount(String filename) {
Map<String, Integer> result = getWordGroupCount(filename);
if (result == null)
return null;
List<String[]> list = new LinkedList<String[]>();
Set<String> keys = result.keySet();
Iterator<String> iter = keys.iterator();
while (iter.hasNext()) {
String key = iter.next();
String[] item = new String[2];
item[1] = key;
item[0] = "" + result.get(key);
list.add(item);
}
list.sort(new Comparator<String[]>() {
public int compare(String[] s1, String[] s2) {
return Integer.parseInt(s2[0])-Integer.parseInt(s1[0]); }
});
return list;
}
针对本程序的简易测试用例自动生成
使用随机数产生随机单词,空格,并在随机位置插入回车
public boolean gernerateWordList(String filePosition, long wordNum) {
FileWriter fw = null;
StringBuffer sb = new StringBuffer("");
try {
fw = new FileWriter(filePosition);
for (int j = 0; j < wordNum; j++) {
int length = (int) (Math.random() * 10 + 1);
for (int i = 0; i < length; i++) {
char ch = (char) ('a' + (int) (Math.random() * 26));
sb.append(ch);
}
fw.write(sb.toString());
fw.write(" ");
sb = new StringBuffer("");
if (wordNum % (int) (Math.random() * 8 + 4) == 0)
fw.write("\n");
}
} catch (IOException e) {
e.printStackTrace();
return false;
} finally {
try {
fw.close();
} catch (IOException e) {
e.printStackTrace();
System.out.println("failed to close file");
return false;
}
}
return false;
}
}
实际用例结果
源自维基百科specification词条节选
"Specification" redirects here. For other uses, see Specification (disambiguation).
There are different types of specifications, which generally are mostly types of documents, forms or orders or relates to information in databases. The word specification is defined as "to state explicitly or in detail" or "to be specific". A specification may refer to a type of technical standard (the main topic of this page). Using a word "specification" without additional information to what kind of specification you refer to is confusing and considered bad practice within systems engineering. A requirement specification is a set of documented requirements to be satisfied by a material, design, product, or service.[1] A functional specification is closely related to the requirement specification and may show functional block diagrams. A design or product specification describes the features of the solutions for the Requirement Specification, referring to the designed solution or final produced solution. Sometimes the term specification is here used in connection with a data sheet (or spec sheet). This may be confusing. A data sheet describes the technical characteristics of an item or product as designed and/or produced. It can be published by a manufacturer to help people choose products or to help use the products. A data sheet is not a technical specification as described in this article. A "in-service" or "maintained as" specification, specifies the conditions of a system or object after years of operation, including the effects of wear and maintenance (configuration changes). Specifications may also refer to technical standards, which may be developed by any of various kinds of organizations, both public and private. Example organization types include a corporation, a consortium (a small group of corporations), a trade association (an industry-wide group of corporations), a national government (including its military, regulatory agencies, and national laboratories and institutes), a professional association (society), a purpose-made standards organization such as ISO, or vendor-neutral developed generic requirements. It is common for one organization to refer to (reference, call out, cite) the standards of another. Voluntary standards may become mandatory if adopted by a government or business contract.
结果截取
a: 16
of: 16
or: 15
to: 14
the: 12
specification: 11
is: 7
A: 7
and: 7
may: 6
in: 5
.: 5
as: 5
be: 5
standards: 4
by: 4
technical: 4
sheet: 4
refer: 4
product: 3
Specification: 3
organization: 3
data: 3
types: 3
an: 2
word: 2
functional: 2
association: 2
It: 2
government: 2
are: 2
national: 2
describes: 2
help: 2
information: 2
developed: 2
group: 2
which: 2
including: 2
this: 2
corporations: 2
for: 2
design: 2
designed: 2
requirement: 2
mostly: 1
practice: 1
bad: 1
products.: 1
considered: 1
type: 1
without: 1
years: 1
professional: 1
可进行的拓展
建立重载函数,添加参数,可以根据用户需要倒序排列统计结果,即按照单词出现次数从少到多排序。
工程源码包地址:https://coding.net/u/jx8zjs/p/wordCount/git
git@git.coding.net:jx8zjs/wordCount.git
英文词频统计的java实现方法的更多相关文章
- 词频统计的java实现方法——第一次改进
需求概要 原需求 1.读取文件,文件内包可含英文字符,及常见标点,空格级换行符. 2.统计英文单词在本文件的出现次数 3.将统计结果排序 4.显示排序结果 新需求: 1.小文件输入. 为表明程序能跑 ...
- 效能分析——词频统计的java实现方法的第一次改进
java效能分析可以使用JProfiler 词频统计处理的文件为WarAndPeace,大小3282KB约3.3MB,输出结果到文件 在程序本身内开始和结束分别加入时间戳,差值平均为480-490ms ...
- Python——字符串、文件操作,英文词频统计预处理
一.字符串操作: 解析身份证号:生日.性别.出生地等. 凯撒密码编码与解码 网址观察与批量生成 2.凯撒密码编码与解码 凯撒加密法的替换方法是通过排列明文和密文字母表,密文字母表示通过将明文字母表向左 ...
- 组合数据类型,英文词频统计 python
练习: 总结列表,元组,字典,集合的联系与区别.列表,元组,字典,集合的遍历. 区别: 一.列表:列表给大家的印象是索引,有了索引就是有序,想要存储有序的项目,用列表是再好不过的选择了.在python ...
- Hadoop的改进实验(中文分词词频统计及英文词频统计)(4/4)
声明: 1)本文由我bitpeach原创撰写,转载时请注明出处,侵权必究. 2)本小实验工作环境为Windows系统下的百度云(联网),和Ubuntu系统的hadoop1-2-1(自己提前配好).如不 ...
- 1.字符串操作:& 2.英文词频统计预处理
1.字符串操作: 解析身份证号:生日.性别.出生地等. ID = input('请输入十八位身份证号码: ') if len(ID) == 18: print("你的身份证号码是 " ...
- Programming | 中/ 英文词频统计(MATLAB实现)
一.英文词频统计 英文词频统计很简单,只需借助split断句,再统计即可. 完整MATLAB代码: function wordcount %思路:中文词频统计涉及到对"词语"的判断 ...
- python字符串操作、文件操作,英文词频统计预处理
1.字符串操作: 解析身份证号:生日.性别.出生地等. 凯撒密码编码与解码 网址观察与批量生成 解析身份证号:生日.性别.出生地等 def function3(): print('请输入身份证号') ...
- python复合数据类型以及英文词频统计
这个作业的要求来自于:https://edu.cnblogs.com/campus/gzcc/GZCC-16SE1/homework/2753. 1.列表,元组,字典,集合分别如何增删改查及遍历. 列 ...
随机推荐
- msfconsole 无法启动,解决办法
今天突然碰上kali msfconsole 无法启动,经过查找资料,现已成功解决该问题,现将解决办法整理如下: service postgresql start # 启动数据库服务 msfdb ini ...
- BZOJ1084_最大子矩阵_KEY
题目传送门 DP. 但要分类讨论,对于M=1和M=2的情况分别讨论. 1>M=1 设f[i][j]表示选了i个矩阵,到第j位.N^3转移.(前缀和) 2>M=2 设f[i][j][k]表示 ...
- 一维码Code 39简介及其解码实现(zxing-cpp)
一维码Code 39:由于编制简单.能够对任意长度的数据进行编码.支持设备广泛等特性而被广泛采用. Code 39码特点: 1. 能够对任意长度的数据进行编码,其局限在于印刷品的长度和条码阅读器的识别 ...
- centos 中sshd莫名其妙不见了?
发现问题 遇到问题:首先莫要慌:事出有因:先检查一波: 首先呢,看一下/var/log/yum.log 是否有误删的记录: 如有被误删的操作的话:可以去看看日志:到底咋回事: 然后么 yum ins ...
- android prgoressBar setProgressDrawable 在4.0系统式正常,在2.3系统上不能正常使用的问题
上次在做一个电池电量的进度显示时,需要根据背景主题色来切换电池电量的进度的颜色, 但是在对prgoressBar的setProgressDrawable进行设置之后发现,在4.0系统上能够正常,而在2 ...
- UWP DEP0700: 应用程序注册失败。[0x80073CF9] 另一个用户已安装此应用的未打包版本。当前用户无法将该版本替换为打包版本。
最近电脑抽风,我在[应用程序和功能]中重置了以下我的App自然灾害,居然,搞出大新闻了. 它居然从列表中消失了... vs再次编译代码的时候,提示 严重性 代码 说明 项目 文件 行 禁止显示状态 错 ...
- 聊聊Http协议
http协议是大家在互联网中最为熟悉的协议,只要上网大家都会遇到,但是,很多人被问道什么是http协议,http协议的内容是什么就懵了.这里,我们随便聊聊http协议. 首先,我们说说协议.我一直觉得 ...
- linux多项目分别使用不同jdk版本(tomcat版)
此操作只针对tomcat 背景:linux服务器普通用户默认版本为jdk6,jboss项目使用jdk6版本 ,但是tomcat需要使用jdk7.当然也可以分开使用不同账户来启用这两个项目,下面主要介绍 ...
- TPO-19 C2 Cafeteria's Food Policy
TPO-19 C2 Cafeteria's Food Policy 第 1 段 1.Listen to a conversation between a student and the directo ...
- Spark聚合操作:combineByKey()
Spark中对键值对RDD(pairRDD)基于键的聚合函数中,都是通过combineByKey()实现的. 它可以让用户返回与输入数据类型不同的返回值(可以自己配置返回的参数,返回的类型) 首先理解 ...