1. char_separator

char_separator有两个构造函数
1. char_separator()
使用函数 std::isspace() 来识别被弃分隔符，同时使用 std::ispunct() 来识别保留分隔符。另外，抛弃空白单词。(见例2)
2. char_separator(// 不保留的分隔符
                               const Char* dropped_delims,
                               // 保留的分隔符
                               const Char* kept_delims = 0,
                               // 默认不保留空格分隔符, 反之加上改参数keep_empty_tokens
                               empty_token_policy empty_tokens = drop_empty_tokens)
该函数创建一个 char_separator 对象，该对象被用于创建一个 token_iterator 或 tokenizer 以执行单词分解。dropped_delims 和 kept_delims 都是字符串，其中的每个字符被用作分解时的分隔符。当在输入序列中遇到一个分隔符时，当前单词即完成，并开始下一个新单词。dropped_delims 中的分隔符不出现在输出的单词中，而 kept_delims 中的分隔符则会作为单词输出。如果 empty_tokens 为 drop_empty_tokens, 则空白单词不会出现在输出中。如果 empty_tokens 为 keep_empty_tokens 则空白单词将出现在输出中。 (见例3)

2. escaped_list_separator

escaped_list_separator有两个构造函数
下面三个字符做为分隔符: '\', ',', '"'
1. explicit escaped_list_separator(Char e = '\\', Char c = ',',Char q = '\"');

参数	描述
e	指定用作转义的字符。缺省使用C风格的\(反斜杠)。但是你可以传入不同的字符来覆盖它。如果你有很多字段是Windows风格的文件名时，路径中的每个\都要转义。你可以使用其它字符作为转义字符。
c	指定用作字段分隔的字符
q	指定用作引号的字符

2. escaped_list_separator(string_type e, string_type c, string_type q):

参数	描述
e	字符串e中的字符都被视为转义字符。如果给定的是空字符串，则没有转义字符。
c	字符串c中的字符都被视为分隔符。如果给定的是空字符串，则没有分隔符。
q	字符串q中的字符都被视为引号字符。如果给定的是空字符串，则没有引号字符。

3. offset_separator

offset_separator 有一个有用的构造函数
template<typename Iter>
offset_separator(Iter begin,Iter end,bool bwrapoffsets = true, bool breturnpartiallast = true);

参数	描述
begin, end	指定整数偏移量序列
bwrapoffsets	指明当所有偏移量用完后是否回绕到偏移量序列的开头继续。例如字符串 "1225200101012002" 用偏移量 (2,2,4) 分解，如果 bwrapoffsets 为 true, 则分解为 12 25 2001 01 01 2002. 如果 bwrapoffsets 为 false, 则分解为 12 25 2001，然后就由于偏移量用完而结束。
breturnpartiallast	指明当被分解序列在生成当前偏移量所需的字符数之前结束，是否创建一个单词，或是忽略它。例如字符串 "122501" 用偏移量 (2,2,4) 分解，如果 breturnpartiallast 为 true，则分解为 12 25 01. 如果为 false, 则分解为 12 25，然后就由于序列中只剩下2个字符不足4个而结束。

例子

void test_string_tokenizer()
{
using namespace boost;
// 1. 使用缺省模板参数创建分词对象, 默认把所有的空格和标点作为分隔符.
{
std::string str("Link raise the master-sword.");
tokenizer<> tok(str);
for (BOOST_AUTO(pos, tok.begin()); pos != tok.end(); ++pos)
std::cout << "[" << *pos << "]";
std::cout << std::endl;
// [Link][raise][the][master][sword]
}
// 2. char_separator()
{
std::string str("Link raise the master-sword.");
// 一个char_separator对象, 默认构造函数(保留标点但将它看作分隔符)
char_separator<char> sep;
tokenizer<char_separator<char> > tok(str, sep);
for (BOOST_AUTO(pos, tok.begin()); pos != tok.end(); ++pos)
std::cout << "[" << *pos << "]";
std::cout << std::endl;
// [Link][raise][the][master][-][sword][.]
}
// 3. char_separator(const Char* dropped_delims,
// const Char* kept_delims = 0,
// empty_token_policy empty_tokens = drop_empty_tokens)
{
std::string str = ";!!;Hello|world||-foo--bar;yow;baz|";
char_separator<char> sep1("-;|");
tokenizer<char_separator<char> > tok1(str, sep1);
for (BOOST_AUTO(pos, tok1.begin()); pos != tok1.end(); ++pos)
std::cout << "[" << *pos << "]";
std::cout << std::endl;
// [!!][Hello][world][foo][bar][yow][baz]
char_separator<char> sep2("-;", "|", keep_empty_tokens);
tokenizer<char_separator<char> > tok2(str, sep2);
for (BOOST_AUTO(pos, tok2.begin()); pos != tok2.end(); ++pos)
std::cout << "[" << *pos << "]";
std::cout << std::endl;
// [][!!][Hello][|][world][|][][|][][foo][][bar][yow][baz][|][]
}
// 4. escaped_list_separator
{
std::string str = "Field 1,\"putting quotes around fields, allows commas\",Field 3";
tokenizer<escaped_list_separator<char> > tok(str);
for (BOOST_AUTO(pos, tok.begin()); pos != tok.end(); ++pos)
std::cout << "[" << *pos << "]";
std::cout << std::endl;
// [Field 1][putting quotes around fields, allows commas][Field 3]
// 引号内的逗号不可做为分隔符.
}
// 5. offset_separator
{
std::string str = "12252001400";
int offsets[] = {2, 2, 4};
offset_separator f(offsets, offsets + 3);
tokenizer<offset_separator> tok(str, f);
for (BOOST_AUTO(pos, tok.begin()); pos != tok.end(); ++pos)
std::cout << "[" << *pos << "]";
std::cout << std::endl;
}
}

【Boost】boost::tokenizer详解的更多相关文章

boost::tokenizer详解
tokenizer 库提供预定义好的四个分词对象, 其中char_delimiters_separator已弃用. 其他如下: 1. char_separator char_separator有两个构 ...
[转] boost::function用法详解
http://blog.csdn.net/benny5609/article/details/2324474 要开始使用 Boost.Function, 就要包含头文件 "boost/fun ...
boost::function用法详解
要开始使用 Boost.Function, 就要包含头文件 "boost/function.hpp", 或者某个带数字的版本,从 "boost/function/func ...
Boost::split用法详解
工程中使用boost库:(设定vs2010环境)在Library files加上 D:\boost\boost_1_46_0\bin\vc10\lib在Include files加上 D:\boost ...
boost::fucntion 用法详解
转载自:http://blog.csdn.net/benny5609/article/details/2324474 要开始使用 Boost.Function, 就要包含头文件 "boost ...
boost库asio详解1——strand与io_service区别
namespace { // strand提供串行执行, 能够保证线程安全, 同时被post或dispatch的方法, 不会被并发的执行. // io_service不能保证线程安全 boost::a ...
Boost::bind使用详解
1.Boost::bind 在STL中,我们经常需要使用bind1st,bind2st函数绑定器和fun_ptr,mem_fun等函数适配器,这些函数绑定器和函数适配器使用起来比较麻烦,需要根据是全局 ...
【Boost】boost库asio详解5——resolver与endpoint使用说明
tcp::resolver一般和tcp::resolver::query结合用,通过query这个词顾名思义就知道它是用来查询socket的相应信息,一般而言我们关心socket的东东有address ...
boost::algorithm用法详解之字符串关系判断
http://blog.csdn.net/qingzai_/article/details/44417937 下面先列举几个常用的: #define i_end_with boost::iends_w ...

随机推荐

搭建Hadoop集群（生产环境）
1.搭建之前:百度copy一下介绍 (本博客几乎全都是生产环境的配置..包括mongo等hbase其他) Hadoop是一个由Apache基金会所开发的分布式系统基础架构. 用户可以在不了解分布式底层 ...
windows系统下mysql-8.0.13-winx64（zip安装）
一.下载地址: http://mirrors.163.com/mysql/Downloads/MySQL-8.0/mysql-8.0.13-winx64.zip 二.安装: 1.解压: mysql根路 ...
JAVA核心技术I---JAVA基础知识（数据结构基础）
一:数组 (一)基本内容是与C一致的 (二)数组定义和初始化 (1)声明 int a[]; //a没有new操作,没有被分配内存,为null int[] b; //b没有new操作,没有被分配内存,为 ...
log4j 基础教程【转】
参考引用自: http://javacrazyer.iteye.com/blog/1135493 我的git地址: https://git.oschina.net/KingBoBo/Log4JDemo ...
有关mysql的innodb_flush_log_at_trx_commit参数
一.参数解释 0:log buffer将每秒一次地写入log file中,并且log file的flush(刷到磁盘)操作同时进行.该模式下在事务提交的时候,不会主动触发写入磁盘的操作. 1:每次事务 ...
row_number()over()使用
语法: ROW_NUMBER ( ) OVER ( [ PARTITION BY value_expression , ... [ n ] ] order_by_clause ) 通过语法可以看出 o ...
Java9+版本中，Interface的内容
使用接口的注意事项: 1.接口没有静态代码块或者构造方法 2.一个类的父类是唯一的,但是一个类可以同时实现多个接口(区别) 3.如果实现类实现多个接口有重名的抽象方法,那么实现类只需要覆盖重写一个即可 ...
jquery 控制 video 视频播放和暂停
$('video').trigger('play'); $('video').trigger('pause'); 参考:https://blog.csdn.net/arvin0/article/det ...
今天终于想明白为什么java包要倒着写
比如 com.baidu.video,因为java内部实际上是以文件夹形式存在的,是按com,baidu,video依次生成文件夹的具体功能的是子文件夹,所以要倒着写.
Sqlserver直接附加数据库和设置sa密码
1.exec sp_attach_db 'test','E:\db\test.mdf','E:\db\test_log.ldf' 2.sp_password Null,'123','sa' 推荐一个微 ...

【Boost】boost::tokenizer详解

1. char_separator

2. escaped_list_separator

3. offset_separator

例子

【Boost】boost::tokenizer详解的更多相关文章

随机推荐

热门专题