【Boost】boost::tokenizer详解
tokenizer 库提供预定义好的四个分词对象, 其中char_delimiters_separator已弃用. 其他如下:
1. char_separator
char_separator有两个构造函数
1. char_separator()
使用函数 std::isspace() 来识别被弃分隔符,同时使用 std::ispunct() 来识别保留分隔符。另外,抛弃空白单词。(见例2)
2. char_separator(// 不保留的分隔符
const Char* dropped_delims,
// 保留的分隔符
const Char* kept_delims = 0,
// 默认不保留空格分隔符, 反之加上改参数keep_empty_tokens
empty_token_policy empty_tokens = drop_empty_tokens)
该函数创建一个 char_separator 对象,该对象被用于创建一个 token_iterator 或 tokenizer 以执行单词分解。dropped_delims 和 kept_delims 都是字符串,其中的每个字符被用作分解时的分隔符。当在输入序列中遇到一个分隔符时,当前单词即完成,并开始下一个新单词。dropped_delims 中的分隔符不出现在输出的单词中,而 kept_delims 中的分隔符则会作为单词输出。如果 empty_tokens 为 drop_empty_tokens, 则空白单词不会出现在输出中。如果 empty_tokens 为 keep_empty_tokens 则空白单词将出现在输出中。 (见例3)
2. escaped_list_separator
escaped_list_separator有两个构造函数
下面三个字符做为分隔符: '\', ',', '"'
1. explicit escaped_list_separator(Char e = '\\', Char c = ',',Char q = '\"');
| 参数 | 描述 |
| e | 指定用作转义的字符。缺省使用C风格的\(反斜杠)。但是你可以传入不同的字符来覆盖它。 如果你有很多字段是Windows风格的文件名时,路径中的每个\都要转义。 你可以使用其它字符作为转义字符。 |
| c | 指定用作字段分隔的字符 |
| q | 指定用作引号的字符 |
2. escaped_list_separator(string_type e, string_type c, string_type q):
| 参数 | 描述 |
| e | 字符串e中的字符都被视为转义字符。如果给定的是空字符串,则没有转义字符。 |
| c | 字符串c中的字符都被视为分隔符。如果给定的是空字符串,则没有分隔符。 |
| q | 字符串q中的字符都被视为引号字符。如果给定的是空字符串,则没有引号字符。 |
3. offset_separator
offset_separator 有一个有用的构造函数
template<typename Iter>
offset_separator(Iter begin,Iter end,bool bwrapoffsets = true, bool breturnpartiallast = true);
| 参数 | 描述 |
| begin, end | 指定整数偏移量序列 |
| bwrapoffsets | 指明当所有偏移量用完后是否回绕到偏移量序列的开头继续。 例如字符串 "1225200101012002" 用偏移量 (2,2,4) 分解, 如果 bwrapoffsets 为 true, 则分解为 12 25 2001 01 01 2002. 如果 bwrapoffsets 为 false, 则分解为 12 25 2001,然后就由于偏移量用完而结束。 |
| breturnpartiallast | 指明当被分解序列在生成当前偏移量所需的字符数之前结束,是否创建一个单词,或是忽略它。 例如字符串 "122501" 用偏移量 (2,2,4) 分解, 如果 breturnpartiallast 为 true,则分解为 12 25 01. 如果为 false, 则分解为 12 25,然后就由于序列中只剩下2个字符不足4个而结束。 |
例子
- void test_string_tokenizer()
- {
- using namespace boost;
- // 1. 使用缺省模板参数创建分词对象, 默认把所有的空格和标点作为分隔符.
- {
- std::string str("Link raise the master-sword.");
- tokenizer<> tok(str);
- for (BOOST_AUTO(pos, tok.begin()); pos != tok.end(); ++pos)
- std::cout << "[" << *pos << "]";
- std::cout << std::endl;
- // [Link][raise][the][master][sword]
- }
- // 2. char_separator()
- {
- std::string str("Link raise the master-sword.");
- // 一个char_separator对象, 默认构造函数(保留标点但将它看作分隔符)
- char_separator<char> sep;
- tokenizer<char_separator<char> > tok(str, sep);
- for (BOOST_AUTO(pos, tok.begin()); pos != tok.end(); ++pos)
- std::cout << "[" << *pos << "]";
- std::cout << std::endl;
- // [Link][raise][the][master][-][sword][.]
- }
- // 3. char_separator(const Char* dropped_delims,
- // const Char* kept_delims = 0,
- // empty_token_policy empty_tokens = drop_empty_tokens)
- {
- std::string str = ";!!;Hello|world||-foo--bar;yow;baz|";
- char_separator<char> sep1("-;|");
- tokenizer<char_separator<char> > tok1(str, sep1);
- for (BOOST_AUTO(pos, tok1.begin()); pos != tok1.end(); ++pos)
- std::cout << "[" << *pos << "]";
- std::cout << std::endl;
- // [!!][Hello][world][foo][bar][yow][baz]
- char_separator<char> sep2("-;", "|", keep_empty_tokens);
- tokenizer<char_separator<char> > tok2(str, sep2);
- for (BOOST_AUTO(pos, tok2.begin()); pos != tok2.end(); ++pos)
- std::cout << "[" << *pos << "]";
- std::cout << std::endl;
- // [][!!][Hello][|][world][|][][|][][foo][][bar][yow][baz][|][]
- }
- // 4. escaped_list_separator
- {
- std::string str = "Field 1,\"putting quotes around fields, allows commas\",Field 3";
- tokenizer<escaped_list_separator<char> > tok(str);
- for (BOOST_AUTO(pos, tok.begin()); pos != tok.end(); ++pos)
- std::cout << "[" << *pos << "]";
- std::cout << std::endl;
- // [Field 1][putting quotes around fields, allows commas][Field 3]
- // 引号内的逗号不可做为分隔符.
- }
- // 5. offset_separator
- {
- std::string str = "12252001400";
- int offsets[] = {2, 2, 4};
- offset_separator f(offsets, offsets + 3);
- tokenizer<offset_separator> tok(str, f);
- for (BOOST_AUTO(pos, tok.begin()); pos != tok.end(); ++pos)
- std::cout << "[" << *pos << "]";
- std::cout << std::endl;
- }
- }
【Boost】boost::tokenizer详解的更多相关文章
- boost::tokenizer详解
tokenizer 库提供预定义好的四个分词对象, 其中char_delimiters_separator已弃用. 其他如下: 1. char_separator char_separator有两个构 ...
- [转] boost::function用法详解
http://blog.csdn.net/benny5609/article/details/2324474 要开始使用 Boost.Function, 就要包含头文件 "boost/fun ...
- boost::function用法详解
要开始使用 Boost.Function, 就要包含头文件 "boost/function.hpp", 或者某个带数字的版本,从 "boost/function/func ...
- Boost::split用法详解
工程中使用boost库:(设定vs2010环境)在Library files加上 D:\boost\boost_1_46_0\bin\vc10\lib在Include files加上 D:\boost ...
- boost::fucntion 用法详解
转载自:http://blog.csdn.net/benny5609/article/details/2324474 要开始使用 Boost.Function, 就要包含头文件 "boost ...
- boost库asio详解1——strand与io_service区别
namespace { // strand提供串行执行, 能够保证线程安全, 同时被post或dispatch的方法, 不会被并发的执行. // io_service不能保证线程安全 boost::a ...
- Boost::bind使用详解
1.Boost::bind 在STL中,我们经常需要使用bind1st,bind2st函数绑定器和fun_ptr,mem_fun等函数适配器,这些函数绑定器和函数适配器使用起来比较麻烦,需要根据是全局 ...
- 【Boost】boost库asio详解5——resolver与endpoint使用说明
tcp::resolver一般和tcp::resolver::query结合用,通过query这个词顾名思义就知道它是用来查询socket的相应信息,一般而言我们关心socket的东东有address ...
- boost::algorithm用法详解之字符串关系判断
http://blog.csdn.net/qingzai_/article/details/44417937 下面先列举几个常用的: #define i_end_with boost::iends_w ...
随机推荐
- linux下统计文本行数的各种方法
方法一:awk awk '{print NR}' test1.txt | tail -n1
- System.Collections里的一些接口
System.Collections 名称空间中的几个接口提供了基本的组合功能: IEnumerable 可以迭代集合中的项. ICollection(继承于IEnumerable)可以获取集合中 ...
- HTML汇总以及CSS的一些开端
一.HTML初探 1.HTML(HyperText Markup Language):超文本标记语言指的就是超跃了txt文件,能够在里面进行一些例如:就是指页面内可以包含图片.链接 .甚至音乐.程序等 ...
- https笔记【转】
图解HTTPS 我们都知道HTTPS能够加密信息,以免敏感信息被第三方获取.所以很多银行网站或电子邮箱等等安全级别较高的服务都会采用HTTPS协议. HTTPS简介 HTTPS其实是有两部分组成:HT ...
- IntelliJ IDEA 2017 配置Tomcat 运行Web项目
以前都用MyEclipse写程序的 突然用了IDEA各种不习惯的说 借鉴了很多网上好的配置办法,感谢各位大神~ 前期准备 IDEA.JDK.Tomcat请先在自己电脑上装好 好么~ 博客图片为主 请多 ...
- ifconfig: command not found(CentOS 7,其他的可以参考)
ifconfig: command not found 查看path配置(echo相当于c中的printf,C#中的Console.WriteLine) 1 echo $PATH 解决方案1:先看看是 ...
- linux磁盘空间占满问题快速定位并解决
经常会遇到这样的场景:测试环境磁盘跑满了,导致系统不能正常运行!此时就需要查看是哪个目录或者文件占用了空间.常使用如下几个命令进行排查:df, lsof,du. 通常的解决步骤如下:1. df -h ...
- 404.17 - 动态内容通过通配符 MIME 映射映射到静态文件处理程序
刚刚重装了系统,原有的ASP.NET工程下面的WebService无法运行,如下: 404.17 - 动态内容通过通配符 MIME 映射映射到静态文件处理程序 微软的提示,是做三项更改,但是我改了之后 ...
- docker之搭建私有镜像仓库和公有仓库
一.搭建私有仓库 1.docker pull registry #下载registry镜像并启动 2. docker run -d -v /opt/registry:/var/lib/registry ...
- dll和lib的关系(转)
转自http://blog.163.com/zhengjiu_520/blog/static/3559830620093583438464/ 前面有一章说编译与链接的,说得很简略,其实应该放到这一章一 ...