java原生类型

除char类型以外,所有的原生类型都有对应的Writable类,并且通过get和set方法可以他们的值。
IntWritable和LongWritable还有对应的变长VIntWritable和VLongWritable类。
固定长度还是变长的选用类似与数据库中的char或者vchar。

Text类型

Text类型使用变长int型存储长度,所以Text类型的最大存储为2G.
Text类型采用标准的utf-8编码,所以与其他文本工具可以非常好的交互,但要注意的是,这样的话就和java的String类型差别就很多了。

检索的不同

Text的chatAt返回的是一个整型,及utf-8编码后的数字,而不是象String那样的unicode编码的char类型。
  1. @Test
  2. public void testTextIndex(){
  3. Text text=new Text("hadoop");
  4. Assert.assertEquals(text.getLength(), 6);
  5. Assert.assertEquals(text.getBytes().length, 6);
  6. Assert.assertEquals(text.charAt(2),(int)'d');
  7. Assert.assertEquals("Out of bounds",text.charAt(100),-1);
  8. }

Text还有个find方法,类似String里indexOf方法

  1. @Test
  2. public void testTextFind() {
  3. Text text = new Text("hadoop");
  4. Assert.assertEquals("find a substring",text.find("do"),2);
  5. Assert.assertEquals("Find first 'o'",text.find("o"),3);
  6. Assert.assertEquals("Find 'o' from position 4 or later",text.find("o",4),4);
  7. Assert.assertEquals("No match",text.find("pig"),-1);
  8. }

Unicode的不同

当uft-8编码后的字节大于两个时,Text和String的区别就会更清晰,因为String是按照unicode的char计算,而Text是按照字节计算。
我们来看下1到4个字节的不同的unicode字符
 
4个unicode分别占用1到4个字节,u+10400在java的unicode字符重占用两个char,前三个字符分别占用1个char
我们通过代码来看下String和Text的不同
  1. @Test
  2. public void string() throws UnsupportedEncodingException {
  3. String str = "\u0041\u00DF\u6771\uD801\uDC00";
  4. Assert.assertEquals(str.length(), 5);
  5. Assert.assertEquals(str.getBytes("UTF-8").length, 10);
  6. Assert.assertEquals(str.indexOf("\u0041"), 0);
  7. Assert.assertEquals(str.indexOf("\u00DF"), 1);
  8. Assert.assertEquals(str.indexOf("\u6771"), 2);
  9. Assert.assertEquals(str.indexOf("\uD801\uDC00"), 3);
  10. Assert.assertEquals(str.charAt(0), '\u0041');
  11. Assert.assertEquals(str.charAt(1), '\u00DF');
  12. Assert.assertEquals(str.charAt(2), '\u6771');
  13. Assert.assertEquals(str.charAt(3), '\uD801');
  14. Assert.assertEquals(str.charAt(4), '\uDC00');
  15. Assert.assertEquals(str.codePointAt(0), 0x0041);
  16. Assert.assertEquals(str.codePointAt(1), 0x00DF);
  17. Assert.assertEquals(str.codePointAt(2), 0x6771);
  18. Assert.assertEquals(str.codePointAt(3), 0x10400);
  19. }
  20. @Test
  21. public void text() {
  22. Text text = new Text("\u0041\u00DF\u6771\uD801\uDC00");
  23. Assert.assertEquals(text.getLength(), 10);
  24. Assert.assertEquals(text.find("\u0041"), 0);
  25. Assert.assertEquals(text.find("\u00DF"), 1);
  26. Assert.assertEquals(text.find("\u6771"), 3);
  27. Assert.assertEquals(text.find("\uD801\uDC00"), 6);
  28. Assert.assertEquals(text.charAt(0), 0x0041);
  29. Assert.assertEquals(text.charAt(1), 0x00DF);
  30. Assert.assertEquals(text.charAt(3), 0x6771);
  31. Assert.assertEquals(text.charAt(6), 0x10400);
  32. }

这样一比较就很明显了。

1.String的length()方法返回的是char的数量,Text的getLength()方法返回的是字节的数量。
2.String的indexOf()方法返回的是以char为单元的偏移量,Text的find()方法返回的是以字节为单位的偏移量。
3.String的charAt()方法不是返回的整个unicode字符,而是返回的是java中的char字符
4.String的codePointAt()和Text的charAt方法比较类似,不过要注意,前者是按char的偏移量,后者是字节的偏移量

Text的迭代

在Text中对unicode字符的迭代是相当复杂的,因为与unicode所占的字节数有关,不能简单的使用index的增长来确定。首先要把Text对象使用ByteBuffer进行封装,然后再调用Text的静态方法bytesToCodePoint对ByteBuffer进行轮询返回unicode字符的code point。看一下示例代码:
  1. package com.sweetop.styhadoop;
  2. import org.apache.hadoop.io.Text;
  3. import java.nio.ByteBuffer;
  4. /**
  5. * Created with IntelliJ IDEA.
  6. * User: lastsweetop
  7. * Date: 13-7-9
  8. * Time: 下午5:00
  9. * To change this template use File | Settings | File Templates.
  10. */
  11. public class TextIterator {
  12. public static void main(String[] args) {
  13. Text text = new Text("\u0041\u00DF\u6771\uD801\udc00");
  14. ByteBuffer buffer = ByteBuffer.wrap(text.getBytes(), 0, text.getLength());
  15. int cp;
  16. while (buffer.hasRemaining() && (cp = Text.bytesToCodePoint(buffer)) != -1) {
  17. System.out.println(Integer.toHexString(cp));
  18. }
  19. }
  20. }

Text的修改

除了NullWritable是不可更改外,其他类型的Writable都是可以修改的。你可以通过Text的set方法去修改去修改重用这个实例。
  1. @Test
  2. public void testTextMutability() {
  3. Text text = new Text("hadoop");
  4. text.set("pig");
  5. Assert.assertEquals(text.getLength(), 3);
  6. Assert.assertEquals(text.getBytes().length, 3);
  7. }

但要注意的就是,在某些情况下Text的getBytes方法返回的字节数组的长度和Text的getLength方法返回的长度不一致。因此,在调用getBytes()方法的同时最好也调用一下getLength方法,这样你就知道在字节数组里有多少有效的字符。

  1. @Test
  2. public void testTextMutability2() {
  3. Text text = new Text("hadoop");
  4. text.set(new Text("pig"));
  5. Assert.assertEquals(text.getLength(),3);
  6. Assert.assertEquals(text.getBytes().length,6);
  7. }

BytesWritable类型

ByteWritable类型是一个二进制数组的封装类型,序列化格式是以一个4字节的整数(这点与Text不同,Text是以变长int开头)开始表明字节数组的长度,然后接下来就是数组本身。看下示例:
  1. @Test
  2. public void testByteWritableSerilizedFromat() throws IOException {
  3. BytesWritable bytesWritable=new BytesWritable(new byte[]{3,5});
  4. byte[] bytes=SerializeUtils.serialize(bytesWritable);
  5. Assert.assertEquals(StringUtils.byteToHexString(bytes),"000000020305");
  6. }

和Text一样,ByteWritable也可以通过set方法修改,getLength返回的大小是真实大小,而getBytes返回的大小确不是。

  1. <span style="white-space:pre">  </span>bytesWritable.setCapacity(11);
  2. bytesWritable.setSize(4);
  3. Assert.assertEquals(4,bytesWritable.getLength());
  4. Assert.assertEquals(11,bytesWritable.getBytes().length);

NullWritable类型

NullWritable是一个非常特殊的Writable类型,序列化不包含任何字符,仅仅相当于个占位符。你在使用mapreduce时,key或者value在无需使用时,可以定义为NullWritable。
  1. package com.sweetop.styhadoop;
  2. import org.apache.hadoop.io.NullWritable;
  3. import org.apache.hadoop.util.StringUtils;
  4. import java.io.IOException;
  5. /**
  6. * Created with IntelliJ IDEA.
  7. * User: lastsweetop
  8. * Date: 13-7-16
  9. * Time: 下午9:23
  10. * To change this template use File | Settings | File Templates.
  11. */
  12. public class TestNullWritable {
  13. public static void main(String[] args) throws IOException {
  14. NullWritable nullWritable=NullWritable.get();
  15. System.out.println(StringUtils.byteToHexString(SerializeUtils.serialize(nullWritable)));
  16. }
  17. }

ObjectWritable类型

ObjectWritable是其他类型的封装类,包括java原生类型,String,enum,Writable,null等,或者这些类型构成的数组。当你的一个field有多种类型时,ObjectWritable类型的用处就发挥出来了,不过有个不好的地方就是占用的空间太大,即使你存一个字母,因为它需要保存封装前的类型,我们来看瞎示例:

  1. package com.sweetop.styhadoop;
  2. import org.apache.hadoop.io.ObjectWritable;
  3. import org.apache.hadoop.io.Text;
  4. import org.apache.hadoop.util.StringUtils;
  5. import java.io.IOException;
  6. /**
  7. * Created with IntelliJ IDEA.
  8. * User: lastsweetop
  9. * Date: 13-7-17
  10. * Time: 上午9:14
  11. * To change this template use File | Settings | File Templates.
  12. */
  13. public class TestObjectWritable {
  14. public static void main(String[] args) throws IOException {
  15. Text text=new Text("\u0041");
  16. ObjectWritable objectWritable=new ObjectWritable(text);
  17. System.out.println(StringUtils.byteToHexString(SerializeUtils.serialize(objectWritable)));
  18. }
  19. }

仅仅是保存一个字母,那么看下它序列化后的结果是什么:

  1. 00196f72672e6170616368652e6861646f6f702e696f2e5465787400196f72672e6170616368652e6861646f6f702e696f2e546578740141

太浪费空间了,而且类型一般是已知的,也就那么几个,那么它的代替方法出现,看下一小节

GenericWritable类型

使用GenericWritable时,只需继承于他,并通过重写getTypes方法指定哪些类型需要支持即可,我们看下用法:
  1. package com.sweetop.styhadoop;
  2. import org.apache.hadoop.io.GenericWritable;
  3. import org.apache.hadoop.io.Text;
  4. import org.apache.hadoop.io.Writable;
  5. class MyWritable extends GenericWritable {
  6. MyWritable(Writable writable) {
  7. set(writable);
  8. }
  9. public static Class<? extends Writable>[] CLASSES=null;
  10. static {
  11. CLASSES=  (Class<? extends Writable>[])new Class[]{
  12. Text.class
  13. };
  14. }
  15. @Override
  16. protected Class<? extends Writable>[] getTypes() {
  17. return CLASSES;  //To change body of implemented methods use File | Settings | File Templates.
  18. }
  19. }

然后输出序列化后的结果

  1. package com.sweetop.styhadoop;
  2. import org.apache.hadoop.io.IntWritable;
  3. import org.apache.hadoop.io.Text;
  4. import org.apache.hadoop.io.VIntWritable;
  5. import org.apache.hadoop.util.StringUtils;
  6. import java.io.IOException;
  7. /**
  8. * Created with IntelliJ IDEA.
  9. * User: lastsweetop
  10. * Date: 13-7-17
  11. * Time: 上午9:51
  12. * To change this template use File | Settings | File Templates.
  13. */
  14. public class TestGenericWritable {
  15. public static void main(String[] args) throws IOException {
  16. Text text=new Text("\u0041\u0071");
  17. MyWritable myWritable=new MyWritable(text);
  18. System.out.println(StringUtils.byteToHexString(SerializeUtils.serialize(text)));
  19. System.out.println(StringUtils.byteToHexString(SerializeUtils.serialize(myWritable)));
  20. }
  21. }

结果是:

  1. 024171
  2. 00024171

GenericWritable的序列化只是把类型在type数组里的索引放在了前面,这样就比ObjectWritable节省了很多空间,所以推荐大家使用GenericWritable

集合类型的Writable

ArrayWritable和TwoDArrayWritable

ArrayWritable和TwoDArrayWritable分别表示数组和二维数组的Writable类型,指定数组的类型有两种方法,构造方法里设置,或者继承于ArrayWritable,TwoDArrayWritable也是一样。
  1. package com.sweetop.styhadoop;
  2. import org.apache.hadoop.io.ArrayWritable;
  3. import org.apache.hadoop.io.Text;
  4. import org.apache.hadoop.io.Writable;
  5. import org.apache.hadoop.util.StringUtils;
  6. import java.io.IOException;
  7. /**
  8. * Created with IntelliJ IDEA.
  9. * User: lastsweetop
  10. * Date: 13-7-17
  11. * Time: 上午11:14
  12. * To change this template use File | Settings | File Templates.
  13. */
  14. public class TestArrayWritable {
  15. public static void main(String[] args) throws IOException {
  16. ArrayWritable arrayWritable=new ArrayWritable(Text.class);
  17. arrayWritable.set(new Writable[]{new Text("\u0071"),new Text("\u0041")});
  18. System.out.println(StringUtils.byteToHexString(SerializeUtils.serialize(arrayWritable)));
  19. }
  20. }

看下输出:

  1. 0000000201710141

可知,ArrayWritable以一个整型开始表示数组长度,然后数组里的元素一一排开。

ArrayPrimitiveWritable和上面类似,只是不需要用子类去继承ArrayWritable而已。

MapWritable和SortedMapWritable

MapWritable对应Map,SortedMapWritable对应SortedMap,以4个字节开头,存储集合大小,然后每个元素以一个字节开头存储类型的索引(类似GenericWritable,所以总共的类型总数只能倒127),接着是元素本身,先key后value,这样一对对排开。
这两个Writable以后会用很多,贯穿整个hadoop,这里就不写示例了。
 
我们注意到没看到set集合和list集合,这个可以代替实现。用MapWritable代替set,SortedMapWritable代替sortedmap,只需将他们的values设置成NullWritable即可,NullWritable不占空间。相同类型构成的list,可以用ArrayWritable代替,不同类型的list可以用GenericWritable实现类型,然后再使用ArrayWritable封装。当然MapWritable一样可以实现list,把key设置为索引,values做list里的元素。

各种类型的Writable(Text、ByteWritable、NullWritable、ObjectWritable、GenericWritable、ArrayWritable、MapWritable、SortedMapWritable)转的更多相关文章

  1. Hadoop Serialization -- hadoop序列化具体解释 (2)【Text,BytesWritable,NullWritable】

    回想: 回想序列化,事实上原书的结构非常清晰,我截图给出书中的章节结构: 序列化最基本的,最底层的是实现writable接口,wiritable规定读和写的游戏规则 (void write(DataO ...

  2. Hadoop Serialization -- hadoop序列化详解 (2)【Text,BytesWritable,NullWritable】

    回顾: 回顾序列化,其实原书的结构很清晰,我截图给出书中的章节结构: 序列化最主要的,最底层的是实现writable接口,wiritable规定读和写的游戏规则 (void write(DataOut ...

  3. mysql列类型char,varchar,text,tinytext,mediumtext,longtext的比较与选择

    储存不区分大小写的字符数据 TINYTEXT 最大长度是 255 (2^8 – 1) 个字符. TEXT 最大长度是 65535 (2^16 – 1) 个字符. MEDIUMTEXT 最大长度是 16 ...

  4. 叼叼叼,HTML5日期(Date)类型和文本(Text)类型互相转换

    <input placeholder="From" class="form-control" type="text" onfocus= ...

  5. 【原创】大叔问题定位分享(12)Spark保存文本类型文件(text、csv、json等)到hdfs时为什么是压缩格式的

    问题重现 rdd.repartition(1).write.csv(outPath) 写文件之后发现文件是压缩过的 write时首先会获取hadoopConf,然后从中获取是否压缩以及压缩格式 org ...

  6. hadoop自带的writable类型

    Hadoop 中,并没有使用Java自带的基本类型类(Integer.Float等),而是使用自己开发的类.Hadoop 自带有很多序列化类型,大致分为以下两种: 实现了WritableCompara ...

  7. Hadoop Serialization -- hadoop序列化详解 (3)【ObjectWritable,集合Writable以及自定义的Writable】

    前瞻:本文介绍ObjectWritable,集合Writable以及自定义的Writable TextPair 回顾: 前面了解到hadoop本身支持java的基本类型的序列化,并且提供相应的包装实现 ...

  8. Mysql 中 text类型和 blog类型的异同

    MySQL存在text和blob: (1)相同 在TEXT或BLOB列的存储或检索过程中,不存在大小写转换,当未运行在严格模式时,如果你为BLOB或TEXT列分配一个超过该列类型的最大长度的值值,值被 ...

  9. mysql的text的类型注意

    不要以为text就只有一种类型! Text也分为四种类型:TINYTEXT.TEXT.MEDIUMTEXT和LONGTEXT 其中 TINYTEXT 256 bytes TEXT 65,535 byt ...

随机推荐

  1. HTTP 协议的历史演变和设计思路

    HTTP 协议是互联网的基础协议,也是网页开发的必备知识,最新版本 HTTP/2 更是让它成为技术热点. 本文介绍 HTTP 协议的历史演变和设计思路. 一.HTTP/0.9 HTTP 是基于 TCP ...

  2. sqlserver总结-视图及存储过程

    视图中不能声明变量,不能调用存储过程,如果写比较复杂的查询,需要应用存储过程 视图也可以和函数结合 存储过程通过select或其他语句返回结果集 除此之外,存储过程返回结果只有两种方式 1 retur ...

  3. LeetCode Count Complete Tree Nodes

    原题链接在这里:https://leetcode.com/problems/count-complete-tree-nodes/ Given a complete binary tree, count ...

  4. POJ 1039问题描述

    Description The GX Light Pipeline Company started to prepare bent pipes for the new transgalactic li ...

  5. ionic环境搭建和安装

    1. 安装node环境 nodeJs环境的安装很简单,去官网下载最新版的NodeJs直接安装即可. Node官网: https://nodejs.org/ 安装完成后配置环境变量,计算机->属性 ...

  6. uwsgi + nigix + django的样式展示

    编辑添加黄色部分  是你的项目目录  在你的项目目录写的静态文件 内的样式调用的是static 如果不是 请改名 [root@ayibang-server s10day11]# vim /etc/ng ...

  7. java类的加载、链接、初始化

    JVM和类的关系 当我们调用JAVA命令运行某个java程序时,该命令将会启动一条java虚拟机进程,不管该java程序有多么复杂,该程序启动了多少个线程,它们都处于该java虚拟机进程里.正如前面介 ...

  8. The L1 Median (Weber 1909)

    The L1 Median (Weber 1909) 链接网址 Derived from a transportation cost minimization problem, the L1 medi ...

  9. jose4j / JWT Examples

    jose4j / JWT Examples View History JSON Web Token (JWT) Code Examples Producing and consuming a sign ...

  10. SQL exist

    EXISTS = IN,意思相同不过语法上有点点区别,好像使用IN效率要差点,应该是不会执行索引的原因SELECT ID,NAME FROM A WHERE ID IN (SELECT AID FRO ...