hive的使用02
1.hive的交互方式
1.1 bin/hive 进入hive交互命令行环境
1.2 bin/hive -e 'select * from hive.student;' (可以通过 > 将结果写入到指定的文件中)
1.3 bin/hive -f /opt/data/hive-select.sql (可以通过 > 将结果写入到指定的文件中)
2.hive命令行模式下如何访问hdfs文件系统和本地文件系统
2.1 访问hdfs文件系统 dfs 命令
2.2 访问本地文件系统 !命令
3.表的创建
3.1 创建普通表存储格式为textfile
create table if not exists hive.employee (id int,name string,job string,manager_id int comment '上级领导id',apply_date string comment '入职时间',salary double,
reward double,dept_id int comment '所在部门id') comment '员工表'
row format delimited
fields terminated by '\t'
stored as textfile;
其中:fields terminated by '\t' 表示字段间以tab分割
stored as textfile 表示存储格式为textfile(默认)
3.1.2 创建复杂结构表
create table if not exists cdh_hive.teachers(
id int,
name string,
flag boolean,
score array<int>,
tech map<string,string>,
other struct<phone:string,email:string>
) row format delimited
fields terminated by '\;'
collection items terminated by ','
map keys terminated by ':';
其中:fields terminated by '\;' 表示字段间以分号分割
collection items terminated by ',' 表示集合字段中各个元素间以,分割
map keys terminated by ':'; 表示map字段中键值直接以:分割
以下是测试数据:
1;test1;false;100,85,52;age:29,sex:m;\N,\N
2;test2;true;25,35;\N;021-111111,192165@qq.com
3;test3;true;\N;age:20,sex:F,height:168;\N,3@qq.com
4;test4;true;\N;\N;021-44444444,\N
5;;true;1,2,3;;021-44444444,

3.2 创建外部表(删除表是不删除数据文件)
create external table if not exists hive.dept(id int,name string,address string) comment '部门表'
row format delimited
fields terminated by '\t'
stored as textfile
location '/user/yanglin/data/hive/dept';
其中 external 表示创建的是外部表,默认不写为管理表
location 表示数据所在hdfs上的目录
3.3 创建分区表
create external table if not exists hive.dept_partition(id int,name string,address string)
partitioned by (day string)
row format delimited
fields terminated by '\t';
其中 partitioned by 指定了分区字段,可以指定多个(mouth stirng,day string)
分区表中数据的加载问题:
3.3.1 load data local inpath '/opt/data/dept.txt' into table dept_partition partition (day='20160901');
3.3.2 直接创建目录,把数据放入到指定的目录中,不能查询到数据
需要进行修复:
修复方式一:msck repair table dept_partition;
修复方式二:alter table dept_partition add partition(day=20160903);
3.4 通过like 语句创建表 create table hive.dept_like like hive.dept;
3.5 通过as select 创建表 create table hive.dept_select as select * from dept;
4.加载数据到表中
4.1 通过load语句进行加载数据
load data [local] inpath 'filepath' [overwrite] into table tablename [partition (partcol1=val1,partcol2=val2 ...)]
其中:1、如果数据文件在本地需要local,如果在hdfs上不加local
2、如果要覆盖表中原有的数据需要 overwrite ,如果在原有数据的基础上添加数据不加 overwrite
3、如果是分区表需要制定partition 分区字段
4.2 通过insert into 语句进行加载数据
insert into table hive.dept_into select * from hive.dept;
4.3 创建表时通过location 指定数据所在的目录,例3.2
4.4 通过import语句进行导入
import table hive.employee_like from '/user/yanglin/data/export/employee';
5.数据的导出
5.1 通过insert 方法将结果导入到指定的文件中
insert overwrite local directory '/opt/data/hive_dept_exp' select * from dept;
我们可以对输出的内容进行格式化:
insert overwrite local directory '/opt/data/hive_dept_exp_format'
row format delimited fields terminated by '\t' collection items terminated by '\n'
select * from dept;
其中,不加local时表示将内容放入到hdfs文件系统中。
5.2 使用 > 将查询结果放入指定的文件中
bin/hive -e 'select * from hive.dept;' > /opt/data/hive_dept_exp_e
bin/hive -f /opt/data/hive-select.sql > /opt/data/hive_user_exp_f
5.3 使用sqoop进行数据导入和导出
在接下来的文章中进行讲解。
5.4 使用export进行导出
export table hive.employee to '/user/yanglin/data/export/employee';
注意:这里的路径为hdfs文件系统上的路径
hive的使用02的更多相关文章
- 单节点伪分布集群(weekend110)的Hive子项目启动顺序
因为,我的mysql是用root用户,在/home/hadoop/app/目录下,创建的. 第一步:开启mysql服务 第二步:启动hive [hadoop@weekend110 app]$ su r ...
- 说说单节点集群里安装hive、3\5节点集群里安装hive的诡异区别
这几天,无意之间,被这件事情给迷惑,不解!先暂时贴于此,以后再解决! 详细问题如下: 在hive的安装目录下(我这里是 /home/hadoop/app/hive-1.2.1),hive的安装目录的l ...
- [Spark][Hive]外部文件导入到Hive的例子
外部文件导入到Hive的例子: [training@localhost ~]$ cd ~[training@localhost ~]$ pwd/home/training[training@local ...
- Hive的单节点集群详细启动步骤
说在前面的话, 在这里,推荐大家,一定要先去看这篇博客,如下 再谈hive-1.0.0与hive-1.2.1到JDBC编程忽略细节问题 Hadoop Hive概念学习系列之hive三种方式区别和搭建. ...
- Caused by: java.net.ConnectException: Connection refused: master/192.168.3.129:7077
1:启动Spark Shell,spark-shell是Spark自带的交互式Shell程序,方便用户进行交互式编程,用户可以在该命令行下用scala编写spark程序. 启动Spark Shell, ...
- hive
Hive Documentation https://cwiki.apache.org/confluence/display/Hive/Home 2016-12-22 14:52:41 ANTLR ...
- hive学习笔记
html,body,div,span,applet,object,iframe,h1,h2,h3,h4,h5,h6,p,blockquote,pre,a,abbr,acronym,address,bi ...
- Hive 笔记
DESCRIBE EXTENDED mydb.employees DESCRIBE EXTENDED mydb.employees DESCRIBE EXTENDED mydb.employees ...
- Hadoop学习笔记—17.Hive框架学习
一.Hive:一个牛逼的数据仓库 1.1 神马是Hive? Hive 是建立在 Hadoop 基础上的数据仓库基础构架.它提供了一系列的工具,可以用来进行数据提取转化加载(ETL),这是一种可以存储. ...
随机推荐
- 开启JMX功能,使JVisvualVM能够连接JVM
-Dcom.sun.management.jmxremote.port=1099 -Dcom.sun.management.jmxremote.ssl=false -Dcom.sun.manageme ...
- vuejs里封装的和IOS,Android通信模块
项目需要,在vuejs开发的web项目中与APP进行通信,实现原理和cordova一致.使用WebViewJavascriptBridge. 其实也是通过拦截url scheme,支持ios6往前的系 ...
- 电脑文件出现“windows-文件发生意外问题-可修复(严禁修改)-错误代码0X00000BF8”错误,怎么办
电脑文件出现"windows-文件发生意外问题-可修复(严禁修改)-错误代码0X00000BF8"错误,怎么办 下载一个"纵情文件修复器"修复一下就可以了 下载 ...
- SSIS2012 项目部署模型
SSIS 2012 支持两种部署模型:项目部署模型和包部署模型. 使用项目部署模型可以将项目部署到 Integration Services 服务器,使用包部署模型可以将单独的包部署到Integrat ...
- RabbitMq中的交换机
Rabbitmq的核心概念(如下图所示):有虚拟主机.交换机.队列.绑定: 交换机可以理解成具有路由表的路由程序,仅此而已.每个消息都有一个称为路由键 ...
- jdbc 数据的增删改查的Statement Resultset PreparedStatement
完成数据库的连接,就马上要对数据库进行增删改查操作了:先来了解一下Statement 通过JDBC插入数据 (这里提供一个查找和插入方法) Statement:用于执行sql语句的对象: *1.通过C ...
- Testlink部署全攻略
部署前准备: xampp,我下载的链接:https://www.apachefriends.org/download.html Testlink,下载地址:https://sourceforge.ne ...
- 2.5 C#的数据类型
在我们定义变量的时候需要使用数据类型,不同数据类型定义的变量,它的值的表现形式不同.比如整型主要表示整数,浮点型表示小数等等. C#中的数据类型有很多同C语言的相同,先学习一些简单的数据类型,其他的以 ...
- mysql时间格式化,按时间段查询的MySQL语句
描述:有一个会员表,有个birthday字段,值为'YYYY-MM-DD'格式,现在要查询一个时间段内过生日的会员,比如'06-03'到'07-08'这个时间段内所有过生日的会员. SQL语句: Se ...
- mono支持gb2312
需要安装mono-locale-extras 输入命令 yum install -y mono-locale-extras 安装即可