XGBoost对波士顿房价进行预测

import numpy as np

import matplotlib as mpl

mpl.rcParams["font.sans-serif"] = ["SimHei"]

import matplotlib.pyplot as plt

import pandas as pd

from sklearn.model_selection  import train_test_split

from sklearn.metrics import mean_squared_error

import xgboost as xgb

def notEmpty(s):

    return s != ''

names = ['CRIM','ZN', 'INDUS','CHAS','NOX','RM','AGE','DIS','RAD','TAX','PTRATIO','B','LSTAT']

path = "datas/boston_housing.data"

## 由于数据文件格式不统一，所以读取的时候，先按照一行一个字段属性读取数据，然后再按照每行数据进行处理

fd = pd.read_csv(path, header=None)

data = np.empty((len(fd), 14))

for i, d in enumerate(fd.values):

    d = map(float, filter(notEmpty, d[0].split(' ')))

    data[i] = list(d)

x, y = np.split(data, (13,), axis=1)

y = y.ravel()

print ("样本数据量:%d, 特征个数：%d" % x.shape)

print ("target样本数据量:%d" % y.shape[0])

样本数据量:506, 特征个数：13

target样本数据量:506

# 查看数据信息

X_DF = pd.DataFrame(x)

X_DF.info()

X_DF.describe().T

X_DF.head()

<class 'pandas.core.frame.DataFrame'>

RangeIndex: 506 entries, 0 to 505

Data columns (total 13 columns):

0     506 non-null float64

1     506 non-null float64

2     506 non-null float64

3     506 non-null float64

4     506 non-null float64

5     506 non-null float64

6     506 non-null float64

7     506 non-null float64

8     506 non-null float64

9     506 non-null float64

10    506 non-null float64

11    506 non-null float64

12    506 non-null float64

dtypes: float64(13)

memory usage: 51.5 KB

#数据的分割，

x_train, x_test, y_train, y_test = train_test_split(x, y, train_size=0.8, random_state=14)

print ("训练数据集样本数目：%d, 测试数据集样本数目：%d" % (x_train.shape[0], x_test.shape[0]))

训练数据集样本数目：404, 测试数据集样本数目：102

# XGBoost将数据转换为XGBoost可用的数据类型

dtrain = xgb.DMatrix(x_train, label=y_train)

dtest = xgb.DMatrix(x_test)

# XGBoost模型构建

# 1. 参数构建

params = {'max_depth':2, 'eta':1, 'silent':1, 'objective':'reg:linear'}

num_round = 2

# 2. 模型训练

bst = xgb.train(params, dtrain, num_round)

# 3. 模型保存

bst.save_model('xgb.model')

# XGBoost模型预测

y_pred = bst.predict(dtest)

print(mean_squared_error(y_test, y_pred))

24.869737956719252

# 4. 加载模型

bst2 = xgb.Booster()

bst2.load_model('xgb.model')

# 5 使用加载模型预测

y_pred2 = bst2.predict(dtest)

print(mean_squared_error(y_test, y_pred2))

24.869737956719252

# 画图

## 7. 画图

plt.figure(figsize=(12,6), facecolor='w')

ln_x_test = range(len(x_test))

plt.plot(ln_x_test, y_test, 'r-', lw=2, label=u'实际值')

plt.plot(ln_x_test, y_pred, 'g-', lw=4, label=u'XGBoost模型')

plt.xlabel(u'数据编码')

plt.ylabel(u'租赁价格')

plt.legend(loc = 'lower right')

plt.grid(True)

plt.title(u'波士顿房屋租赁数据预测')

plt.show()

from xgboost import plot_importance

from matplotlib import pyplot

# 找出最重要的特征

plot_importance(bst,importance_type = 'cover')

pyplot.show()

XGBoost对波士顿房价进行预测的更多相关文章

波士顿房价预测 - 最简单入门机器学习 - Jupyter
机器学习入门项目分享 - 波士顿房价预测该分享源于Udacity机器学习进阶中的一个mini作业项目,用于入门非常合适,刨除了繁琐的部分,保留了最关键.基本的步骤,能够对机器学习基本流程有一个最清晰 ...
Tensorflow之多元线性回归问题（以波士顿房价预测为例）
一.根据波士顿房价信息进行预测,多元线性回归+特征数据归一化 #读取数据 %matplotlib notebook import tensorflow as tf import matplotlib. ...
机器学习实战二：波士顿房价预测 Boston Housing
波士顿房价预测 Boston housing 这是一个波士顿房价预测的一个实战,上一次的Titantic是生存预测,其实本质上是一个分类问题,就是根据数据分为1或为0,这次的波士顿房价预测更像是预测一 ...
AdaBoost 算法-分析波士顿房价数据集
公号:码农充电站pro 主页:https://codeshellme.github.io 在机器学习算法中,有一种算法叫做集成算法,AdaBoost 算法是集成算法的一种.我们先来看下什么是集成算法. ...
《用Python玩转数据》项目—线性回归分析入门之波士顿房价预测（二）
接上一部分,此篇将用tensorflow建立神经网络,对波士顿房价数据进行简单建模预测. 二.使用tensorflow拟合boston房价datasets 1.数据处理依然利用sklearn来分训练集 ...
机器学习之路：python 集成回归模型随机森林回归RandomForestRegressor 极端随机森林回归ExtraTreesRegressor GradientBoostingRegressor回归预测波士顿房价
python3 学习机器学习api 使用了三种集成回归模型 git: https://github.com/linyi0604/MachineLearning 代码: from sklearn.dat ...
机器学习之路: python 回归树 DecisionTreeRegressor 预测波士顿房价
python3 学习api的使用 git: https://github.com/linyi0604/MachineLearning 代码: from sklearn.datasets import ...
机器学习之路：python k近邻回归预测波士顿房价
python3 学习机器学习api 使用两种k近邻回归模型分别是平均k近邻回归和距离加权k近邻回归进行预测 git: https://github.com/linyi0604/Machine ...
机器学习之路: python 线性回归LinearRegression, 随机参数回归SGDRegressor 预测波士顿房价
python3学习使用api 线性回归,和随机参数回归 git: https://github.com/linyi0604/MachineLearning from sklearn.datasets ...

随机推荐

C# 读取Excel 单元格是日期格式
原文地址:https://www.cnblogs.com/liu-xia/p/5230768.html DateTime.FromOADate(double.Parse(range.Value2.To ...
微信小程序与云开发
微信小程序基础概念小程序云开发的三大基础能力:云数据库.云函数.云存储 Java.NodeJS.JavaScript.HTML5.CSS3.VueJs.ReactJs.前端工程化.前端架构小程序开 ...
[golang]Go常见问题：# command-line-arguments: ***: undefined: ***
今天遇见一个很蛋疼的问题,不知道是不是我配置的问题,IDE直接run就报错. 问题描述在开发代码过程中,经常会因为逻辑处理而对代码进行分类,放进不同的文件里面:像这样,同一个包下的两个文件,点击id ...
@Aspect注解并不属于@Component的一种
也就是一个类单纯如果只添加了@Aspect注解,那么它并不能被context:component-scan标签扫描到. 想要被扫描到的话,需要追加一个@Component注解
【牛客】小w的魔术扑克（并查集？？树状数组）
题目描述小w喜欢打牌,某天小w与dogenya在一起玩扑克牌,这种扑克牌的面值都在1到n,原本扑克牌只有一面,而小w手中的扑克牌是双面的魔术扑克(正反两面均有数字,可以随时进行切换),小w这个人就准 ...
Artifact tlks: com.intellij.javaee.oss.admin.jmx.JmxAdminException: com.intellij.execution.ExecutionException: E:\IDEAspace\tlksArtfacts\tlks.war not found for the web module.
传送门:https://www.cnblogs.com/eatkid/p/8064763.html intellij idea tomcat 启动不生成war包 intellij idea tom ...
Calcite分析 - RelTrait
RelTrait 表示RelNode的物理属性由RelTraitDef代表RelTrait的类型 /** * RelTrait represents the manifestation of a r ...
SQLServer 截取函数 substring函数
declare @name char(1000) --注意:char(10)为10位,要是位数小了会让数据出错 set @name='s{sss}fc{fggh}dghdf{cccs}x' selec ...
Post Setting Proxy 设置代理
postman的代理使用篇(四) - codingstudy - SegmentFault 思否https://segmentfault.com/a/1190000012024844 postman ...
用Java和Nodejs获取http30X跳转后的url
用Java和Nodejs获取http30X跳转后的url 转 https://calfgz.github.io/blog/2018/05/http-redirect-java-node.html 30 ...

XGBoost对波士顿房价进行预测

24.869737956719252

24.869737956719252

XGBoost对波士顿房价进行预测的更多相关文章

随机推荐

热门专题