【udacity】机器学习-波士顿房价预测

import numpy as np

import pandas as pd

from Udacity.model_check.boston_house_price import visuals as vs # Supplementary code

from sklearn.model_selection import ShuffleSplit

# Pretty display for notebooks

# 让结果在notebook中显示

# Load the Boston housing dataset

# 载入波士顿房屋的数据集

data = pd.read_csv('housing.csv')

prices = data['MEDV']

features = data.drop('MEDV', axis=1)

# print(data.describe())

# Success

# 完成

print("Boston housing dataset has {} data points with {} variables each.".format(*data.shape))

# 目标：计算价值的最小值

minimum_price = np.min(data['MEDV'])

# 目标：计算价值的最大值

maximum_price = np.max(data['MEDV'])

# 目标：计算价值的平均值

mean_price = np.mean(data['MEDV'])

# 目标：计算价值的中值

median_price = np.median(data['MEDV'])

# 目标：计算价值的标准差

std_price = np.std(data['MEDV'])

# 目标：输出计算的结果

print("Statistics for Boston housing dataset:\n")

print("Minimum price: ${:,.2f}".format(minimum_price))

print("Maximum price: ${:,.2f}".format(maximum_price))

print("Mean price: ${:,.2f}".format(mean_price))

print("Median price ${:,.2f}".format(median_price))

print("Standard deviation of prices: ${:,.2f}".format(std_price))

# RM,LSTAT,PTRATIO,MEDV

"""

初步分析结果是

1.RM越大MEDV越大

2.LSTATA越大MEDV越小

3.PTRATIO越大MEDV越小

"""

# TODO: Import 'r2_score'

def performance_metric(y_true, y_predict):

    """ Calculates and returns the performance score between

        true and predicted values based on the metric chosen. """

    from sklearn.metrics import r2_score

    # TODO: Calculate the performance score between 'y_true' and 'y_predict'

    score = r2_score(y_true,y_predict)

    # Return the score

    return score

# score = performance_metric([3, -0.5, 2, 7, 4.2], [2.5, 0.0, 2.1, 7.8, 5.3])

# print ("Model has a coefficient of determination, R^2, of {:.3f}.".format(score))

from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(features, prices, test_size=0.80, random_state=1)

# Success

print ("Training and testing split was successful.")

# vs.ModelLearning(features, prices)

def fit_model(X, y):

    """ Performs grid search over the 'max_depth' parameter for a

        decision tree regressor trained on the input data [X, y]. """

    from sklearn.tree import DecisionTreeRegressor

    from sklearn.model_selection import KFold

    # Create cross-validation sets from the training data

    cross_validator = KFold(10)

    # cv_sets = ShuffleSplit(X.shape[0],  test_size=0.20, random_state=0)

    # TODO: Create a decision tree regressor object

    regressor = DecisionTreeRegressor()

    # TODO: Create a dictionary for the parameter 'max_depth' with a range from 1 to 10

    max_depth = [1,2,3,4,5,6,7,8,9,10]

    params = {"max_depth":max_depth}

    from sklearn.metrics import make_scorer

    # TODO: Transform 'performance_metric' into a scoring function using 'make_scorer'

    scoring_fnc = make_scorer(performance_metric)

    from sklearn.model_selection import GridSearchCV

    # TODO: Create the grid search object

    grid = GridSearchCV(regressor,params,scoring_fnc,cv=cross_validator)

    # Fit the grid search object to the data to compute the optimal model

    grid = grid.fit(X, y)

    # Return the optimal model after fitting the data

    return grid.best_estimator_

reg = fit_model(X_train, y_train)

# Produce the value for 'max_depth'

print ("Parameter 'max_depth' is {} for the optimal model.".format(reg.get_params()['max_depth']))

client_data = [[5, 17, 15], # Client 1

               [4, 32, 22], # Client 2

               [8, 3, 12]]  # Client 3

# Show predictions

for i, price in enumerate(reg.predict(client_data)):

    print ("Predicted selling price for Client {}'s home: ${:,.2f}".format(i+1, price))

【udacity】机器学习-波士顿房价预测的更多相关文章

【udacity】机器学习-波士顿房价预测小结
Evernote Export 机器学习的运行步骤 1.导入数据没什么注意的,成功导入数据集就可以了,打印看下数据的标准格式就行用个info和describe 2.分析数据这里要详细分析数据的内 ...
波士顿房价预测 - 最简单入门机器学习 - Jupyter
机器学习入门项目分享 - 波士顿房价预测该分享源于Udacity机器学习进阶中的一个mini作业项目,用于入门非常合适,刨除了繁琐的部分,保留了最关键.基本的步骤,能够对机器学习基本流程有一个最清晰 ...
机器学习实战二：波士顿房价预测 Boston Housing
波士顿房价预测 Boston housing 这是一个波士顿房价预测的一个实战,上一次的Titantic是生存预测,其实本质上是一个分类问题,就是根据数据分为1或为0,这次的波士顿房价预测更像是预测一 ...
Python之机器学习-波斯顿房价预测
目录波士顿房价预测导入模块获取数据打印数据特征选择散点图矩阵关联矩阵训练模型可视化波士顿房价预测导入模块 import pandas as pd import numpy as ...
Tensorflow之多元线性回归问题（以波士顿房价预测为例）
一.根据波士顿房价信息进行预测,多元线性回归+特征数据归一化 #读取数据 %matplotlib notebook import tensorflow as tf import matplotlib. ...
《用Python玩转数据》项目—线性回归分析入门之波士顿房价预测（二）
接上一部分,此篇将用tensorflow建立神经网络,对波士顿房价数据进行简单建模预测. 二.使用tensorflow拟合boston房价datasets 1.数据处理依然利用sklearn来分训练集 ...
chapter02 回归模型在''美国波士顿房价预测''问题中实践
#coding=utf8 # 从sklearn.datasets导入波士顿房价数据读取器. from sklearn.datasets import load_boston # 从sklearn.mo ...
基于sklearn的波士顿房价预测_线性回归学习笔记
> 以下内容是我在学习https://blog.csdn.net/mingxiaod/article/details/85938251 教程时遇到不懂的问题自己查询并理解的笔记,由于sklear ...
02-11 RANSAC算法线性回归(波斯顿房价预测)
目录 RANSAC算法线性回归(波斯顿房价预测) 一.RANSAC算法流程二.导入模块三.获取数据四.训练模型五.可视化更新.更全的<机器学习>的更新网站,更有python.go ...

随机推荐

Co-prime
Co-prime Time Limit: 2000/1000 MS (Java/Others) Memory Limit: 32768/32768 K (Java/Others) Problem ...
oracle到mysql的导数据方式（适用于任意数据源之间的互导）
http://www.wfuyu.com/Internet/19955.html 为了生产库释放部份资源, 需要将API模块迁移到mysql中,及需要导数据. 尝试了oracle to mysql工具 ...
FZU - 1606 - Format the expression
先上题目: Problem 1606 Format the expression Accept: 87 Submit: 390Time Limit: 1000 mSec Memory Li ...
Linux去重命令uniq（转）
注意:需要先排序sort才能使用去重. Linux uniq命令用于检查及删除文本文件中重复出现的行列. uniq可检查文本文件中重复出现的行列. 语法 uniq [-cdu][-f<栏位> ...
例说linux内核与应用数据通信（三）：读写内核设备驱动文件
[版权声明:尊重原创,转载请保留出处:blog.csdn.net/shallnet.文章仅供学习交流.请勿用于商业用途] 读写设备文件也就是调用系统调用read()和write(),系 ...
64bit Centos6.4搭建hadoop-2.5.1
64bit Centos6.4搭建hadoop-2.5.1 1.分布式环境搭建採用4台安装Linux环境的机器来构建一个小规模的分布式集群. 当中有一台机器是Master节点,即名称节点,另外三台是 ...
Online Object Tracking: A Benchmark 论文笔记
Factors that affect the performance of a tracing algorithm 1 Illumination variation 2 Occlusion 3 Ba ...
循环神经网络(RNN, Recurrent Neural Networks)——无非引入了环，解决时间序列问题
摘自:http://blog.csdn.net/heyongluoyao8/article/details/48636251 不同于传统的FNNs(Feed-forward Neural Networ ...
lightoj--1005--Rooks（组合数）
Rooks Time Limit: 1000MS Memory Limit: 32768KB 64bit IO Format: %lld & %llu Submit Status De ...
Path Sum II 总结DFS
https://oj.leetcode.com/problems/path-sum-ii/ Given a binary tree and a sum, find all root-to-leaf p ...

【udacity】机器学习-波士顿房价预测

【udacity】机器学习-波士顿房价预测的更多相关文章

随机推荐

热门专题