机器学习和引荐系统(二十二)分类模型 – KNN代码完成(下)
作者头像
  • AI科技评论
  • 2020-05-11 21:29:19 2

分类模型 – KNN代码实现(下)

1. 引入依赖

首先,我们需要导入一些必要的库:

```python import numpy as np import pandas as pd

导入鸢尾花数据集

from sklearn.datasets import load_iris

划分数据集为训练集和测试集

from sklearn.modelselection import traintest_split

用于计算分类预测准确率

from sklearn.metrics import accuracy_score ```

2. 数据加载与预处理

接下来,我们加载鸢尾花数据集并对其进行预处理:

```python iris = load_iris()

将数据集转换为DataFrame格式

df = pd.DataFrame(data=iris.data, columns=iris.featurenames) df['class'] = iris.target df['class'] = df['class'].map({0: iris.targetnames[0], 1: iris.targetnames[1], 2: iris.targetnames[2]}) ```

查看数据描述统计信息:

python df.describe()

划分训练集和测试集:

```python x = iris.data y = iris.target.reshape(-1, 1)

划分数据集

xtrain, xtest, ytrain, ytest = traintestsplit(x, y, testsize=0.3, randomstate=35, stratify=y) ```

3. 核心算法实现

定义间隔函数和KNN分类器:

```python

间隔函数定义

def l1_distance(a, b): return np.sum(np.abs(a - b), axis=1)

def l2_distance(a, b): return np.sqrt(np.sum((a - b)**2, axis=1))

KNN分类器

class KNN: def init(self, nneighbors=1, distfunc=l1distance): self.nneighbors = nneighbors self.distfunc = dist_func

def fit(self, x, y):
    self.x_train = x
    self.y_train = y

def predict(self, x):
    y_predict = np.zeros((x.shape[0], 1), dtype=self.y_train.dtype)

    for i, x_test in enumerate(x):
        distance = self.dist_func(self.x_train, x_test)
        nn_index = np.argsort(distance)
        nn_y = self.y_train[nn_index[:self.n_neighbors]].ravel()
        y_predict[i] = np.argmax(np.bincount(nn_y))
    return y_predict

```

4. 测试

通过实例化KNN对象来验证模型性能:

```python

实例化KNN对象

knn = KNN(n_neighbors=3)

训练模型

knn.fit(xtrain, ytrain)

预测测试数据

ypredict = knn.predict(xtest)

计算预测准确率

accuracy = accuracyscore(ytest, y_predict) print("预测准确率:", accuracy) ```

为了评估不同参数的效果,我们可以尝试不同的距离函数和K值:

```python

存储结果

result_list = []

针对不同的参数进行预测

for p in [1, 2]: knn.distfunc = l1distance if p == 1 else l2_distance

# 测试不同的K值
for k in range(1, 10, 2):
    knn.n_neighbors = k
    y_predict = knn.predict(x_test)
    accuracy = accuracy_score(y_test, y_predict)
    result_list.append([k, 'l1_distance' if p == 1 else 'l2_distance', accuracy])

df = pd.DataFrame(result_list, columns=['k', '距离函数', '预测准确率']) ```

通过以上步骤,我们完成了KNN算法的实现,并对其进行了详细的测试与评估。

    本文来源:图灵汇
责任编辑: : AI科技评论
声明:本文系图灵汇原创稿件,版权属图灵汇所有,未经授权不得转载,已经协议授权的媒体下载使用时须注明"稿件来源:图灵汇",违者将依法追究责任。
    分享
引荐模型机器完成代码学习分类系统KNN二十二
    下一篇