머신러닝 회귀 코드순서
1. 라이브러리 불러오기
2. 데이터 불러오기
3. 데이터 전처리
4. 입력값(X)과 정답값(y) 분리
5. 학습용·테스트용 데이터 분리
6. 모델 생성 및 학습
7. 예측
8. 성능 평가
9. 결과 시각화 및 변수 중요도 확인
회귀 : 숫자예측
X : 정보
y : 맞일 데이터
X_train, X_test : 모델 학습용 데이터
y_train, y_test : 모델 성능 확인 데이터
[1, 라이브러리 작성]
import pandas as pd
import seaborn as sns
import matplotlib.pyplot as plt
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeRegressor
from sklearn.ensemble import RandomForestRegressor
from sklearn.metrics import root_mean_squared_error
base_path = r"데이터 위치"
file_name = r"데이터 이름.csv"
file_path = os.path.join(base_path, file_name)
df = pd.read_csv(file_path)
df.head()
df.info()
df.describe()
df.drop(columns="열이름") # 특정 열삭제
df[df["열이름"] < 기준값] # 조건에 맞는 열만 남김
df["새로운 열"] = 계산식 # 새로운 열 만들기
df = pd.get_dummies({참고데이터, columns=[바꿀 열이름], drop_tirst=True, dtype=int})
X = df.drop(columns=예측대상의 열이름)
y = df[예측 대상 열이름]
X_train, y_train, y_test, y_train = train_test_split(
X,
y,
test_size=0.2,
random_state=2026
)
tree_model = DecisionTreeRagressor(
max_depth=6, # 최대 깊이
min_samples_leaf=5, # 최소 데이터 수
random_state=2026
)
tree_model.fit(X_train, y_train) ⭐️ .fit :모델 학습
(예측과 평가)
tree_pred = tree_model.predict(X_test) ⭐️ .predict : 예측
tree_rmse = root_mean_squared_error(y_test, tree_pred)
실제 가격과 모델이 예측한 가격차이가 얼마나 차이나는지 계산(확인)
실제 정답 모델의 예측
↓ ↓
y_test tree_pred
└─────────┬─────────────┘
↓
root_mean_squared_error()
↓
오차 계산
↓
RMSE
rf_model = RandomForestRegressor(
n_estimators=200, # 결정트리 200생성
random_state=2026 # 실행결과 고정
)
# 2. 학습
rf_model✨.fit(X_train, y_test) # rf_model: 학습이 끝난 랜덤포레스트
# 3. 예측
rf_pred = rf_model✨.predict(X_test) # predict: 예측해라, X_test:
# 4. 평가
rf_rmse = root_mean_squared_error(y_test, rf_pred) # rf_pred:모델이 예측한 문제
# ✨실제값과 예측값 비교-> RMSE 계산
# 5. 출력
print(f"랜덤 포레스트 RMSE: {rf_rmse:.3f}")
# 소수점 3자리 까지 표시
(✨ = 중요표시 )
-⭐️ 회귀 랜덤 포레스트에서는 각 트리가 내놓은 예측값을 평균함
입력 데이터
↓
결정 트리 1 ─┐
결정 트리 2 ─┼─ 평균 → 최종 가격 예측
결정 트리 3 ─┘
sns.scatterplot(x=y_test, y=rf_pred)
plt.plot(
[y_test.min(), y_test.max()],
[y_test.min(), y_test.max()],
color="red"
)
plt.xlabel("실제 판매 가격")
plt.ylabel("예측 판매 가격")
plt.show()
feature_importance = pd.DataFrame({
"feature": X.columns, # "feature"열에 X.columns 데이터 들어감 (X.columns : ✨변수명)
"importance": rf_model.feature_importances_
# ^(랜덤 포레스트가 학습한)변수✨중요도
}).sort_valies("importance", ascending=False)
# sort_valies() : 값을 기준으로 ✨정렬하는 함수
-> "importance"값을 기준으로 정렬, ascending=False 내림차순
print(feature_importance)
sns.barplot(
data=feature_importance,
x="importance",
y="feature"
)
pit.title("변수 중요도")
plt.show()