머신러닝 정리

이예나·3일 전

머신러닝 회귀 코드순서
1. 라이브러리 불러오기
2. 데이터 불러오기
3. 데이터 전처리
4. 입력값(X)과 정답값(y) 분리
5. 학습용·테스트용 데이터 분리
6. 모델 생성 및 학습
7. 예측
8. 성능 평가
9. 결과 시각화 및 변수 중요도 확인

  • 회귀 : 숫자예측

    	X : 정보 
    	y : 맞일 데이터 
    	X_train, X_test : 모델 학습용 데이터
    	y_train, y_test : 모델 성능 확인 데이터 

[1, 라이브러리 작성]

import pandas as pd
import seaborn as sns
import matplotlib.pyplot as plt

from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeRegressor
from sklearn.ensemble import RandomForestRegressor
from sklearn.metrics import root_mean_squared_error
  1. 데이터 불러오기
		
	base_path = r"데이터 위치"
	file_name = r"데이터 이름.csv"
	file_path = os.path.join(base_path, file_name) 
		
	df = pd.read_csv(file_path)
	df.head()
  1. 전처리 작성
	df.info()
	df.describe()    
  • 불필요한 정보 제거 (이름, 이상치 ...)
  • 필요없는 열 제거
    ✨[자주사용]
df.drop(columns="열이름")  # 특정 열삭제
df[df["열이름"] < 기준값]    # 조건에 맞는 열만 남김 
df["새로운 열"] = 계산식     #  새로운 열 만들기 
  1. 범주형 데이터 변환
  • 분류형 -> 수치형
df = pd.get_dummies({참고데이터, columns=[바꿀 열이름], drop_tirst=True, dtype=int})
  1. X와 y분리 작성
X = df.drop(columns=예측대상의 열이름)
y = df[예측 대상 열이름]
  1. 학습용, 테스트용 데이터 분리
X_train, y_train, y_test, y_train = train_test_split(
	X,
    y,
    test_size=0.2,
    random_state=2026
)
  1. 결정 트리
  • 작성법
    (예시)
tree_model = DecisionTreeRagressor(
	max_depth=6,               # 최대 깊이
    min_samples_leaf=5,        # 최소 데이터 수 
    random_state=2026
)
tree_model.fit(X_train, y_train)  ⭐️ .fit :모델 학습

(예측과 평가)

tree_pred = tree_model.predict(X_test)     ⭐️ .predict : 예측 
tree_rmse = root_mean_squared_error(y_test, tree_pred)

실제 가격과 모델이 예측한 가격차이가 얼마나 차이나는지 계산(확인)

실제 정답              모델의 예측
   ↓                       ↓
y_test                 tree_pred
   └─────────┬─────────────┘
             ↓
root_mean_squared_error()
             ↓
          오차 계산
             ↓
            RMSE
            
  1. 랜덤 포래스트 작성법
  • 여러개의 결정 트리를 만든 뒤 각 트리의 예측을 평균 내는 모델
    (예시)
rf_model = RandomForestRegressor(
	n_estimators=200,               # 결정트리 200생성
    random_state=2026				# 실행결과 고정
)
# 2. 학습
rf_model✨.fit(X_train, y_test)		# rf_model: 학습이 끝난 랜덤포레스트 

# 3. 예측 
rf_pred = rf_model✨.predict(X_test)	# predict: 예측해라, X_test: 

# 4. 평가
rf_rmse = root_mean_squared_error(y_test, rf_pred)	# rf_pred:모델이 예측한 문제
							# ✨실제값과 예측값 비교-> RMSE 계산 
# 5. 출력
print(f"랜덤 포레스트 RMSE: {rf_rmse:.3f}")
						# 소수점 3자리 까지 표시
(✨ = 중요표시 )

-⭐️ 회귀 랜덤 포레스트에서는 각 트리가 내놓은 예측값을 평균함

  • rf_pred : X_test 정체에 개한 예측값 묶음-> 리스트로 반환됨
    (rf_pred = rf_model.predict(X_test) ->[1, 2, 3 ...])
  • predict(X): 입력된 샘플 각각에 대한 예측값을 반환
    모델의 예측 -> rf_pred
    실제정답 -> y_test

입력 데이터
↓
결정 트리 1 ─┐
결정 트리 2 ─┼─ 평균 → 최종 가격 예측
결정 트리 3 ─┘

  1. 실제값과 예측값 시각화
sns.scatterplot(x=y_test, y=rf_pred)

plt.plot(
    [y_test.min(), y_test.max()],
    [y_test.min(), y_test.max()],
    color="red"
)

plt.xlabel("실제 판매 가격")
plt.ylabel("예측 판매 가격")
plt.show()
  1. 변수 중요도 작성법
feature_importance = pd.DataFrame({
	"feature": X.columns, 			# "feature"열에 X.columns 데이터 들어감 (X.columns : ✨변수명)
    "importance": rf_model.feature_importances_ 
    				# ^(랜덤 포레스트가 학습한)변수✨중요도 
}).sort_valies("importance", ascending=False)
	# sort_valies() : 값을 기준으로 ✨정렬하는 함수 
    	-> "importance"값을 기준으로 정렬, ascending=False 내림차순

print(feature_importance)

sns.barplot(
	data=feature_importance,
    x="importance",
    y="feature"
)
pit.title("변수 중요도")
plt.show()
  • featureimportances : 랜덤 포레스트가 예측할 때 어떤 변수가 상대적으로 중요하게 사용했는지 보여줌 (값이 클수록 랜덤 포레스트가 예측할 때 그 변수를 많이 활용했다는 의미)

0개의 댓글