어제 데이터 전처리 후에도 ML modeling을 진행했으나 성능이 큰 차이를 보이지 않았고 통계적 가설 검정을 통해 종속 변수와 독립 변수 간의 관계를 살펴보려고 한다.
| 종속변수 | 독립변수 | 분석 방법 |
|---|---|---|
| 범주형 | 범주형 | 카이제곱 검정 (Chi-square test) |
| 범주형 | 연속형 | 로지스틱 회귀 분석 (Logistic Regression) |
from scipy.stats import chi2_contingency
target_col = '인지장애 경험 여부'
categorical_col = [
col for col in df_old.columns
if col not in continuous_col and col != target_col
]
chi2_results = []
for col in categorical_col:
contingency = pd.crosstab(df_old[col], df_old[target_col])
if contingency.shape[0] >= 2 and contingency.shape[1] >= 2:
chi2, p, dof, expected = chi2_contingency(contingency)
chi2_results.append({
'독립변수': col,
'Chi2': chi2,
'p-value': p,
'유의': 'Yes' if p < 0.05 else 'No'
})
result_df = pd.DataFrame(chi2_results)
result_df = result_df.sort_values(by='p-value')
print(result_df)