빅데이터분석기사 실기 3유형 (Python)

송범·2024년 11월 27일

데이터 자격 검정 사이트에서 제공해주는 실습 환경

https://dataq.goorm.io/exam/3/%EC%B2%B4%ED%97%98%ED%95%98%EA%B8%B0/quiz/1

import pandas as pd

df = pd.read_csv("data/Titanic.csv")

# 사용자 코딩

# 1번 
from scipy.stats import chi2_contingency, ttest_1samp, ttest_ind, ttest_rel, chisquare

table = pd.crosstab(df['Gender'],df['Survived'])
# statistic, p ,df , expected = chi2_contingency(table)/
# print(round(statistic,3))  260.717

# 2번 
from statsmodels.api import Logit
from sklearn.preprocessing import LabelEncoder
import statsmodels.api as sm

encoder = LabelEncoder()
df['Gender'] = encoder.fit_transform(df['Gender'])

X = df[['Gender','SibSp','Parch','Fare']]
X = sm.add_constant(X)
y = df['Survived']

model = Logit(y,X)

results = model.fit()
print(results.summary()) # -0.201

# 2번 다른 풀이 방법 
# Logit(로지스틱 회귀) OLS(선형회귀)
from statsmodels.api import Logit, OLS
import statsmodels.api as sm

formula = "Survived ~ Gender + SibSp + Parch + Fare"
results = Logit.from_formula(formula,df).fit()
print(results.summary()) # 결과 확인


# 3번 
import numpy as np

print(round(np.exp(results.params['SibSp']),3)) # 0.702
print(round(np.exp(-0.3539),3)) # 0.702 반환 2번에서 나온 결과를 토대로 동일한 값 확인 가능


import pandas as pd

df = pd.read_csv("sample_data/churn.csv")

from statsmodels.formula.api import logit
model = logit(formula="Churn~AccountWeeks+ContractRenewal+DataPlan+DataUsage+CustServCalls+DayMins+DayCalls+MonthlyCharge+OverageFee+RoamMins",data=df).fit()

# print(model.summary()) # DataUsage+DayMins

model = logit(formula="Churn~DataUsage+DayMins",data=df).fit()


print(model.summary())

print(-0.0039+-0.1697+-1.0395)
import numpy as np


round(np.exp(-0.1697*5),3)
Optimization terminated successfully.
         Current function value: 0.393603
         Iterations 6
Optimization terminated successfully.
         Current function value: 0.397599
         Iterations 6
                           Logit Regression Results                           
==============================================================================
Dep. Variable:                  Churn   No. Observations:                 1000
Model:                          Logit   Df Residuals:                      997
Method:                           MLE   Df Model:                            2
Date:                Thu, 28 Nov 2024   Pseudo R-squ.:                 0.01375
Time:                        09:56:19   Log-Likelihood:                -397.60
converged:                       True   LL-Null:                       -403.14
Covariance Type:            nonrobust   LLR p-value:                  0.003908
==============================================================================
                 coef    std err          z      P>|z|      [0.025      0.975]
------------------------------------------------------------------------------
Intercept     -1.0395      0.303     -3.434      0.001      -1.633      -0.446
DataUsage     -0.1697      0.071     -2.376      0.017      -0.310      -0.030
DayMins       -0.0039      0.002     -2.264      0.024      -0.007      -0.001
==============================================================================
-1.2131
0.428
profile
BackEnd&Data Scientist가 되고 싶은 개발 기록 노트

0개의 댓글