Stage 1: Overall Data Overview - Descriptive Statistics
Stage 2: Exploring Differences by Target - Group Analysis
Stage 3: Verifying the Authenticity of Differences - Statistical Hypothesis Testing
Stage 4: Identifying Relationships Between Variables - Correlation Analysis
index : the range and data type of the index
entries : the total number of rows in the dataset<
data columns : the total count and names of each individual column
non-null count : the number of actual data points, excluding missing values(=NaN)
Dytpe : the data type of each column
Memory usage : the total amount of memory occupied by the DataFrame
Detailed metrics for numerical data
continuous_vars = [
'BMI',
'MentHlth',
'PhysHlth',
'Age'
]
dfmean=df[continuous_vars].mean() # type(dfmean) - Series -(4.1)
dfmedian=df[continuous_vars].median() # type(dfmedian) - Series -(4.1)
dfmode=df[continuous_vars].mode() # type(dfmode) -dataframe - (1.4)