Data Analysis and Visualization class-2

DODO·2026년 3월 14일

Review_for class

목록 보기
11/20

my review

EDA(Exploratory Data Analysis)

Stage 1: Overall Data Overview - Descriptive Statistics

Stage 2: Exploring Differences by Target - Group Analysis

Stage 3: Verifying the Authenticity of Differences - Statistical Hypothesis Testing

Stage 4: Identifying Relationships Between Variables - Correlation Analysis

Environment Setup

수업 전 준비 (Pre-class Preparation)
  • !apt-get update -qq !apt-get install -y fonts-nanum* > /dev/null !fc-list :lang=ko !rm -rf ~/.cache/matplotlib/*
  • !ls /usr/share/fonts/truetype/nanum
  • import pandas as pd, import numpy as np
  • import matplotlib.pyplot as plt
  • import matplotlib.pyplot as plt, from scipy import stats
  • !pip install kagglehub, import pandas as pd, import kagglehub

basic words

index : the range and data type of the index
entries : the total number of rows in the dataset<
data columns : the total count and names of each individual column
non-null count : the number of actual data points, excluding missing values(=NaN)
Dytpe : the data type of each column
Memory usage : the total amount of memory occupied by the DataFrame

Descriptive analysis - describe()

Descriptive Statistics: describe()

Detailed metrics for numerical data

  • Count: Number of actual data points (excluding Null).
  • Mean & Std: Average and standard deviation (dispersion).
  • Min & Max: The range of the data.
  • Quartiles (25%, 50%, 75%): Points that divide the data into four equal parts (Q1, Q2/Median, Q3).

example

  • import pandas as pd

    continuous_vars = [
    'BMI',
    'MentHlth',
    'PhysHlth',
    'Age'
    ]

    dfmean=df[continuous_vars].mean() # type(dfmean) - Series -(4.1)

    dfmedian=df[continuous_vars].median() # type(dfmedian) - Series -(4.1)

    dfmode=df[continuous_vars].mode() # type(dfmode) -dataframe - (1.4)

0개의 댓글