Data Analysis and Visualization class-3

DODO·2026년 3월 14일

Review_for class

목록 보기
12/20

review

Data Preprocessing

first step : Data cleaning
-Missing Values - NaN or None or Null
second step : Basic Transformation

Data Preprocessing - isnull, notnull
isnull : Null=> return True / isna == isnull
notnull : Null=> return False / notna == notnull


Handling missing values - dropna()
axis=0 del row , axis=1 del column
subset=['column's name']
inplace = False, inplace= True


Filling missing values - fillna()
fillna: constant imputation
ffill: forward fill
bfill: backward fill


Detecting Duplicate data - duplicated()
Ture: duplicated data
False: the first occurrence of the data, the original entry


delete duplicated row - drop_dulicates
keep option
1. first : keep the first occurrence of the row(basic value)
2. last : keep the last occurrence of the row

Outliers

ex: bp 300/200, bst 700
Data Error -> drop, clinical extreme -> capping

Search Outliters to use IQR
IQR < Q1-1.5IQR & IQR > Q3+1.5IQR
Q1 = df.quantile(0.25) , Q3 = df.quantile(0.75)
IQR = Q3-Q1

Problem about data types
astype() - type casting
pd.to_numeric() | pd.to_datetime()

0개의 댓글