
๐ A key to supervised data mining is that we have some target quantity we would like to predict or to otherwise understand better
Fundamental concepts: Identifying informative attributes of the entities described by the data.
A model is a simplified representation of reality created to serve a purpose.
A predictive model may be judged solely on its predictive performance
Supervised learning(์ง๋ํ์ต) is model creation where the model describes a relationship between a set of selected variables (attributes or features) and a predefined variable called the target variable.
An instance or example represents a fact or a data point.โข An instance is described by a set of attributes(fields, columns, variables, or features.)โข An instance is also called a feature vector,because it can be represented as a fixed-length ordered collection (vector) of feature values
โข Model induction: The creation of models from data is known as model induction.(๋ชจ๋ธ์ ๋)
โข Induction algorithm or learner: The procedure that creates the model from the data is called the induction algorithm or learner.
โข Induction: Induction refers to generalizing from specific cases to general rules (๊ตฌ์ฒด์ ์ฌ์ค์ ์ผ๋ฐํ, ๊ท๋ฉ๋ฒ)
โข Deduction: Deduction starts with general rules and specific facts, and creates other specific facts from them (์ผ๋ฐ์ ์ฌ์ค์ ๊ตฌ์ฒดํ, ์ฐ์ญ๋ฒ)
โขTraining data: The input data for the induction algorithm, used for inducing the model, are called the training data.
โข Labeled data: The training data are called labeled data because the value for the target variable (the label) is known.
A predictive model focuses on estimating the value of some particular target variable of interest
๐ค How can we judge whether a variable contains important information about the target variable?
๐ Entropy ๋ก ์ฐพ์.
Entropy is a measure of disorder(๋ฌด์ง์ ์ ๋) that can beapplied to a set, such as one of our individual segments

-> ์ํธ๋กํผ๊ฐ ๋์์๋ก, ๋ฌด์ง์๊ฐ ๋์ ๊ฒ์ด๋ค.
Entropy ,only tells us how impure one individual subset is ๋ฐ๋ฉด์
Information gain(IG) measures how much an attribute improves (decreases) entropy over the whole segmentation it creates.
IG measures the change in entropy due to any amount of new information being added (์๋ก์ด ์ ๋ณด์ ์ถ๊ฐ์ ๋ฐ๋ฅธ ์ํธ๋กํผ์ ๋ณํ๋)
IG is a function of both a parent set and the children resulting from some partitioning of the parent set โ how much information has this attribute provided?


So this split reduces entropy substantially. In predictive modeling terms, the attribute provides a lot of information on the value of the target.
๐ค ๋ค๋ฅด๊ฒ ๋๋ ์ ์์ง ์์๊ฐ?
๐ ๋๋ ์๋ ์๋ค. ํ์ง๋ง ์ ๋ณด์ด๋์ด ์ ๋ค.์ ๋ณด์ด๋์ด ์ต๋๊ฐ ๋๋ ์์ฑ(attributes)์ฐพ์์ split ํ๋๊ฒ ๋ชฉํ์ธ ๊ฒ์ด๋ค.

๋ฐ๋ผ์
For a dataset with instances described by attributes and a target variable, we can determine which attribute is the most informative with respect to estimating the value of the target variable
๐ We also can rank a set of attributes by their informativeness, in particular by their information gain
-> IG ๊ฐ ๊ฐ์ฅ ๋์ attribute๋ฅผ ๊ณจ๋ผ์, ๊ฐ์ฅ ์ข์ ์์ฑ์ ์ ํํ์ฌ ์ด ์์ฑ์ ๊ธฐ์ค์ผ๋ก edible ํ์ง ์ํ์ง ๋๋์ด๋ณธ๋ค.

๐์ด ๊ทธ๋ํ๋ค์ ๊ฐ ์์ฑ ๊ฐ์ ๋ฐ๋ผ ๋ฐ์ดํฐ์ ์์ธก ๊ฐ๋ฅ์ฑ์ ๊ธฐ์ฌํ๋ ์ ๋๋ฅผ ์ํธ๋กํผ๋ฅผ ํตํด ์๊ฐํํ๊ณ ์์ต๋๋ค. Odor ์์ฑ์ฒ๋ผ ๋น์จ์ด ๋๊ณ ์ํธ๋กํผ๊ฐ ๋ฎ์ ์์ฑ์ ๋ฐ์ดํฐ์ ๋ถ๋ฅ์ ๋ ๋์์ด ๋ ์ ์์ผ๋ฉฐ, ๋ฐ๋๋ก ์ํธ๋กํผ๊ฐ ๋๊ณ ๋ค์ํ ๊ฐ์ด ๋ถํฌํ ์์ฑ์ ์์ธก์ ์์ด ์ ๋ณด์ ๋ถ์ฐ๋๊ฐ ํฌ๋ค๋ ์ ์ ์ ์ ์์ต๋๋ค.

The tree is made up of nodes,
Each interior node in the tree contains a test of an attribute
each path eventually terminates at a terminal node, or leaf.
Classificiation Tree

The tree is a supervised segmentation(์ง๋ ์ธ๋ถํ),because each leaf contains a value for the target variable.
such a tree is called a classification tree or more loosely a decision tree.

-> x์ถ y์ถ์ผ๋ก ๊ฐ ๋ถ๋ถ์ด ์ด๋ป๊ฒ ๋๋์๋์ง ๊ฐ๋จํ๊ฒ ์ดํด๋ณผ ์ ์๋ ์ฅ์ ์ด ์๋ค.


If we assign the same class probability to every member of the segment corresponding to a tree leaf, we can use instance counts at each leaf to compute a class probability estimate.
A frequency-based estimate of class membership probability may be overly optimistic about the probability of class membership for segments with very small numbers of instance. (์์ ์ํ์ ๋ํด์๋ ๊ณผ๋ํ๊ฒ ๋์ฒ์ ์ผ ์๋ ์์)
<- ์ด๋ฅผ ํด๊ฒฐํ๊ธฐ ์ํด Laplace ๋ณด์ ์ด ๋ํ๋๊ฒ ํจ.
Instead of simply computing the frequency, we use a โsmoothing(ํํํ)โ version of the frequency-based estimate (Laplace correction); the purpose of which is to moderate the influence of leaves with only a few instance
Laplace ๋ณด์ ์ ์ฃผ๋ก ๋์ด๋ธ ๋ฒ ์ด์ฆ(Naive Bayes) ๋ถ๋ฅ๊ธฐ์์ ์ฌ์ฉ๋๋ ๊ธฐ๋ฒ์ผ๋ก, ํ๋ จ ๋ฐ์ดํฐ์์ ๊ด์ฐฐ๋์ง ์์ ์ฌ๊ฑด์ ๋ํด 0์ด ์๋ ํ๋ฅ ์ ํ ๋นํ๋ ๋ฐฉ๋ฒ์.
Laplace ๋ณด์ ์ ์๋ ๋ฐฉ์
๊ธฐ๋ณธ ์์ด๋์ด: ๋ชจ๋ ๊ฐ๋ฅํ ๊ฒฐ๊ณผ์ ์์ ์์๊ฐ(๋ณดํต 1)์ ๋ํฉ๋๋ค.
์์:
์ฌ๊ธฐ์:
: ๋จ์ด
: ํด๋์ค
: ํด๋์ค c์์ ๋จ์ด w์ ์ถํ ํ์
: ํด๋์ค c์ ์ด ๋จ์ด ์
ฮฑ: ์ค๋ฌด๋ฉ ํ๋ผ๋ฏธํฐ (๋ณดํต 1)
: ๊ณ ์ ํ ๋จ์ด์ ์ด ๊ฐ์


As the number of instance increases, the Laplace equation converges to the frequency-based
estimate.
For each ratio the solid horizontal line shows the uncorrected (constant) estimate, while the corresponding dashed line shows the estimate with the Laplace correction applied.
The uncorrected line is the asymptote of the Laplace correction as the number of instances goes to infinity
๊ธ์ 100%, 80%, 66% - ์ค์ ์ผ๋ก Frequency Estimation ๋ํ๋ด๋ฉฐ
๊ทธ ๋ฐ์ ์ ์ ์ด Laplace Estimation ์ด๋ค.
instance ์๊ฐ ์ ์ ๋, Laplace ๋ณด์ ์ ๊ณผ๋ํ ํ๋ฅ ์ ์ํํ๋ ์ญํ ์ ํ๋ค.
๐ ํนํ ๋ฐ์ดํฐ์ ์๊ฐ ๋ถ์กฑํ ์ํฉ์์ ๊ณผ์ ํฉ์ ๋ฐฉ์ง.

๐ Information Gain์ ํตํด ์ค์ํ Attributes๊ฐ ๋ฌด์์ธ์ง ํ์
ํ ์ ์๊ฒ ๋์๋ค. House์ OVERAGE ํน์ฑ์ด ๊ณ ๊ฐ ํน์ฑ์์ ๊ฐ์ฅ ์ค์ํ ์์ธ์์ ์๋ ค์ค๋ค.

์์ฌ๊ฒฐ์ ํธ๋ฆฌ๊ฐ ์๊ฐํ๋์ด ์๋ค. House ์ ๊ฐ 600469์ ๋ฐ๋ผ์ ์ฒซ๋ฒ์งธ ๋ถ๊ธฐ๊ฐ ๋๋ฉฐ, ๊ฐ ๋ถ๊ธฐ์ ํน์ฑ์ ๋ฐ๋ผ์ ๋ค์ ๋๋ ์ ์๋ค.
a model can estimate the value of a target variable for a new unseen example.
์ ๋ณธ ๋ฐ์ดํฐ์์๋ predict ํจ
โ Use information gain, which is based on purity measure called entropy.
โ Given a large collection of data, we can find those variables that correlate with or give us information about another variable of interest.
โ Tree induction recursively finds informative attributes for subsets of the data.( ๋ฐ์ดํฐ์ ํ์ ์งํฉ์ ๋ํด ์ ๋ณด๋ฅผ ์ ๊ณตํ๋ ์์ฑ์ ์ฌ๊ท์ ์ผ๋ก ์ฐพ์ต๋๋ค)
โ The partitioning is โsupervisedโ in that it tries to find segments that give increasingly precise information about the quantity to be predicted, the target.
โ The resulting tree-structured model partitions the space of all possible instances into a set of segments with different predicted values for the target.(๊ฒฐ๊ณผ์ ์ผ๋ก ํธ๋ฆฌ ๊ตฌ์กฐ์ ๋ชจ๋ธ์ ๋ชจ๋ ๊ฐ๋ฅํ ์ธ์คํด์ค ๊ณต๊ฐ์ ์ฌ๋ฌ ์ธ๊ทธ๋จผํธ๋ก ๋๋๋ฉฐ, ๊ฐ ์ธ๊ทธ๋จผํธ๋ ๋ชฉํ์ ๋ํ ์๋ก ๋ค๋ฅธ ์์ธก ๊ฐ์ ๊ฐ์ง๋๋ค.)
ex) ๊ณ ๊ฐ ์ดํ ์์ธก์ ํ๋ค๊ณ ๊ฐ์ ํด๋ณด์.
์ฒซ๋ฒ์งธ ๋ถ๊ธฐ๋ ์๋์ผ๋ก ๋๋ฒ์งธ ๋ถ๊ธฐ๋ ํตํ์๊ฐ, ์ด๋ ๊ฒ ํ๋ฉด 4๊ฐ์ง์ ์ธ๊ทธ๋จผํธ๊ฐ ๋ง๋ค์ด์ง๋ค.
์ด๋ ์ธ๊ทธ๋จผํธ๋ณ๋ก ๋ค๋ฅธ ์ดํํ๋ฅ ์ ๊ฐ์ง๊ฒ ๋๋ค.
๋ฐ๋ผ์"๋ชจ๋ ๊ฐ๋ฅํ ์ธ์คํด์ค ๊ณต๊ฐ์ ๋๋๋ค"๋ ๋ง์, ๋ชจ๋ ๊ณ ๊ฐ ๋ฐ์ดํฐ๋ฅผ ์กฐ๊ฑด์ ๋ฐ๋ผ ๊ทธ๋ฃน์ผ๋ก ๋๋๊ณ , ๊ทธ๋ฃน๋ณ๋ก ์ดํ ํ๋ฅ ๋ฑ์ ์์ธก ๊ฐ์ ์ค์ ํ๋ ๊ฒ์ ์๋ฏธํฉ๋๋ค