3. Introduction to Predictive Modeling

Leejaegunยท2024๋…„ 9์›” 30์ผ
post-thumbnail

0. Intro

  • Predictive modeling as supervised segmentation

๐Ÿ‘‰ A key to supervised data mining is that we have some target quantity we would like to predict or to otherwise understand better

Fundamental concepts: Identifying informative attributes of the entities described by the data.

  • We would like to find knowable attributes that correlate with the target of interest(ํƒ€์ผ“๋ณ€์ˆ˜์™€ ๊ฐ€์žฅ ๊ด€๋ จ์ด ์žˆ๋Š” ์†์„ฑ ์ฐพ๊ธฐ)
  • Information is a quantity that reduces uncertainty about something
  • Having a target variable crystalizes our notion of finding informative attributes.(๋ชฉํ‘œ๋ณ€์ˆ˜๋ฅผ ๊ฐ–๊ณ  ์žˆ์œผ๋ฉด, ์œ ์šฉํ•œ ์†์„ฑ์„ ์ฐพ๋Š” ๊ฐœ๋…์ด ๊ตฌ์ฒดํ™”๋จ)

1. Models, Induction and Prediction

A model is a simplified representation of reality created to serve a purpose.

A predictive model may be judged solely on its predictive performance

Supervised learning(์ง€๋„ํ•™์Šต) is model creation where the model describes a relationship between a set of selected variables (attributes or features) and a predefined variable called the target variable.

What is "instance"?

An instance or example represents a fact or a data point.โ€ข An instance is described by a set of attributes(fields, columns, variables, or features.)โ€ข An instance is also called a feature vector,because it can be represented as a fixed-length ordered collection (vector) of feature values

Terminology

โ€ข Model induction: The creation of models from data is known as model induction.(๋ชจ๋ธ์œ ๋„)

โ€ข Induction algorithm or learner: The procedure that creates the model from the data is called the induction algorithm or learner.

โ€ข Induction: Induction refers to generalizing from specific cases to general rules (๊ตฌ์ฒด์  ์‚ฌ์‹ค์˜ ์ผ๋ฐ˜ํ™”, ๊ท€๋‚ฉ๋ฒ•)

โ€ข Deduction: Deduction starts with general rules and specific facts, and creates other specific facts from them (์ผ๋ฐ˜์  ์‚ฌ์‹ค์˜ ๊ตฌ์ฒดํ™”, ์—ฐ์—ญ๋ฒ•)

โ€ขTraining data: The input data for the induction algorithm, used for inducing the model, are called the training data.

โ€ข Labeled data: The training data are called labeled data because the value for the target variable (the label) is known.

2. Supervised Segmentation

A predictive model focuses on estimating the value of some particular target variable of interest

๐Ÿค” How can we judge whether a variable contains important information about the target variable?

๐Ÿ‘‰ Entropy ๋กœ ์ฐพ์ž.

3.Entropy

3.1 what is Entropy?

Entropy is a measure of disorder(๋ฌด์งˆ์„œ ์ •๋„) that can beapplied to a set, such as one of our individual segments

-> ์—”ํŠธ๋กœํ”ผ๊ฐ€ ๋†’์„์ˆ˜๋ก, ๋ฌด์งˆ์„œ๊ฐ€ ๋†’์€ ๊ฒƒ์ด๋‹ค.

3.2 ์ˆ˜์‹

H(X)=โˆ’โˆ‘xโˆˆXp(x)logbย p(x)H(X) = - \displaystyle \sum_{x \in X}p(x) log_{b}\ p(x)

  • H(X)H(X)๋Š” ๋žœ๋ค ๋ณ€์ˆ˜ XX์˜ ์—”ํŠธ๋กœํ”ผ
  • p(x)p(x)๋Š” XX๊ฐ€ ๊ฐ’ xx๋ฅผ ๊ฐ€์งˆ ํ™•๋ฅ 
  • bb๋Š” ๋กœ๊ทธ์˜ ๋ฐ‘์œผ๋กœ, ๋ณดํ†ต 2๋ฅผ ์‚ฌ์šฉํ•˜๋ฉฐ ์ด ๊ฒฝ์šฐ ์—”ํŠธ๋กœํ”ผ์˜ ๋‹จ์œ„๋Š” ๋น„ํŠธ(bits)๊ฐ€ ๋ฉ๋‹ˆ๋‹ค.

3.3 Information Gain

Entropy ,only tells us how impure one individual subset is ๋ฐ˜๋ฉด์—

Information gain(IG) measures how much an attribute improves (decreases) entropy over the whole segmentation it creates.

  • IG measures the change in entropy due to any amount of new information being added (์ƒˆ๋กœ์šด ์ •๋ณด์˜ ์ถ”๊ฐ€์— ๋”ฐ๋ฅธ ์—”ํŠธ๋กœํ”ผ์˜ ๋ณ€ํ™”๋Ÿ‰)

  • IG is a function of both a parent set and the children resulting from some partitioning of the parent set โ€“ how much information has this attribute provided?

3.3.2 IG_Example

So this split reduces entropy substantially. In predictive modeling terms, the attribute provides a lot of information on the value of the target.

๐Ÿค” ๋‹ค๋ฅด๊ฒŒ ๋‚˜๋ˆŒ ์ˆ˜ ์žˆ์ง€ ์•Š์€๊ฐ€?
๐Ÿ‘‰ ๋‚˜๋ˆŒ ์ˆ˜๋Š” ์žˆ๋‹ค. ํ•˜์ง€๋งŒ ์ •๋ณด์ด๋“์ด ์ ๋‹ค.์ •๋ณด์ด๋“์ด ์ตœ๋Œ€๊ฐ€ ๋˜๋Š” ์†์„ฑ(attributes)์ฐพ์•„์„œ split ํ•˜๋Š”๊ฒŒ ๋ชฉํ‘œ์ธ ๊ฒƒ์ด๋‹ค.

๋”ฐ๋ผ์„œ
For a dataset with instances described by attributes and a target variable, we can determine which attribute is the most informative with respect to estimating the value of the target variable

๐Ÿš€ We also can rank a set of attributes by their informativeness, in particular by their information gain

3.4 Mushroom Example

3.4.1 Problem situation

  • This is a classification problem because we have a target variable, called edible?, with two values yes (edible) and no (poisonous), specifying our two classes
  • We will use information gain to answer the question:
    โ€œWhat single attribute is the most useful for distinguishing edible
    (edible?=Yes) mushrooms from poisonous (edible?=No) ones?โ€

3.4.2 Solved

-> IG ๊ฐ€ ๊ฐ€์žฅ ๋†’์€ attribute๋ฅผ ๊ณจ๋ผ์„œ, ๊ฐ€์žฅ ์ข‹์€ ์†์„ฑ์„ ์„ ํƒํ•˜์—ฌ ์ด ์†์„ฑ์„ ๊ธฐ์ค€์œผ๋กœ edible ํ•œ์ง€ ์•ˆํ•œ์ง€ ๋‚˜๋ˆ„์–ด๋ณธ๋‹ค.

๐Ÿ‘‰์ด ๊ทธ๋ž˜ํ”„๋“ค์€ ๊ฐ ์†์„ฑ ๊ฐ’์— ๋”ฐ๋ผ ๋ฐ์ดํ„ฐ์˜ ์˜ˆ์ธก ๊ฐ€๋Šฅ์„ฑ์— ๊ธฐ์—ฌํ•˜๋Š” ์ •๋„๋ฅผ ์—”ํŠธ๋กœํ”ผ๋ฅผ ํ†ตํ•ด ์‹œ๊ฐํ™”ํ•˜๊ณ  ์žˆ์Šต๋‹ˆ๋‹ค. Odor ์†์„ฑ์ฒ˜๋Ÿผ ๋น„์œจ์ด ๋†’๊ณ  ์—”ํŠธ๋กœํ”ผ๊ฐ€ ๋‚ฎ์€ ์†์„ฑ์€ ๋ฐ์ดํ„ฐ์˜ ๋ถ„๋ฅ˜์— ๋” ๋„์›€์ด ๋  ์ˆ˜ ์žˆ์œผ๋ฉฐ, ๋ฐ˜๋Œ€๋กœ ์—”ํŠธ๋กœํ”ผ๊ฐ€ ๋†’๊ณ  ๋‹ค์–‘ํ•œ ๊ฐ’์ด ๋ถ„ํฌํ•œ ์†์„ฑ์€ ์˜ˆ์ธก์— ์žˆ์–ด ์ •๋ณด์˜ ๋ถ„์‚ฐ๋„๊ฐ€ ํฌ๋‹ค๋Š” ์ ์„ ์•Œ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

3.5 Supervisied segmentation with Tree-Structured Models

  • The tree is made up of nodes,

  • Each interior node in the tree contains a test of an attribute

  • each path eventually terminates at a terminal node, or leaf.

Classificiation Tree

  • The tree is a supervised segmentation(์ง€๋„ ์„ธ๋ถ„ํ™”),because each leaf contains a value for the target variable.

  • such a tree is called a classification tree or more loosely a decision tree.

3.6 Visualizing Segmentations


-> x์ถ• y์ถ•์œผ๋กœ ๊ฐ ๋ถ€๋ถ„์ด ์–ด๋–ป๊ฒŒ ๋‚˜๋‰˜์—ˆ๋Š”์ง€ ๊ฐ„๋‹จํ•˜๊ฒŒ ์‚ดํŽด๋ณผ ์ˆ˜ ์žˆ๋Š” ์žฅ์ ์ด ์žˆ๋‹ค.

3.7 Trees as Sets of Rules

  • Every classification tree can be expressed as a set of rules.

3.8 Probability Estimation

  • In the context of supervised segmentation, we would like to each segment to be assigned an estimate of the probability of membership in the different classes.

4. Frequency-based estimation

If we assign the same class probability to every member of the segment corresponding to a tree leaf, we can use instance counts at each leaf to compute a class probability estimate.

A frequency-based estimate of class membership probability may be overly optimistic about the probability of class membership for segments with very small numbers of instance. (์ž‘์€ ์ƒ˜ํ”Œ์— ๋Œ€ํ•ด์„œ๋Š” ๊ณผ๋„ํ•˜๊ฒŒ ๋‚™์ฒœ์ ์ผ ์ˆ˜๋„ ์žˆ์Œ)
<- ์ด๋ฅผ ํ•ด๊ฒฐํ•˜๊ธฐ ์œ„ํ•ด Laplace ๋ณด์ •์ด ๋‚˜ํƒ€๋‚˜๊ฒŒ ํ•จ.

5. Laplace Correction

Instead of simply computing the frequency, we use a โ€œsmoothing(ํ‰ํ™œํ™”)โ€ version of the frequency-based estimate (Laplace correction); the purpose of which is to moderate the influence of leaves with only a few instance

Laplace ๋ณด์ •์€ ์ฃผ๋กœ ๋‚˜์ด๋ธŒ ๋ฒ ์ด์ฆˆ(Naive Bayes) ๋ถ„๋ฅ˜๊ธฐ์—์„œ ์‚ฌ์šฉ๋˜๋Š” ๊ธฐ๋ฒ•์œผ๋กœ, ํ›ˆ๋ จ ๋ฐ์ดํ„ฐ์—์„œ ๊ด€์ฐฐ๋˜์ง€ ์•Š์€ ์‚ฌ๊ฑด์— ๋Œ€ํ•ด 0์ด ์•„๋‹Œ ํ™•๋ฅ ์„ ํ• ๋‹นํ•˜๋Š” ๋ฐฉ๋ฒ•์ž„.


Laplace ๋ณด์ •์˜ ์ž‘๋™ ๋ฐฉ์‹

๊ธฐ๋ณธ ์•„์ด๋””์–ด: ๋ชจ๋“  ๊ฐ€๋Šฅํ•œ ๊ฒฐ๊ณผ์— ์ž‘์€ ์ƒ์ˆ˜๊ฐ’(๋ณดํ†ต 1)์„ ๋”ํ•ฉ๋‹ˆ๋‹ค.
์ˆ˜์‹: P(wโˆฃc)=count(w,c)+ฮฑcount(c)+ฮฑKP(w|c) = \frac{count(w,c) + \alpha}{count(c) + \alpha K}

์—ฌ๊ธฐ์„œ:
ww: ๋‹จ์–ด
cc: ํด๋ž˜์Šค
count(w,c)count(w,c): ํด๋ž˜์Šค c์—์„œ ๋‹จ์–ด w์˜ ์ถœํ˜„ ํšŸ์ˆ˜
count(c)count(c): ํด๋ž˜์Šค c์˜ ์ด ๋‹จ์–ด ์ˆ˜
ฮฑ: ์Šค๋ฌด๋”ฉ ํŒŒ๋ผ๋ฏธํ„ฐ (๋ณดํ†ต 1)
KK: ๊ณ ์œ ํ•œ ๋‹จ์–ด์˜ ์ด ๊ฐœ์ˆ˜


5.1 The Effect of Laplace Smoothing

  • As the number of instance increases, the Laplace equation converges to the frequency-based
    estimate.

  • For each ratio the solid horizontal line shows the uncorrected (constant) estimate, while the corresponding dashed line shows the estimate with the Laplace correction applied.

  • The uncorrected line is the asymptote of the Laplace correction as the number of instances goes to infinity

์œ„ ์ด๋ฏธ์ง€ ์„ค๋ช…

๊ธ์ • 100%, 80%, 66% - ์‹ค์„ ์œผ๋กœ Frequency Estimation ๋‚˜ํƒ€๋‚ด๋ฉฐ
๊ทธ ๋ฐ‘์— ์ ์„ ์ด Laplace Estimation ์ด๋‹ค.

instance ์ˆ˜๊ฐ€ ์ ์„ ๋•Œ, Laplace ๋ณด์ •์€ ๊ณผ๋„ํ•œ ํ™•๋ฅ ์„ ์™„ํ™”ํ•˜๋Š” ์—ญํ• ์„ ํ•œ๋‹ค.
๐Ÿ‘‰ ํŠนํžˆ ๋ฐ์ดํ„ฐ์˜ ์ˆ˜๊ฐ€ ๋ถ€์กฑํ•œ ์ƒํ™ฉ์—์„œ ๊ณผ์ ํ•ฉ์„ ๋ฐฉ์ง€.

5.2 Example: Addressing the Churn Problem with Tree Induction


๐Ÿ‘‰ Information Gain์„ ํ†ตํ•ด ์ค‘์š”ํ•œ Attributes๊ฐ€ ๋ฌด์—‡์ธ์ง€ ํŒŒ์•…ํ•  ์ˆ˜ ์žˆ๊ฒŒ ๋˜์—ˆ๋‹ค. House์™€ OVERAGE ํŠน์„ฑ์ด ๊ณ ๊ฐ ํŠน์„ฑ์—์„œ ๊ฐ€์žฅ ์ค‘์š”ํ•œ ์š”์ธ์ž„์„ ์•Œ๋ ค์ค€๋‹ค.


์˜์‚ฌ๊ฒฐ์ •ํŠธ๋ฆฌ๊ฐ€ ์‹œ๊ฐํ™”๋˜์–ด ์žˆ๋‹ค. House ์˜ ๊ฐ’ 600469์— ๋”ฐ๋ผ์„œ ์ฒซ๋ฒˆ์งธ ๋ถ„๊ธฐ๊ฐ€ ๋˜๋ฉฐ, ๊ฐ ๋ถ„๊ธฐ์— ํŠน์„ฑ์— ๋”ฐ๋ผ์„œ ๋‹ค์‹œ ๋‚˜๋ˆŒ ์ˆ˜ ์žˆ๋‹ค.

6. Summary

Predictive modeling:

a model can estimate the value of a target variable for a new unseen example.
์•ˆ ๋ณธ ๋ฐ์ดํ„ฐ์—์„œ๋„ predict ํ•จ

How to find and select informative attributes

โ€“ Use information gain, which is based on purity measure called entropy.
โ€“ Given a large collection of data, we can find those variables that correlate with or give us information about another variable of interest.

Tree induction

โ€“ Tree induction recursively finds informative attributes for subsets of the data.( ๋ฐ์ดํ„ฐ์˜ ํ•˜์œ„ ์ง‘ํ•ฉ์— ๋Œ€ํ•ด ์ •๋ณด๋ฅผ ์ œ๊ณตํ•˜๋Š” ์†์„ฑ์„ ์žฌ๊ท€์ ์œผ๋กœ ์ฐพ์Šต๋‹ˆ๋‹ค)

โ€“ The partitioning is โ€œsupervisedโ€ in that it tries to find segments that give increasingly precise information about the quantity to be predicted, the target.

โ€“ The resulting tree-structured model partitions the space of all possible instances into a set of segments with different predicted values for the target.(๊ฒฐ๊ณผ์ ์œผ๋กœ ํŠธ๋ฆฌ ๊ตฌ์กฐ์˜ ๋ชจ๋ธ์€ ๋ชจ๋“  ๊ฐ€๋Šฅํ•œ ์ธ์Šคํ„ด์Šค ๊ณต๊ฐ„์„ ์—ฌ๋Ÿฌ ์„ธ๊ทธ๋จผํŠธ๋กœ ๋‚˜๋ˆ„๋ฉฐ, ๊ฐ ์„ธ๊ทธ๋จผํŠธ๋Š” ๋ชฉํ‘œ์— ๋Œ€ํ•œ ์„œ๋กœ ๋‹ค๋ฅธ ์˜ˆ์ธก ๊ฐ’์„ ๊ฐ€์ง‘๋‹ˆ๋‹ค.)

ex) ๊ณ ๊ฐ ์ดํƒˆ ์˜ˆ์ธก์„ ํ•œ๋‹ค๊ณ  ๊ฐ€์ •ํ•ด๋ณด์ž.
์ฒซ๋ฒˆ์งธ ๋ถ„๊ธฐ๋Š” ์†Œ๋“์œผ๋กœ ๋‘๋ฒˆ์งธ ๋ถ„๊ธฐ๋Š” ํ†ตํ™”์‹œ๊ฐ„, ์ด๋ ‡๊ฒŒ ํ•˜๋ฉด 4๊ฐ€์ง€์˜ ์„ธ๊ทธ๋จผํŠธ๊ฐ€ ๋งŒ๋“ค์–ด์ง„๋‹ค.
์ด๋Š” ์„ธ๊ทธ๋จผํŠธ๋ณ„๋กœ ๋‹ค๋ฅธ ์ดํƒˆํ™•๋ฅ ์„ ๊ฐ€์ง€๊ฒŒ ๋œ๋‹ค.

๋”ฐ๋ผ์„œ"๋ชจ๋“  ๊ฐ€๋Šฅํ•œ ์ธ์Šคํ„ด์Šค ๊ณต๊ฐ„์„ ๋‚˜๋ˆˆ๋‹ค"๋Š” ๋ง์€, ๋ชจ๋“  ๊ณ ๊ฐ ๋ฐ์ดํ„ฐ๋ฅผ ์กฐ๊ฑด์— ๋”ฐ๋ผ ๊ทธ๋ฃน์œผ๋กœ ๋‚˜๋ˆ„๊ณ , ๊ทธ๋ฃน๋ณ„๋กœ ์ดํƒˆ ํ™•๋ฅ  ๋“ฑ์˜ ์˜ˆ์ธก ๊ฐ’์„ ์„ค์ •ํ•˜๋Š” ๊ฒƒ์„ ์˜๋ฏธํ•ฉ๋‹ˆ๋‹ค

profile
Lee_AA

0๊ฐœ์˜ ๋Œ“๊ธ€