์์ฐ์ด ์ฒ๋ฆฌ(Natural Language Processing) ๋ ์ธ๊ฐ ์ธ์ด๋ฅผ ์ปดํจํฐ๊ฐ ์ดํดํ๊ณ ์ฒ๋ฆฌํ๋๋ก ๋ง๋๋ ๊ธฐ์ ์ด๋ค.
ํ
์คํธ๋ ์์ฑ ๋ฐ์ดํฐ๋ฅผ ์ํ์ ์ผ๋ก ๋ณํํด ์๋ฏธ๋ฅผ ์ถ์ถํ๋ค.
| ๋จ๊ณ | ๋ชฉ์ | ์์ |
|---|---|---|
| ํ ์คํธ ์ ์ฒ๋ฆฌ | ๋ถํ์ํ ๋จ์ด ์ ๊ฑฐ, ์ ๊ทํ | โIโm happy!!!โ โ โi happyโ |
| ๋ฒกํฐํ | ๋ฌธ์ฅ์ ์ซ์๋ก ํํ | Bag-of-Words, TF-IDF |
| ํ์ต | ๊ฐ์ฑ๋ถ์ยทํ ํฝ๋ถ๋ฅ ๋ฑ | ๊ธ์ /๋ถ์ ๋ถ๋ฅ๊ธฐ |
| ์์ฉ | ์ฑ๋ด, ๋ฒ์ญ, ์์ฝ ๋ฑ | ChatGPT, DeepL, BERT |
| ์ ํ | ์ค๋ช | ๋ํ ๊ธฐ๋ฒ |
|---|---|---|
| ํด๋์ ๋ชจ๋ธ | ๊ท์นยทํต๊ณ ๊ธฐ๋ฐ | BoW, TF-IDF, Naive Bayes |
| ๋ฅ๋ฌ๋ ๋ชจ๋ธ | ์๋ฏธยท๋ฌธ๋งฅ ๊ธฐ๋ฐ | RNN, LSTM, Transformer |
๋ฌธ์๋ฅผ ๋จ์ด์ โ๊ฐ๋ฐฉ(bag)โ์ผ๋ก ๋ณด๊ณ ,
๋จ์ด์ ์ถํ ํ์๋ฅผ ๋ฒกํฐ ํํ๋ก ํํํ๋ค.
์์:
๋ฌธ์ฅ1: โI love NLPโ
๋ฌธ์ฅ2: โI love Pythonโ
โ ๋จ์ด ์งํฉ(Vocabulary): [I, love, NLP, Python]
โ BoW ๋ฒกํฐ
๋จ์ํ์ง๋ง, ๋จ์ด ์์๋ ์๋ฏธ๋ฅผ ๊ณ ๋ คํ์ง ์๋๋ค.
import re
import nltk
from nltk.corpus import stopwords
from nltk.stem.porter import PorterStemmer
corpus = []
for i in range(0, 1000):
review = re.sub('[^a-zA-Z]', ' ', dataset['Review'][i])
review = review.lower()
review = review.split()
ps = PorterStemmer()
review = [ps.stem(word) for word in review if not word in set(stopwords.words('english'))]
review = ' '.join(review)
corpus.append(review)
| ๋จ๊ณ | ์ค๋ช |
|---|---|
| ์ ๊ท์ ์ ๊ฑฐ | ํน์๋ฌธ์, ์ซ์ ์ ๊ฑฐ |
| ์๋ฌธ์ ๋ณํ | ๋์๋ฌธ์ ์ผ๊ด์ฑ |
| ํ ํฐํ(Tokenization) | ๋จ์ด ๋จ์ ๋ถ๋ฆฌ |
| ๋ถ์ฉ์ด ์ ๊ฑฐ(Stopwords) | ์๋ฏธ ์๋ ๋จ์ด ์ ๊ฑฐ (โtheโ, โisโ) |
| ์ด๊ฐ ์ถ์ถ(Stemming) | ๋จ์ด์ ๊ธฐ๋ณธํ์ผ๋ก ์ถ์ (โlovedโโโloveโ) |
BoW๊ฐ ๋จ์ด์ โ์กด์ฌโ๋ง ๋ณธ๋ค๋ฉด,
TF-IDF(Term Frequency-Inverse Document Frequency) ๋
๋จ์ด์ ์ค์๋๋ฅผ ๋ฐ์ํ๋ค.
๊ฒฐ๊ณผ์ ์ผ๋ก,
from sklearn.feature_extraction.text import TfidfVectorizer
tfidf = TfidfVectorizer(max_features=1500)
X = tfidf.fit_transform(corpus).toarray()
y = dataset.iloc[:, 1].values # ๊ธ์ (1)/๋ถ์ (0)
โ X๋ 1000ร1500 ์ฐจ์์ ๋ฒกํฐ,
๊ฐ ๋จ์ด์ ์ค์๋(weight)๊ฐ ๋ฐ์๋ ์์น ํ๋ ฌ.
๋ฆฌ๋ทฐ๋ ๋๊ธ์ ๊ธ์ ยท๋ถ์ ๊ฐ์ ์ ๋ถ๋ฅํ๋ค.
์ง๋ ํ์ต ๋ถ๋ฅ ๋ชจ๋ธ(Logistic, Naive Bayes, SVM ๋ฑ)์ ์ฌ์ฉ.
from sklearn.model_selection import train_test_split
from sklearn.naive_bayes import GaussianNB
from sklearn.metrics import confusion_matrix, accuracy_score
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=0)
classifier = GaussianNB()
classifier.fit(X_train, y_train)
y_pred = classifier.predict(X_test)
cm = confusion_matrix(y_test, y_pred)
print("Accuracy:", accuracy_score(y_test, y_pred))
๐ก TF-IDF + Naive Bayes = ๊ฐ๋จํ์ง๋ง ๊ฐ๋ ฅํ ํ
์คํธ ๋ถ๋ฅ ์กฐํฉ.
๋ฆฌ๋ทฐ ๋ฐ์ดํฐ(โ์ข์์โ, โ๋ณ๋ก์์โ)๋ฅผ ํ์ตํด ์ค์๊ฐ ๊ฐ์ ํ๋ณ ๊ฐ๋ฅ.
library(tm)
library(SnowballC)
library(e1071)
corpus = VCorpus(VectorSource(dataset$Review))
corpus = tm_map(corpus, content_transformer(tolower))
corpus = tm_map(corpus, removePunctuation)
corpus = tm_map(corpus, removeNumbers)
corpus = tm_map(corpus, removeWords, stopwords())
corpus = tm_map(corpus, stemDocument)
dtm = DocumentTermMatrix(corpus)
dtm = removeSparseTerms(dtm, 0.99)
dataset_sparse = as.data.frame(as.matrix(dtm))
dataset_sparse$Liked = dataset$Liked
classifier = naiveBayes(Liked ~ ., data = dataset_sparse)
pred = predict(classifier, dataset_sparse)
| ํญ๋ชฉ | Bag-of-Words | TF-IDF |
|---|---|---|
| ํน์ง | ๋จ์ ์ถํ ํ์ | ์ถํ + ์ค์๋ ๋ฐ์ |
| ๊ณ์ฐ๋ | ์์ | ํผ |
| ์ค๋ณต ๋จ์ด ์ํฅ | ํผ | ์์ |
| ์ถ์ฒ ์ฌ์ฉ | ๊ธฐ๋ณธ baseline | ๊ณ ๊ธ ํํ, ๊ฐ์ฑ ๋ถ์์ ์ ํฉ |
| ์งํ | ์๋ฏธ |
|---|---|
| ์ ํ๋(Accuracy) | ์ ์ฒด ์ค ๋ง๊ฒ ์์ธกํ ๋น์จ |
| ์ ๋ฐ๋(Precision) | ๊ธ์ ์ด๋ผ ์์ธกํ ๊ฒ ์ค ์ค์ ๊ธ์ |
| ์ฌํ์จ(Recall) | ์ค์ ๊ธ์ ์ค ๋ง๊ฒ ์์ธกํ ๋น์จ |
| F1 Score | ์ ๋ฐ๋ยท์ฌํ์จ์ ์กฐํ ํ๊ท |
from sklearn.metrics import classification_report
print(classification_report(y_test, y_pred))
| ๋จ๊ณ | ๋ํ ๊ธฐ๋ฒ | ์ค๋ช |
|---|---|---|
| ๋จ์ด ์๋ฒ ๋ฉ | Word2Vec, GloVe | ๋จ์ด ์๋ฏธ ๋ฒกํฐํ |
| ๋ฌธ๋งฅ ๋ชจ๋ธ | BERT, GPT | ๋ฌธ์ฅ ์์ค ์๋ฏธ |
| ๊ฐ์ฑ ๋ถ์ ๊ณ ๋ํ | LSTM, Transformer | ๋ฌธ๋งฅ ์์กด ๊ฐ์ ํ์ |
์์ฐ์ด ์ฒ๋ฆฌ๋ ํ
์คํธ๋ฅผ ์์นํํ์ฌ
๊ฐ์ ยท์๋ยท์๋ฏธ๋ฅผ ๋ฐ์ดํฐ๋ก ํด์ํ๋ ๊ธฐ์ ์ด๋ค.
๐ ํต์ฌ ์์ฝ
| ๋จ๊ณ | ํต์ฌ ๊ธฐ์ | ๋ชฉ์ |
|---|---|---|
| ํ ์คํธ ์ ์ | Tokenization, Stopword ์ ๊ฑฐ | ๋ ธ์ด์ฆ ์ต์ํ |
| ๋ฒกํฐํ | BoW, TF-IDF | ์์น ๋ณํ |
| ๋ชจ๋ธ๋ง | Naive Bayes, SVM | ๊ฐ์ ยท์ฃผ์ ๋ถ๋ฅ |