"The animal didn't cross the street because it was too tired"라는 문장에서 "it"은 무엇을 가리킬까? 인간에게는 쉽지만 알고리즘에서는 알아내기 쉽지 않다.
이 때 self-attention은 "it"을 "animal"과 연관시킬 수 있게 한다.



Score 계산
Query vector를 Key vector와 내적한다. 결과값이 클수록 벡터가 유사하다.
(Key vector의 dimension)의 루트값으로 score를 나눠줌
이 과정은 gradient가 더 안정적이게 해준다.
softmax
score를 normalize하여 합이 1이 되도록 한다.

Value vector를 이 softmax score와 곱해준다.
이는 relevent words를 살리고 나머지는 drown-out 되도록 해준다.
위의 과정에서 나온 weighted value vector들을 모두 합한다.
이것이 self-attention 레이어의 결과(z)다!

위의 과정을 차원을 높여 행렬로 진행하면 다음과 같다.


전체적인 self-attention 메커니즘은 다음과 같이 요약된다.
