논문 리뷰 - Attention Is All You Need

svenskpotatis·2024년 8월 16일

Abstract, introduction

1. 기존 모델의 한계

  • 주로 사용되는 sequence transduction 모델들은 Encoder와 Decoder를 포함하는 복잡한 Recurrent 혹은 Convolutional Neural Networks 기반

2. 새로 제안하는 모델 new simple network architecture, the Transformer

  • 전적으로 Attention 메커니즘에 기반하고, Recurrence와 Convolution을 완전히 제거

  • 성능

    • parallelizable(병렬) -> 학습 시간 단축
    • WMT 2014 English-to-German translation task: 28.4 BLEU (ensemble 보다도 우위)
    • WMT 2014 English-to-French translation task: 41.4 BLEU (single model SOTA)

3. Sequential computation 의 단점

  • RNN, LSTM, gated RNN
  • sequential nature: 병렬 불가
  • memory constraints

4. Recent works achieved

  • Computional efficiency (계산 효율성: factorization tricks, conditional computation)
  • Improving model performance

-> 그러나, Sequential computation의 근본적인 제약 여전히 존재

5. Transformer

  • Recurrence를 배제하고 전적으로 Attention 메커니즘에 의존하는 모델 아키텍처
  • draw global depencies between input and output(입력과 출력 간의 전역적인 종속성을 도출)
  • more parallelization, train faster, reach SOTA

Model Architecture

  • Encoder-Decoder 구조
  • Encoder: 입력 시퀀스의 symbol 표현을 연속적인 표현(z)의 시퀀스로 mapping
  • Decoder: z -> 한 번에 하나씩 출력 시퀀스의 symbol 생성
  • each step: auto-regressive 방식으로, 이전에 생성된 symbol들을 추가적인 입력으로 사용
  • encoder, decoder both: stacked Self-Attention 및 point-wise, fully connected layers로 구성

Encoder and Decoder Stacks

Encoder

  • stack of N=6 identical layers

  • each layer has 2 sub-layers

    1. multi-head attention
    2. position-wise fully connected feed-forward network
  • 각 sub-layer 주위에 residual connection을 추가하여 LayerNorm(x + Sublayer(x)) 형식으로 layer normalization 적용

  • 출력 차원: 512 (embedding layer, sub-layers 모두)

Decoder

  • encoder와 동일, N=6 identical layers

  • 3 sub-layers

    • encoder 와 동일한 sublayer에 하나의 sublayer 추가: Encoder stack의 출력에 대해 Multi-Head Attention 수행

Attention

  • Query와 Key-Value pair 집합을 입력으로 받아 출력(벡터)을 생성

  • Output: weighted sum of the values(V)

    • Weight: Query와 해당하는 Key 간의 유사도 함수(compatibility function)로 계산

Scaled Dot-Product Attention

input: Query와 Key는 차원이 dkd_k, Value는 차원이 dvd_v인 벡터들로 구성

Attention(Q,K,V)=softmax(QKTdk)V\text{Attention}(Q, K, V) = \text{softmax}\left(\frac{QK^T}{\sqrt{d_k}}\right) V
  • QQ: Query 행렬
  • KK: Key 행렬
  • VV: Value 행렬
  • dkd_k: Key의 차원
  • Query와 Key의 내적 수행 -> 그 값을 Key의 차원의 제곱근으로 나눔 -> softmax 함수 적용하여 Attention Weight 계산 -> Weight를 Value 행렬에 곱해 최종 Attention 출력 생성

cf. Attention functions

  • Additive Attention

    • Query와 Key를 합하여 비선형 변환을 거친 후, softmax를 적용하여 Attention Weight 계산
    • 계산량 많음
  • Dot-Product (Multiplicative) Attention

    • Query와 Key의 내적 통해 유사도를 계산
    • 비교적 계산이 간단하지만, 차원이 큰 경우 값이 커져 학습이 어려워질 수 있음
  • Scaled Dot-Product Attention

    • Dot-Product Attention에 scaling 추가
    • 차원이 큰 경우 발생할 수 있는 문제 해결

Multi-Head Attention

  • 여러 개(h개)의 attention head 병렬로 사용
  • 각 head 에서 독립적으로 Q, W, V 계산(scaled dot-product attention)
  • 병렬 수행 결과 concatenate -> linear -> final values
  • 수행 이유: 다양한 시점에서의 정보를 보다 풍부하게 학습하기 위해
MultiHead(Q,K,V)=Concat(head1,head2,…,headh)WO\text{MultiHead}(Q, K, V) = \text{Concat}(\text{head}_1, \text{head}_2, \dots, \text{head}_h)W^O

각 Attention head:

headi=Attention(QWiQ,KWiK,VWiV)\text{head}_i = \text{Attention}(QW_i^Q, KW_i^K, VW_i^V)
  • QQ, KK, VV: Query, Key, Value 행렬
  • WiQW_i^Q, WiKW_i^K, WiVW_i^V: Query, Key, Value 행렬에 각각 적용되는 학습 가능한 가중치 행렬
  • WOW^O: 각 헤드의 출력을 결합한 후 적용되는 출력 가중치 행렬
  • headi\text{head}_i: ii번째 Attention 헤드의 출력
  • Concat\text{Concat}: 여러 Attention 헤드 결합
  • hh: Attention 헤드의 수

Multi-head

  • Query, Key, Value 행렬을 각각의 Attention 헤드에 대해 다른 가중치 행렬(WiQW_i^Q, WiKW_i^K, WiVW_i^V)로 변환하여, 병렬로 여러 Attention 연산 수행

Scaled Dot-Product Attention

  • 각 헤드에서 변환된 Query, Key, Value를 사용하여 Scaled Dot-Product Attention을 계산
  • 이때 각 헤드는 다른 가중치 행렬을 사용하므로, 서로 다른 특성을 학습할 수 있음

Concatenate

  • 여러 Attention 헤드의 출력 concatenate 하나의 큰 벡터로 만듭니다. 이는 서로 다른 정보가 함께 합쳐진 상태를 의미합니다.

Linear layer

  • 최종 출력 생성

Applications of Attention in our Model

Transfor가 multi-head attention을 사용하는 세 가지 방법

Encoder-decoder attention layer:

  • Encoder와 Decoder 사이의 정보 연결
  • Decoder의 Query가 Encoder에서 생성된 Key, Value와 상호작용하여 Decoder가 입력 시퀀스에서 중요한 정보에 집중할 수 있게 함
  • 이를 통해 번역 작업이나 sequence-to-sequence 문제에서 이전에 입력된 단어들과 관련된 정보를 잘 포착하는 데에 도움 줌

Self attention layer in encoder:

  • Encoder 내부에서 Self-Attention 사용
  • 입력 시퀀스 내에서 각 단어가 다른 단어들과 어떻게 연관되어 있는지를 계산하여 문맥 정보 포착

Self attention layer in decoder:

  • 이전에 생성된 단어들 간의 관계를 학습하고, 이를 바탕으로 다음 단어 예측
  • Masked Self-Attention을 사용하여, 현재 단어 이후의 정보는 참조하지 않음

Position-wise Feed-Forward Networks

  • 각 위치에서 독립적으로, 동일한 Fully Connected Layer 적용
  • 두 개의 Linear Layer와 ReLU 활성화 함수
  • 입력 벡터 -> linear layer -> ReLU(비선형성 추가) -> linear layer -> 출력 벡터
  • 각 위치에서 병렬적으로 수행
FFN(x)=max(0,xW1+b1)W2+b2\text{FFN}(x) = \text{max}(0, xW_1 + b_1)W_2 + b_2

Embeddings and Softmax

  • 단어의 정수 인덱스를 실수 벡터로 매핑
  • 출력층: softmax 함수 적용하여 각 단어가 다음에 나올 확률 계산

Positional Encoding

  • Transformer는 순서를 고려하지 않는 Self-Attention 메커니즘을 사용하기 때문에, positional encoding 통해 입력 시퀀스의 위치 정보를 추가적으로 인코딩
  • sin, cos
  • 학습x, 고정된 방식으로 계산
  • 각 입력 벡터에 더해져 위치정보 반영

Why Self-Attention

  • 각 요소가 다른 모든 요소와 상호작용: global dependency
  • sequential computation 한계 극복
  • 병렬 -> 학습 속도 향상
  • 긴 의존성 효과적으로 처리: 긴 sequence에서 우수한 성능

Training

  • optimizer: adam
  • regularization(과적합 방지): dropout, label smoothing

Results

  • model variations: layer 수, attention head 수, hidden layer 크기 변경 통해 최적화 가능
  • 영어 구문 분석(constisuency parsing)에서 우수한 성능

references
스케일드 닷-프로덕트 어텐션(Scaled dot-product Attention)

0개의 댓글