# 트랜스포머(Transformer)

[원문](https://www.linkymeadow.com/w/498d84/)

마지막 수정: 2026-10-05 20:20 (KST)

트랜스포머(Transformer)는 어텐션을 이용해 입력 요소들이 서로 어떤 정보를 참고할지 계산하는 신경망 구조입니다. 텍스트에서는 글자·단어 조각 등의 토큰을 벡터로 바꾸고, 셀프어텐션으로 토큰 사이의 관계를 반영하며 위치 정보와 피드포워드 네트워크 등도 함께 사용합니다.

2017년 원논문은 번역을 위한 인코더-디코더 구조를 제안했습니다. 이후 인코더 중심의 BERT, 디코더 중심의 GPT 계열과 이미지 패치를 다루는 Vision Transformer(ViT) 등으로 활용 범위가 넓어졌습니다. 아래 자료는 개념 설명부터 구조별 차이, 논문과 구현 순서로 읽을 수 있습니다.

## 처음 읽기

- [트랜스포머는 어떻게 동작하나요? — Hugging Face](https://huggingface.co/learn/llm-course/ko/chapter1/4) · 한국어

  트랜스포머의 등장 배경, 어텐션과 인코더·디코더의 역할, 사전학습·미세조정의 개념을 설명합니다.
- [Transformer: A Novel Neural Network Architecture for Language Understanding — Google Research](https://research.google/blog/transformer-a-novel-neural-network-architecture-for-language-understanding/) · 영문

  트랜스포머 연구진이 순환 신경망과의 차이, 어텐션의 동작과 번역 모델의 구조를 소개합니다.

## 셀프어텐션과 전체 구조

- [The Illustrated Transformer — Jay Alammar](https://jalammar.github.io/illustrated-transformer/) · 영문

  그림으로 임베딩, Query·Key·Value, 멀티헤드 어텐션, 위치 정보와 인코더·디코더의 계산 흐름을 설명합니다.
- [The Transformer Architecture — Dive into Deep Learning](https://d2l.ai/chapter_attention-mechanisms-and-transformers/transformer.html) · 영문

  잔차 연결·층 정규화·피드포워드 네트워크와 인코더·디코더를 수식·코드로 구성합니다.

## 인코더·디코더·인코더-디코더

- [인코더 모델 — Hugging Face](https://huggingface.co/learn/llm-course/ko/chapter1/5) · 한국어

  입력의 앞뒤 문맥을 함께 살피는 인코더 구조와 문장 분류·개체명 인식 등의 활용을 설명합니다.
- [디코더 모델 — Hugging Face](https://huggingface.co/learn/llm-course/ko/chapter1/6) · 한국어

  이전 토큰을 바탕으로 다음 토큰을 예측하는 디코더 구조와 텍스트 생성의 관계를 설명합니다.
- [시퀀스-투-시퀀스 모델 — Hugging Face](https://huggingface.co/learn/llm-course/ko/chapter1/7) · 한국어

  인코더가 처리한 입력을 디코더가 참조하며 번역·요약 등의 출력을 생성하는 구조를 설명합니다.

## 원논문과 대표 모델

- [Attention Is All You Need — Vaswani 외](https://arxiv.org/abs/1706.03762) · 영문

  2017년 트랜스포머를 제안한 논문으로, 어텐션을 중심으로 한 인코더·디코더 구조와 기계 번역 실험을 다룹니다.
- [BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding — Devlin 외](https://arxiv.org/abs/1810.04805) · 영문

  양방향 문맥을 활용하는 트랜스포머 인코더의 사전학습과 여러 언어 이해 과제로의 미세조정을 제시합니다.
- [Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer — Raffel 외](https://arxiv.org/abs/1910.10683) · 영문

  여러 자연어 처리 과제를 텍스트 입력·출력 형식으로 통일한 T5와 전이학습 방법을 비교합니다.
- [An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale — Dosovitskiy 외](https://arxiv.org/abs/2010.11929) · 영문

  이미지를 패치 단위의 시퀀스로 바꿔 트랜스포머에 입력하는 Vision Transformer(ViT)를 제안합니다.

## 코드로 따라가기

- [The Annotated Transformer — Harvard NLP](https://nlp.seas.harvard.edu/annotated-transformer/) · 영문

  원논문의 주요 구성 요소와 학습·추론 과정을 PyTorch 코드와 해설로 따라가는 자료입니다.
- [언어 이해를 위한 변환기 모델 — TensorFlow](https://www.tensorflow.org/text/tutorials/transformer?hl=ko) · 한국어

  위치 인코딩·마스킹·어텐션부터 번역 모델의 학습·추론까지 구현하는 튜토리얼입니다.
- [Accelerating PyTorch Transformers by replacing nn.Transformer with Nested Tensors and torch.compile() — PyTorch](https://docs.pytorch.org/tutorials/intermediate/transformer_building_blocks.html) · 영문

  어텐션 구성 요소, Nested Tensor와 컴파일 기능을 이용해 트랜스포머 구현을 최적화하는 심화 튜토리얼입니다.

## 강연·영상

- 3Blue1Brown · Transformers, the tech behind LLMs | Deep Learning Chapter 5 — [원본](https://www.youtube.com/watch?v=wjZofJX0v4M) / [(한국어)](https://www.youtube.com/watch?v=g38aoGttLhI)
- 3Blue1Brown · Attention in transformers, step-by-step | Deep Learning Chapter 6 — [원본](https://www.youtube.com/watch?v=eMlx5fFNoYc) / [(한국어)](https://www.youtube.com/watch?v=_Z3rXeJahMs)
- [Jay Alammar · The Narrated Transformer Language Model](https://www.youtube.com/watch?v=-QH8fRhqFHM) · 영문 영상

## 연관 문서

- [AI·에이전트](https://www.linkymeadow.com/w/dd12dc/)
- [한국어 자연어 처리·음성 AI 자료](https://www.linkymeadow.com/w/b575a1/)
- [AI 에이전트와 프롬프트 인젝션 보안](https://www.linkymeadow.com/w/bf45e6/)
