Text ─────────┐
Image ────────┼──→ Multimodal Image Model → Image
Image ────────┤
Conversation ─┘
- GPT Image 2는 공식적으로 text input + image input/output + editing + high-fidelity image input을 지원한다.
- Nano Banana 2는 아예 Gemini 계열의 natively multimodal reasoning model이고 text와 image를 입력받는다.
- Seedream 5.0 Pro 는 ByteDance가 multimodal image creation model이라고 정의한다.
Text
│
Reference Image(s) ───┤
│
▼
Multimodal Conditioning
│
▼
┌────────────────────────┐
│ Efficient Diffusion │
│ Transformer (DiT) │
└───────────┬────────────┘
│
Image Latents
│
▼
High-compression VAE
│
▼
Image
Reference Image = 생성 결과가 참고해야 할 시각적 정보(visual evidence/context/source)
Text = 그 시각적 정보를 어떻게 해석하고 선택·변경·보존·조합할지 지정하는 언어적 instruction
최근에 사용해본 2가지 이미지 모델에 대해 best practices 를 정리해 보았다.
Seedream 4.5 edit
Figure 1 → Base image
Figure 2 → Product reference
Figure 3 → Text reference
"Replace the product in Figure 1
with that in Figure 2.
For the title, copy the text
in Figure 3 to the top..."
Image 1 = base/content reference
Image 2 = style reference
→ Image 2의 style을
Image 1에 적용
→ Image 1의 layout/object는 유지