26F12C

Young-Kyoo Kim·2026년 2월 11일

인프라 및 네트워크 회의록


1. 회의 배경 및 현황

1.1 현재 인프라 상태 (2026년 2월 기준)

지난 방문 이후 변경사항

벤더 확인:

"I think there are no changes in general from September, except for 
the BGP. So they applied BGP for internal network, and they'll do it 
for the external network this or next month."

주요 변경:

  • BGP 프로토콜 도입: 내부 네트워크에 이미 적용
  • 외부 네트워크: 향후 1개월 내 적용 예정
  • OS/Kubernetes: 변경 없음

1.2 인프라 최적화 목표

고객사 질문:

"is this optimal setup, or is there anything that they can, you know?"

핵심 확인사항:
1. 현재 네트워크 구성이 최적인가?
2. 개선 가능한 부분은 무엇인가?
3. 향후 확장 시 고려사항은?


2. 네트워크 아키텍처: Spine-Leaf 토폴로지

2.1 아키텍처 개요

기본 구조

구성요소:

  • Spine Switch: 상위 계층 스위치
  • Leaf Switch: 하위 계층 스위치 (노드 직접 연결)
  • 이중 Spine: North-South + East-West 트래픽 처리

회의 중 확인:

"right now for the bonds, you know, the east, west traffic, 
everything is over Leaf switches only. There's no spine behind 
those leaves... All of these, the north south is so that they 
can talk."

2.2 포트 구성 및 사용률

Spine-Leaf 연결

고객사 현황:

"the port that's connected to spine from leaf for public, 
it's 26 ports. Private, it's 28 ports."

포트 배분:

  • Public 네트워크: Leaf → Spine 26 포트
  • Private 네트워크: Leaf → Spine 28 포트
  • 스위치 규격: 64포트 스위치

벤더 질문:

"if this 64 port switch space, which then they're using half 
for public traffic, and then they have for [private]"

이중 Spine 활용

현재 상태:

"They have double, you know, two spines, which means they're 
only utilizing half of [capacity]"

의미:

  • 2개의 Spine 스위치 보유
  • 각 Spine당 약 50% 활용률
  • 이중화 및 부하 분산

2.3 Oversubscription (과다 구독) 분석

Oversubscription이란?

벤더 설명:

"oversubscribed could mean if they're have a spine switch where 
they are splitting one port into two, and then they're taking both 
parts into different like leaves. So now the leaves have 64, 64 
but the spine only has..."

기술적 정의:

  • Spine 스위치의 1개 포트를 2개로 분할
  • 2개의 서로 다른 Leaf에 연결
  • Leaf는 64포트 용량이지만 Spine은 32포트만 실제 처리

예시:

"let's say the spine has 64 port, 32 ports, and just splitting 
them to 64 and 64 so that means every port is being used by two 
different leaves."

결과:

"the total throughput the spine can provide is is already over 
committed."

현재 상태 확인

벤더 평가:

"as long as things are not oversubscribed... this design will stay. 
Yes, is fine. It's stable enough, right? And there's no over 
subscription or anything."

결론:

  • Oversubscription 없음 확인
  • 안정적인 네트워크 구성
  • 추가 확장 여력 있음

2.4 확장 계획 시 고려사항

Pool 1 추가 (3월)

계획:

  • 54개 노드 추가
  • 기존 54개 → 총 108개 노드

벤더 조언:

"if these are 64 port switches, there'll be plenty. You can stay 
on, leave only, and connect on all the nodes. But if you're 
planning to add another 54 nodes, you're definitely gonna need 
spines."

필요 조치:

  • 현재: Leaf만으로 충분
  • 108개 노드: Spine 필수
  • 네트워크 확장 계획 수립

2.5 Fabric 구성

Single Fabric vs Dual Fabric

벤더 질문:

"if you actually have two separate fabrics, fabric for compute 
and fabric for internal storage. That was my curiosity. But if 
everything is connected on top of the same... leaf switches are 
created to the same spine switches, then it's just one fabric."

현재 구성:

  • Single Fabric: 하나의 통합 네트워크
  • 논리적 분리: Public/Private 파티셔닝
  • 물리적 공유: 동일한 Spine-Leaf 인프라

장점:

  • 관리 단순화
  • 비용 효율적
  • 리소스 공유

3. BGP (Border Gateway Protocol) 도입

3.1 BGP 적용 현황

도입 시기 및 범위

고객사 설명:

"we updated the network Kubernetes so we BGP for the gateway protocol 
using that part so that main point is exposes the path IP is a switch 
level. So faster than 3% before then before architecture."

적용 상태:

  • 내부 네트워크: 이미 적용 완료
  • 외부 네트워크: Bond 0 스케줄링 중 (이번 달 완료 예정)
  • 성능 개선: 기존 대비 3% 향상

3.2 BGP의 작동 원리

경로 노출 (Path Exposure)

특징:

"main point is exposes the path IP is a switch level"

메커니즘:

  • MinIO 노드 IP를 스위치 레벨에서 노출
  • 클라이언트가 직접 노드 선택 가능
  • Load Balancer 우회

라우팅 최적화

벤더 설명:

"BGP might be blind to the latency that the service... every 
node has [but] it will be continually checking... every 
application will be like, [if a] node is slow, I'm not gonna 
touch it"

특징:

  • 인프라 레벨 라우팅
  • 지연시간 인지 불가 (Latency-blind)
  • 네트워크 계층 부하 분산

4. Sidekick vs BGP 비교

4.1 Sidekick 개념

정의 및 목적

벤더 설명:

"sidekick, we wrote to supplement spark. We work for Spark. 
Because spark can only take one input, it cannot take a 
collection of inputs."

개발 배경:

  • Spark의 단일 입력 제약 해결
  • 클라이언트 측 로드 밸런싱
  • 지연시간 최적화

작동 방식

실시간 헬스체크:

"Sidekick will be continually checking... So every application 
will be like, [this] node is slow. I'm not gonna touch it right. 
So, but this slow, this node 54 is very fast, so I'll push through 
that node"

특징:
1. 클라이언트 측 로직: 각 애플리케이션이 판단
2. 실시간 모니터링: 노드 성능 지속 체크
3. 동적 라우팅: 빠른 노드 우선 사용

4.2 BGP vs Sidekick 비교표

항목BGPSidekick
레벨인프라/네트워크애플리케이션/클라이언트
지연시간 인지없음 (Latency-blind)있음 (실시간 체크)
의사결정 주체네트워크 장비각 클라이언트
설정 변경중앙집중식분산형 (각 클라이언트)
확장 시간편 (클라이언트 재설정 불필요)복잡 (모든 클라이언트 재설정)
단일 장애점Load Balancer 제거없음
관리 복잡도낮음높음

4.3 Tesla 사례: Sidekick의 한계

대규모 환경 도전

벤더 경험:

"Tesla tried to go all the way. Tesla is [using] sidekick... 
they try to send sidekick to the car. So the idea is Tesla is 
uploading video, and they wanted to have the Tesla car see 
1000 servers and then decide which one to go."

문제점:

"The problem was propagating configuration... if they go from 
1200 nodes to 2400 nodes, [it's like] How do we propagate the 
information?"

설정 전파 문제:
1. 1,200 노드 → 2,400 노드 확장
2. 모든 클라이언트(Tesla 차량)에 새 설정 전파 필요
3. 설정 배포 복잡도 급증

벤더 조언:

"Single IP address will always, like, hide a lot of complexity 
for application. But if you [know] right now, we know it's 100 
nodes, 54 nodes, we're going to make [108] nodes, and we're 
going to put that configuration."

4.4 권장 사항

BGP 사용 시나리오

적합한 경우:
1. 대규모 환경: 노드 수 많음
2. 인프라 안정적: 네트워크 장비 신뢰 가능
3. 중앙 관리 선호: 단일 지점 설정 변경
4. 확장 빈번: 노드 추가/제거 자주 발생

벤더 평가:

"BGP [is] very fast as well, but of course, you have a component 
that's taking care of the routing, right?"

Sidekick 사용 시나리오

적합한 경우:
1. 지연시간 민감: Latency-critical 애플리케이션
2. 소규모 환경: 노드 수 제한적
3. 예측 가능한 확장: 노드 변경 드뭄
4. 클라이언트 통제 가능: 설정 전파 가능

벤더 추천:

"[if you have] spark... sidekick [is good]. Every Spark job, 
every driver, every worker, will be load balancing... it will 
[choose] whoever has the best latency response"

4.5 고객사 질문 및 답변

애플리케이션별 차별화

질문:

"they have Trino and spark, right? So, like, and now that they 
already have BGP. Like, maybe, is it better if, like, [Trino 
uses] BGP and [Spark uses] sidekick?"

벤더 답변:

"Just experimentation, will just show them the results. I mean, 
BGP [is] very fast as well... every application, every client, 
is making the decision on how to route on their own, right?"

실험 권장:

  • 두 방식 모두 테스트
  • 성능 벤치마크 비교
  • 애플리케이션 특성 고려

현재 권장사항

BGP 이미 구축 완료:

"BGP will be there. You still recommend them to use sidecar? 
No, you don't need sidekick. Yeah. They have BGP, one or the 
other, yeah"

결론:

  • BGP 채택 권장: 이미 구축했으므로
  • Sidekick 불필요: 중복 투자 회피
  • 필요 시 병행: 특정 애플리케이션만 Sidekick

5. 하드웨어 사양

5.1 현재 노드 스펙 (Pool 0)

스토리지

드라이브 구성:

  • 개수: 노드당 20개 NVMe 드라이브
  • 용량: 드라이브당 7TB (7 terabyte)
  • 총 용량: 54 노드 × 20 드라이브 × 7TB = 7.56PB (raw)

벤더 언급:

"they can't even get it in the market... Lenovo and Cisco were 
hoarding all the eight and 16 terabyte... eight and 15 terabyte 
drives"

시장 상황:

  • 7.5TB, 15TB NVMe 공급 부족
  • Cisco, Lenovo가 재고 확보
  • 고객사들이 해당 벤더 선택

네트워크

현재:

  • 대역폭: 40 Gbps
  • 초당 전송: 약 6~7 GB/s (클러스터 전체)

벤더 평가:

"the average is consistent with the 40 gigabit... I've seen 
[in 100Gbps environments] the size like 80, 90 gigs a second."

비교:

  • 40 Gbps → 6~7 GB/s
  • 100 Gbps → 80~90 GB/s
  • 네트워크 대역폭에 비례

5.2 AI Factory (용인) 계획 스펙

Unit Cluster 구성

계획:

"for unit cluster... they will have 50 nodes, 30 terabyte NVMe, 
24 [drives per node]. But the realistic issue is, there's no 30 
terabyte NVMe. They're working on it."

사양:

  • 노드: 50개
  • 드라이브: 노드당 24개
  • 용량: 드라이브당 30TB (목표)
  • 총 용량: 50 × 24 × 30TB = 36PB (raw)
  • 실제 용량: 2.5~3PB (usable, EC 8:3)

현실적 문제:

  • 30TB NVMe 아직 시장에 없음
  • 대안 검토 필요
  • 예산 제약

네트워크 (AI Factory)

GPU 노드:

"VT content has eight GPUs... Each of these GPUs is connected 
into its own [NIC]... these ones are 400 [Gbps]"

사양:

  • GPU: 노드당 8개
  • NIC: GPU당 전용 NIC
  • 대역폭: 400 Gbps (ConnectX 인코더)
  • 백엔드: 400G InfiniBand 확정

스토리지 네트워크:

"if you have two 400 [Gbps NICs] per server with very few 
[MinIO] nodes, you can satisfy [requirements]"

권장:

  • 서버당 2 × 400 Gbps NIC
  • 소수의 MinIO 노드로 충분

5.3 드라이브 권장 사항

PCIe Generation

중요 경고:

"PCI Express, Gen four or better... Gen fives, you get smaller 
than this, you're gonna go to Gen threes, and then they really 
they just suck... you're gonna be running at half speed."

권장:

  • Gen 4 이상: 필수
  • Gen 5: 최적
  • Gen 3: 피해야 함 (50% 성능 손실)

벤더 조언:

"Don't go dumpster diving. Don't look at the older stuff, 
because you're just not gonna get any performance out of it."

용량 옵션

벤더 제안:

"Depending on the size of the [NVMe] that you pick. If you 
stick to seven terabyte volumes, you're going to need more 
servers. But if you go to 15 terabyte volumes, 31 terabyte 
volumes, even less nodes."

전략:

  • 7TB: 더 많은 노드 필요
  • 15TB: 노드 수 감소
  • 30TB: 최소 노드 (아직 미출시)

고려사항:

  • 노드 수 vs 드라이브 용량 트레이드오프
  • 예산 및 공급 가능성
  • 네트워크 처리량 영향

6. 용량 및 확장성

6.1 파일 시스템 한계 (XFS)

Inode 제약

벤더 설명:

"So you'll run out of inode before you run up. You'll run out 
of inode before you'll run out of objects... The Linux File 
System is a bigger limitation."

기술적 배경:

"when you format XFS, one quarter is reserved for inodes... 
If you're writing a lot of small files, you'll only see 60% 
capacity available."

XFS 특성:

  • 포맷 시 25% inode용 예약
  • 실제로는 동적 할당
  • 작은 파일 많으면 용량 60%만 사용 가능

객체 수 계산

Pool 0 예시:

  • Raw 용량: 7.56PB
  • EC 8:3: 63% 효율
  • Usable: ~5PB
  • 최대 객체: 약 94~100억 개

벤더 계산:

"It's about 10 million objects in the cluster, give or take, 
and it's [based on] the limitation... based on the size of 
the drives. So the more the larger the drives are, the more 
inodes you have."

공식:

  • 1TB = 10M 파일 (inode 기준)
  • Pool 0: 7,560TB × 10M/TB ≈ 75.6B 파일 (raw)
  • EC 적용: ~50B 파일 (usable)

6.2 Pool 추가 시 효과

용량 및 객체 수

벤더 설명:

"In their case, as soon as they have another Pool, they get 
another 4.5 billion [objects]."

Pool 1 추가 (3월):

  • 추가 용량: ~5PB
  • 추가 객체: ~50억 개
  • 총 객체: ~100억 개

6.3 확장 임계값

용량 기준

권장:

"80% [disk usage is the] expansion trigger"

임계값:

  • 70%: Warning
  • 80%: 확장 시작
  • 85%: Critical

파일 수 기준

  • 작은 파일 위주: 60% 용량 도달 시 inode 고갈 가능
  • 큰 파일 위주: 80% 용량까지 여유

7. 네트워크 성능 최적화

7.1 Erasure Coding과 네트워크

EC 8:3의 네트워크 영향

벤더 설명:

"when you do [EC 8:3], let's assume initial set size equals four... 
you'll only be reading four drives. So at the same time, either 
these four nodes or these four nodes will be [reading] at the same 
rate."

특징:

  • 읽기: 8개 중 5개 드라이브만 필요
  • 분산: 최대 54개 노드에 분산 가능
  • 병렬: 여러 노드 동시 읽기

노드 분산의 이점

벤더 설명:

"So in MinIO, every node can take any request... out of 50 nodes, 
54 nodes, five or eight can attend any request for the same file."

메커니즘:
1. 요청 수신: 54개 노드 중 아무 노드
2. 데이터 읽기: 8개 노드 중 5개에서
3. 네트워크 기여: 모든 노드의 NIC 활용

효과:

"the bandwidth their network cards will contribute... they'll 
say, Okay, I'll receive your request and I'll start reading."

7.2 핫스팟 (Hot Spot) 처리

문제 상황

고객사 우려:

"if they're going to, like, read a very specific data... that 
data set over and over and over, like, that'll hit the same pool, 
like, same location of the disk, like, over and over and over."

시나리오:

  • Team A가 특정 데이터셋 반복 읽기
  • 동일 Pool, 동일 드라이브 집중 액세스
  • 성능 저하 우려

실제 동작

벤더 설명:

"[Hot spot] only happens if it's the same file... as soon as 
it's two objects, those are most likely going to be [in] 
different erasure sets."

핵심:

  • 동일 파일: 핫스팟 발생 가능
  • 다른 파일: 자동 분산
  • Rebalancing 불필요: 자연스럽게 해결

해결 방법

GET + PUT 재배치:

"if you have data that's hot... you can just get put the file 
in place, and MinIO automatically will move it to wherever there's 
capacity."

Disney+ 사례:

  • 인기 영화 → 핫스팟 발생
  • GET + PUT으로 재배치
  • 용량 여유 있는 곳으로 이동

Rebalancing이 필요한 경우:

"rebalances makes more sense if everything is continuously 
being accessed."
  • 모든 데이터를 균등하게 지속 액세스
  • Rakuten 사례: 10개 Pool, 분기마다 추가
  • 3주간 Rebalancing 수행

7.3 Rebalancing 비용

시간 계산

벤더 제공:

"rebalancing 1.5 petabytes of data on their network should 
take between six and 12 days."

현재 환경:

  • 데이터: 1.5PB
  • 네트워크: 25 Gbps (추정)
  • 소요 시간: 6~12일

Rakuten 사례:

"Rakuten... the [final] rebalance... took like three weeks."
  • 10개 Pool
  • 대규모 데이터
  • 3주 소요

권장사항

벤더 조언:

"rebalancing... is going to take a while, especially because 
of the [consumption] of the network. So unless you already are 
constrained by access pattern, I will not recommend rebalance"

이유:
1. 네트워크 부하: 장기간 대역폭 소모
2. 성능 영향: Rebalancing 중 서비스 저하
3. 자동 균등화: 신규 데이터는 자동 분산

대안:

"new files coming in will automatically start getting load 
balanced... the weighted placement algorithm will start putting 
the data wherever there's [no] space."

8. 멀티 Pool 성능

8.1 Pool 성능 비교 테스트

테스트 계획

고객사 요청:

"they want to actually test if it really doesn't affect the 
performance. So for example, like they can do like 12 nodes, 
pool 0, 8 nodes [pool 1]"

사용 가능 자원:

  • 20개 노드 (테스트용)
  • 노드당 20개 드라이브
  • 3월까지 사용 가능

테스트 시나리오

벤더 권장:

"we do two warps, one small objects and one large objects on 
the large 20 node pool, single pool... and then we destroy it. 
And then we can do [multiple pools]"

방법론:
1. Single Pool (20 노드)

  • Small objects Warp
  • Large objects Warp
  • 성능 측정
  1. Multiple Pools

    • 4 노드 Pool × 5개
    • 또는 5 노드 Pool × 4개
    • 동일 테스트 수행
  2. 비교 분석

    • Throughput: 동일 예상
    • Request/sec: 1~2% 차이 예상

벤더 예상:

"single pool should be like this, and like multiple should be 
slightly behind, but like pretty close, like one 2% difference, 
mainly on the request per second. Throughput should be the same"

8.2 실제 벤치마크 사례

JP Morgan Chase

테스트:

"we even tested that with JP Morgan Chase. There was no there 
was like, 1% performance difference between... an empty cluster 
and a half cluster"

결과:

  • 빈 클러스터 vs 50% 찬 클러스터
  • 성능 차이: 1% (오차 범위)
  • Rebalancing 불필요 입증

벤더 평가:

"It's like a rounding error. It's line noise. There's going to 
be more of a degradation from moving data around... because you 
think you need balance, versus actually just using [it]"

Texas Instruments

벤더 언급:

"Disney+/JPMC operations, Tesla... benchmarks prove identical 
performance"

공통 결론:

  • Pool 개수와 성능 무관
  • Erasure Set 크기가 핵심
  • 네트워크 카드가 동일하면 성능 동일

9. GPU 클러스터 네트워크 요구사항

9.1 Training vs Inference 차이

Training 요구사항

Nvidia SuperPOD 사양:

"Nvidia says, take one of these [VT content with 8 GPUs] and 
put 72 together, right? 72 [nodes], which [is] 576 GPUs."

네트워크 처리량:

  • GPU당: 100~500 MB/s
  • 72 노드: 64~288 GB/s
  • MinIO 노드: 4~6개 충분

벤더 계산:

"assuming MinIO has [2 × 400 gigabit NICs]... 72 nodes of 
[VT content] for training alone will be four to six [MinIO nodes]"

Inference 요구사항

KV Cache 처리:

"KB cache... between 18 gigabytes a second to 70 gigabytes"

단일 배치 (32 사용자):

  • 18~70 GB/s
  • Training보다 낮음
  • 모델 크기에 따라 변동

9.2 네트워크 병목 경고

40 Gbps의 한계

벤더 경고:

"GPUs can consume very a lot of data, but if the networking 
still stays at 40 gigabits... you got a lot of more, way more 
storage nodes to satisfy that"

의미:

  • GPU 데이터 소비 >> 네트워크 처리
  • 40 Gbps로는 많은 스토리지 노드 필요
  • 100 Gbps+ 권장

대역폭과 노드 수 관계

공식:

"what matters... is the overall throughput. And if you have 
two 400 [Gbps NICs] per server with very few [MinIO] nodes, 
you can satisfy [requirements]"

최적화:

  • 높은 NIC 대역폭 → 적은 노드
  • 낮은 NIC 대역폭 → 많은 노드
  • 400 Gbps 권장

9.3 AI Factory 네트워크 전략

백엔드: InfiniBand

확정 사항:

"they will use [400G] InfiniBand as backend... that's confirmed."

InfiniBand 특징:

  • 초저지연
  • GPU 간 통신 최적화
  • NVIDIA 에코시스템

스토리지: Ethernet

트렌드 분석:

"Ethernet transition trend (Meta/xAI: complexity avoidance, 
Tesla: standardized Ethernet adoption, InfiniBand vendor 
lock-in avoidance)"

Ethernet 장점:

  • 표준화
  • 벤더 독립
  • 관리 용이
  • 비용 효율

트레이드오프:

  • InfiniBand: 최고 성능 / 높은 비용
  • Ethernet: 표준화 / 관리 편의 / 벤더 독립

10. 액션 아이템

10.1 벤더 측 (MinIO/AIStore)

즉시 실행 (1주일 내)

  1. 네트워크 최적화 가이드

    • BGP 설정 모범 사례 문서
    • Sidekick vs BGP 비교 가이드
    • Oversubscription 진단 방법
  2. 성능 테스트 지원

    • Warp 테스트 시나리오 제공
    • Single vs Multi-pool 벤치마크 스크립트
    • 결과 분석 템플릿
  3. 하드웨어 권장사항

    • PCIe Gen 4/5 선택 가이드
    • NVMe 용량별 노드 수 계산기
    • 네트워크 대역폭 산정 도구

단기 (2주일 내)

  1. AI Factory 아키텍처 설계

    • GPU 클러스터 네트워크 다이어그램
    • MinIO 노드 수 계산 (Training/Inference)
    • InfiniBand + Ethernet 통합 설계
  2. 확장 계획 지원

    • Pool 1 추가 시 네트워크 체크리스트
    • Spine 확장 필요성 평가
    • 108 노드 구성 네트워크 설계

중기 (1개월 내)

  1. 레퍼런스 아키텍처
    • Tesla 네트워크 구성 사례
    • JP Morgan Chase 벤치마크 결과
    • Disney+ 핫스팟 처리 사례

10.2 고객사 측 (SK Hynix)

즉시 실행

  1. 현재 네트워크 검증

    • Oversubscription 여부 재확인
    • Spine-Leaf 포트 사용률 모니터링
    • BGP 성능 측정 (내부 네트워크)
  2. BGP 완료

    • 외부 네트워크 BGP 적용 (Bond 0)
    • 성능 비교 (적용 전후)
    • 안정성 검증
  3. 테스트 환경 준비

    • 20 노드 Warp 테스트 계획
    • Single vs Multi-pool 벤치마크
    • 결과 문서화

단기

  1. Pool 1 확장 준비

    • 네트워크 용량 확인 (Spine 필요성)
    • 스위치 포트 확보
    • 케이블링 계획
  2. AI Factory 기획

    • 30TB NVMe 공급 대안 검토
    • 네트워크 대역폭 요구사항 정리
    • GPU 수량 및 구성 확정

중기

  1. 장기 확장 전략
    • 3년 후 3,000 노드 네트워크 설계
    • GPU Farm (1,000~5,000 GPU) 네트워크
    • 용인 AI Factory 전력 및 네트워크 인프라

10.3 공동 작업

성능 테스트 (2주 내)

  1. Warp 벤치마크

    • 일정: Pool 1 추가 전
    • 내용: Single vs Multi-pool 성능 비교
    • 참여: 양측 기술팀
  2. BGP 성능 검증

    • 일정: 외부 네트워크 BGP 완료 후
    • 내용: 기존 대비 성능 개선 측정
    • 비교: Latency, Throughput, Error rate

아키텍처 리뷰 (1개월 내)

  1. AI Factory 설계 검토
    • GPU 클러스터 네트워크
    • MinIO 스토리지 클러스터
    • InfiniBand + Ethernet 통합

11. 네트워크 모범 사례

11.1 Spine-Leaf 설계 원칙

확장성

  1. Oversubscription 회피

    • Spine 포트를 충분히 확보
    • 1:1 비율 유지 (Leaf:Spine)
  2. 이중화

    • 최소 2개 Spine
    • 단일 장애 대응
  3. 성장 여력

    • 현재 사용률 50% 이하 권장
    • 향후 2배 확장 고려

11.2 BGP 운영

설정 관리

  1. 중앙집중식

    • 네트워크 장비에서 통합 관리
    • 클라이언트 재설정 불필요
  2. 모니터링

    • 경로 변화 추적
    • 장애 탐지 자동화
  3. 문서화

    • IP 범위 및 경로 정보
    • 변경 이력 관리

11.3 용량 계획

드라이브 선택

  1. PCIe Gen 4+: 필수
  2. 7~15TB: 현실적 선택
  3. 30TB: 미래 대비

노드 수 계산

  • 공식: (Target Capacity / Drive Size / Drives per Node) / EC Efficiency
  • 예시: (5PB / 7TB / 20) / 0.63 ≈ 57 노드

네트워크 대역폭

  • 최소: 40 Gbps (현재)
  • 권장: 100 Gbps (일반)
  • GPU: 400 Gbps (AI Factory)

12. 주요 기술 결정 사항

12.1 BGP 채택

결정: BGP 사용 (Sidekick 대신)

  • 이유: 이미 구축, 관리 간편, 확장성
  • 시기: 내부 완료, 외부 1개월 내

12.2 네트워크 구성

결정: Single Fabric 유지

  • 이유: 충분한 성능, 비용 효율
  • 현재: Oversubscription 없음

12.3 Rebalancing

결정: 당분간 불필요

  • 이유: 자동 균등화, 높은 비용
  • 조건: 핫스팟 발생 시 GET+PUT

12.4 AI Factory 네트워크

결정: 400G InfiniBand + Ethernet

  • GPU 백엔드: InfiniBand
  • 스토리지: Ethernet
  • 대역폭: 400 Gbps

13. 위험 요소 및 완화 방안

13.1 네트워크 병목

위험

  • GPU 처리량 >> 네트워크 대역폭
  • 40 Gbps로는 부족

완화

  • 100 Gbps 이상 업그레이드 검토
  • 스토리지 노드 수 증가
  • GPU 워크로드 분산

13.2 Spine 용량

위험

  • Pool 1 추가 시 Spine 부족
  • 108 노드 수용 불가

완화

  • Spine 추가 계획 수립
  • 포트 사용률 지속 모니터링
  • 사전 용량 확보

13.3 하드웨어 공급

위험

  • 30TB NVMe 미출시
  • 대용량 드라이브 품귀

완화

  • 15TB 대안 검토
  • Cisco/Lenovo 재고 확인
  • 다중 공급선 확보

14. 참고 자료

14.1 벤더 사례

  • Tesla: Sidekick 대규모 환경 한계
  • JP Morgan Chase: Multi-pool 성능 검증 (1% 차이)
  • Disney+: 핫스팟 GET+PUT 해결
  • Rakuten: 10 Pool 운영 (3주 Rebalancing)

14.2 기술 문서

  • BGP 설정 가이드
  • Spine-Leaf 설계 원칙
  • Warp 벤치마크 방법론
  • GPU 네트워크 요구사항

15. 결론 및 다음 단계

15.1 핵심 합의사항

  1. BGP 채택: Sidekick 대신 BGP 사용
  2. 네트워크 안정: Oversubscription 없음, 확장 가능
  3. Rebalancing 불필요: 자동 균등화 의존
  4. 40 Gbps 한계: GPU 워크로드 시 업그레이드 필요
  5. AI Factory 설계: 400G InfiniBand + Ethernet

15.2 즉시 실행 항목

  • BGP 외부 네트워크 완료 (1개월 내)
  • Warp 성능 테스트 (Pool 1 전)
  • AI Factory 네트워크 설계

15.3 성공 기준

  • ✅ Pool 1 추가 시 성능 저하 없음 (< 2%)
  • ✅ BGP 성능 개선 확인 (3% 이상)
  • ✅ AI Factory 네트워크 400 Gbps 달성
  • ✅ 108 노드 확장 시 Oversubscription 없음

문서 버전: 1.0
최종 수정일: 2026년 2월 6일
다음 리뷰: Pool 1 추가 후 성능 검증 (2026년 3월)


부록: 용어 정리

  • Spine-Leaf: 계층형 네트워크 토폴로지
  • BGP: Border Gateway Protocol (경계 게이트웨이 프로토콜)
  • Sidekick: 클라이언트 측 로드 밸런서
  • Oversubscription: 과다 구독 (포트 분할 사용)
  • Erasure Coding (EC): 데이터 보호 방식
  • InfiniBand: 고성능 네트워크 기술
  • NVMe: 고속 SSD 인터페이스
  • PCIe Gen: PCI Express 세대

0개의 댓글