26F12D

Young-Kyoo Kim·2026년 2월 11일

클러스터 확장 및 리밸런싱 회의록


1. 회의 배경 및 현황

1.1 현재 클러스터 상태

Pool 0 (기존)

구성:

  • 노드 수: 54개
  • 드라이브: 노드당 20개 NVMe
  • 용량: 약 5PB (usable)
  • 데이터: ETL 데이터 (이미 마이그레이션 완료)
  • 사용률: ~60%

회의 중 확인:

"as is right now they have, like, one pool, 45 clusters. And you 
see the pool zero is where they have right now, yeah, 54"

1.2 확장 계획

Pool 1 (2026년 3월)

계획:

"in March there will be, like similar spike, same number, same 
memory. Yeah, more memory, and it'll be pool one."

사양:

  • 노드 수: 54개 (Pool 0와 동일)
  • 메모리: Pool 0보다 증가
  • 용량: ~5PB (usable, 예상)
  • 데이터: Hadoop 데이터 마이그레이션

Pool 2 (시기 미정)

계획:

"They don't know when there will be a third pool coming in... 
it'll be around 20 nodes, so it'll be lower number of nodes."

사양:

  • 노드 수: 약 20개
  • 특징: 소규모 Pool
  • 시기: 미확정

1.3 확장 목표 및 우려사항

목표

  1. 무중단 확장: Production 환경에서 다운타임 없이
  2. 성능 유지: 확장 후에도 동일한 성능
  3. 최적화: 데이터 분산 최적화

고객사 질문:

"how do we get pool in production? Like, you know, like, without 
no downtime... what are the checklists? Like, what is a driver?"

우려사항

  1. 리밸런싱 필요성: Pool 0에 ETL 데이터 100%, Pool 1 비어있음
  2. 핫스팟: 특정 데이터셋 반복 액세스 시 성능 저하
  3. 네트워크 부하: 리밸런싱 시 네트워크 대역폭 소비

2. Pool 확장 메커니즘

2.1 Pool 추가의 동작 원리

Single Pool vs Multi-Pool

벤더 설명:

"the behavior of one pool cluster and two pool cluster is different, 
right? But once you go to two, pool is a two, pool, three, pool four, 
pool five, pool is the same behavior."

핵심:

  • 1 Pool → 2 Pool: 동작 방식 변화
  • 2 Pool → 10 Pool: 동일한 동작 방식
  • 오버헤드: Pool 수 증가에 따라 약간 증가

요청 처리 메커니즘

Single Pool:

"when the request comes to any node in the system, this node, even 
though I don't have the data, it will know in constant time where 
the file is immediately just [from] its key, its name"

특징:

  • 상수 시간(Constant time) 조회
  • 중앙 메타데이터 불필요
  • 즉시 파일 위치 파악

Multi-Pool:

"when the request arrives to the server, we'll make the request to 
all the pools, right? So pool one through pool 10 at the same time 
fan out... So latency wise, is the same cost"

메커니즘:
1. 요청이 임의 노드 도착
2. 모든 Pool에 동시 Fan-out 쿼리
3. 파일 소유 Pool 응답
4. 해당 Pool에서 데이터 스트리밍

지연시간:

"that extra step taxes the performance. But of course, it's only 
millisecond tax. So if you're latency sensitive, you wouldn't want 
this. But in reality, it's not that big of an impact."
  • 추가 지연: 수 밀리초
  • 실제 영향: 미미
  • Latency-sensitive 애플리케이션만 고려 필요

2.2 크기가 다른 Pool 추가

성능 영향

고객사 질문:

"when there are different number [of nodes]... when pool zero is 
50 nodes... and then pool 1 [is] 25 way less notes."

벤더 답변:

"even though the pool is smaller than the original what matters is 
the set size. Set Size, same size will be same performance given 
the same network cards and similar NVMe."

핵심 원리:

  • Erasure Set 크기: 성능 결정 요소
  • Pool 크기: 성능과 무관 (Set 크기 동일 시)
  • 예시: EC 8:3, Set Size 4
    • Pool 0: 50 노드, Set Size 4
    • Pool 1: 25 노드, Set Size 4
    • 성능 동일

기술적 설명

데이터 읽기:

"let's assume initial set size equals four. When you do [a read], 
let's say your request arrives on the first node... you'll only 
be reading four drives. So at the same time, either these four 
nodes or these four nodes will be [reading] at the same rate."

병렬 처리:

  • EC 8:3: 8개 드라이브에 분산, 5개 읽으면 복원
  • 4개 노드에서 동시 읽기
  • Pool 크기와 무관하게 동일한 4개 노드 활용

AI Factory 예시

계획:

  • Pool 0: 50 노드
  • Pool 1: 10 노드
  • Pool 2: 10 노드

벤더 확인:

"the cluster will actually behave as if now it has 75 [nodes]... 
But both writes and reads have a small latency increase... the 
behavior of two pools, three pools, 10 pools, is the same as 
only two pools."

결과:

  • 총 70 노드처럼 동작
  • 지연 시간 미세 증가 (1~2%)
  • Throughput 동일

3. 자동 로드 밸런싱

3.1 Weighted Placement Algorithm

작동 원리

벤더 설명:

"new files coming in will automatically start getting load balanced. 
So the problem, even if you have the problem right now, as the time 
goes by, the problem will naturally start balancing"

메커니즘:

"the weighted placement algorithm will start putting the data 
wherever there's [most] space. When they start balancing their 
[capacity], yeah."

핵심 개념:
1. 신규 데이터: 자동으로 용량 여유 있는 Pool로
2. 가중치 기반: 빈 공간 많은 Pool에 높은 가중치
3. 자연스러운 균등화: 시간 경과에 따라 자동 밸런싱

3.2 시나리오 분석

초기 상태 (Pool 1 추가 직후)

상황:

  • Pool 0: 60% 사용 (ETL 데이터)
  • Pool 1: 0% 사용 (완전히 비어있음)

고객사 우려:

"if you don't do any rebalancing, it'll probably become like, 20% 
will be Hadoop [in Pool 0], and here [Pool 1], like, you know, 
like 60% can be like Hadoop and maybe a little bit of ETL"

예상 결과 (리밸런싱 없이):

  • Pool 0: 60% ETL + 20% Hadoop = 80%
  • Pool 1: 10% ETL + 60% Hadoop = 70%

고객사 질문:

"wouldn't it be better if they have like, 50-50?"

벤더 답변: 실질적 차이 없음

핵심 설명:

"Think about it this way. It's a cluster. It's nothing [but] a 
collection of erasure sets. An erasure set has a collection of 
drives. So even within the ETL data set, data is scattered."

기술적 이유:
1. Erasure Coding: 데이터는 이미 여러 노드/드라이브에 분산
2. Set 기반: 파일은 Erasure Set 단위로 분산
3. Pool 추가: 더 많은 Set 선택지 제공

결론:

"The presence of another pool means more [erasure sets], just more 
[sets] for us to grab data from... every single ETL or [Hadoop] 
file will be scattered somewhere else."
  • 50:50 vs 80:20 → 실질적 차이 없음
  • 모든 파일이 여러 Set에 분산
  • Pool 간 자연스럽게 균등 액세스

3.3 네트워크 카드 활용

자동 부하 분산

벤더 설명:

"as soon as this becomes a pool, they will naturally have [access], 
yes, because you have twice as many network cards to serve the ETL 
traffic as if you were treating them as separate."

메커니즘:
1. 요청 분산: 모든 노드가 요청 수신 가능
2. NIC 활용: 108개 노드의 모든 NIC 활용
3. 자연스러운 균등화: 네트워크 레벨에서 자동 분산

고객사 시나리오:

"you definitely want both applications [ETL and Hadoop]... the 
request to travel to both [pools]... whichever server gets the 
request... will retrieve the data from the other pool"

결과:

  • ETL 요청 → Pool 1 노드 도착 → Pool 0에서 데이터 가져옴
  • Hadoop 요청 → Pool 0 노드 도착 → Pool 1에서 데이터 가져옴
  • 자동으로 모든 Pool의 네트워크 활용

4. 리밸런싱 (Rebalancing)

4.1 리밸런싱이란?

정의

기존 데이터를 Pool 간 재분배하는 작업

벤더 설명:

"the only thing that will happen is that [Pool 0] will do some stuff, 
and then you get a tiny bit [in Pool 1]... This can only happen after 
rebalance. New files coming in, they'll be automatically placed into 
both [pools]"

특징:

  • 기존 데이터: 리밸런싱 필요
  • 신규 데이터: 자동 균등 분산
  • 목적: 기존 데이터 재배치

4.2 리밸런싱 불필요 권장

벤더 권장사항

명확한 입장:

"rebalancing doesn't make sense when you have rebalance. I mean, 
you're already petabyte size, petabyte rebalance is going to take 
a while, especially because of the [consumption] of the network. 
So unless you already are constrained by access pattern, I will 
not recommend rebalance"

이유:
1. 시간 소요: 매우 오래 걸림
2. 네트워크 부하: 대역폭 집중 소비
3. 실익 없음: 자동 균등화로 충분

비용 계산

현재 환경:

"rebalancing 1.5 petabytes of data on their network should take 
between six and 12 days."

조건:

  • 데이터: 1.5PB (Pool 0의 절반)
  • 네트워크: 25 Gbps (추정)
  • 소요 시간: 6~12일

대규모 환경 시:

"imagine when you have 20 petabytes, and you think, Oh, let me 
rebalance 10 petabytes, it will take even longer. So unless the 
networking is faster."
  • 20PB 환경 → 10PB 리밸런싱
  • 소요 시간: 수 주일
  • 네트워크 업그레이드 없으면 더 오래

사용자 영향

서비스 저하:

"but then [users] will be like, this is slow. It's like, oh, 
we are rebalancing. Sorry, yeah"

대안:

"if you [did], let's start rebalance at night and stop rebalancing 
in the morning. Maybe [users] will notice that you're slowly shifting 
the data. Could work"
  • 야간 리밸런싱 가능
  • 주간 중단
  • 점진적 데이터 이동

4.3 리밸런싱이 필요한 경우

예외적 상황

벤더 설명:

"rebalances makes more sense if everything is continuously being 
accessed."

조건:

  • 모든 데이터: 지속적으로 균등하게 액세스
  • 모든 Pool: 동일한 읽기 빈도
  • 예시: Rakuten 사례

5. 핫스팟 (Hot Spot) 처리

5.1 핫스팟이란?

정의 및 발생 조건

고객사 시나리오:

"Team A... wants to... read a very specific data over and over 
and over... that'll hit the same pool, like, same location of 
the disk, like, over and over and over."

우려:

  • 특정 데이터셋 반복 읽기
  • 동일 Pool/드라이브 집중 액세스
  • 성능 저하 가능성

실제 발생 조건

벤더 설명:

"[Hot spot] only happens if it's the same file. So let's say you 
do compaction, the file becomes one gigabyte, and now you're only 
reading that over and over and over."

핵심:

  • 동일 파일: 핫스팟 발생
  • 다른 파일: Erasure Set 분산으로 자동 해결
  • 확률: 2개 파일이 동일 Set일 가능성 낮음

결론:

"as soon as it's two objects, those are most likely going to be 
[in different] erasure sets... most likely they're going to be 
separate"

5.2 리밸런싱으로 해결 불가

벤더 명확화:

"if that's the case, rebalancing will eventually not solve the 
problem... if hotspot only happens for the [same] files, it 
doesn't really help"

이유:

  • 동일 파일 → 동일 Erasure Set
  • 리밸런싱해도 동일 노드/드라이브
  • 핫스팟 지속

5.3 해결 방법: GET + PUT

Disney+ 사례

벤더 소개:

"the history of rebalance was [Disney+]... they thought they needed 
to... rebalance their clusters because of the video. But turns out... 
they found out that there's a hot file, hot content, they can just 
get put the file in place, and MinIO automatically will move it to 
wherever there's capacity."

Disney+ 상황:

  • 인기 영화 출시
  • 모든 사용자가 동일 파일 시청
  • 특정 노드/드라이브 집중 액세스

해결책:

"if there's only a movie, very popular movie, everyone wants to 
watch the movie, then they [do] the get put, so it's [redistributed]"

메커니즘:
1. GET: 인기 파일 읽기
2. PUT: 동일 파일 다시 쓰기
3. 자동 재배치: MinIO가 용량 여유 있는 곳에 자동 배치

효과:

"it'll then overwrite wherever there is [capacity], and that's 
going to be on the new pool."
  • 새 Pool로 자동 이동
  • 핫스팟 해소
  • 리밸런싱 불필요

Disney+ 결과

벤더 평가:

"they rarely rebalance because [it] didn't make sense."
  • GET + PUT로 충분
  • 리밸런싱 거의 안 함
  • 비용 대비 효과 우수

6. 고객 사례 연구

6.1 Disney+ (디즈니 플러스)

배경

서비스 특성:

  • 전 세계 비디오 스트리밍
  • Pool = 추가 대역폭
  • 대량 동시 접속

초기 계획:

"when they [added pools], they thought they needed to... rebalance 
their clusters because of the video."
  • Pool 추가 = 대역폭 증가
  • 리밸런싱 필요하다고 판단

실제 운영

벤더가 구축:

"we built this for them... But so far, they just, they never used 
[rebalancing]"

이유:

  • GET + PUT 방식 효과적
  • 인기 콘텐츠만 재배치
  • 전체 리밸런싱 불필요

결론:

  • 리밸런싱 기능 있음
  • 실제로는 거의 사용 안 함

6.2 Rakuten (라쿠텐)

배경

공격적 확장:

"Rakuten... were adding pools like every quarter. At some point 
they had 10 pools."

특징:

  • 분기마다 Pool 추가
  • 최대 10개 Pool 운영
  • 모든 데이터 균등 액세스

리밸런싱 수행

이유:

"they will actually read from everywhere. So they will continuously 
rebalance."
  • 모든 Pool에서 지속적 읽기
  • 균등한 데이터 액세스 패턴
  • 리밸런싱 효과 있음

최종 리밸런싱

벤더 경험:

"the [final] rebalance was the heavy one. I think it took like 
three weeks."

과정:

  • 최종 대규모 리밸런싱: 3주 소요
  • 이후: 점진적 소량 리밸런싱
  • 최적화: 느리고 지속적으로

벤더 권고

Pool 병합:

"At some point they reach [optimal] pool size, we told them start 
merging"
  • 10개 Pool → 병합 권장
  • 관리 복잡도 감소
  • 성능 최적화

결론

유일한 예외:

"that was the only case where it made sense, but they were reading 
very aggressively from everywhere."
  • 모든 데이터 균등 액세스 → 리밸런싱 효과
  • 대부분의 경우 → 리밸런싱 불필요

6.3 JP Morgan Chase

벤치마크 테스트

목적:

  • 빈 클러스터 vs 50% 찬 클러스터
  • 리밸런싱 필요성 검증

벤더 설명:

"we even tested that with JP Morgan Chase. There was [like] 1% 
performance difference between... an empty cluster and a half cluster"

결과:

  • 성능 차이: 1% (오차 범위)
  • 결론: 리밸런싱 불필요

고객 반응:

"they were really, really angry when we told them you don't need 
to rebalance, because they're so used to rebalancing on their 
NetApp and on their Pure and on their EMC"

기존 스토리지 습관:

  • NetApp, Pure, EMC → 리밸런싱 필수
  • MinIO → Erasure Coding으로 불필요
  • 개념 전환 어려움

실증:

"when we... because we're using erasure code. If I'm using four 
drives across four nodes over here and four drives across four 
nodes over here, the capacity doesn't matter. It's the number of 
drives across the number of nodes."

최종 합의:

"That takes away the need for any rebalancing."
  • 실제 벤치마크로 증명
  • 리밸런싱 불필요 인정

6.4 Texas Instruments

벤더 언급:

  • 유사한 테스트 수행
  • 동일한 결론
  • 리밸런싱 불필요

7. Pool 병합 (Merging)

7.1 Thin Provisioning을 활용한 병합

개념

벤더 설명:

"on the AI store directly, there's thin provisioning of volumes. 
So even if the drive doesn't have capacity for a new volume, you 
can allocate a volume that's not really provisioned"

Thin Provisioning:

  • 물리적 공간 없이 볼륨 할당
  • 실제 사용 시 공간 할당
  • DirectPV 기능

병합 프로세스

시나리오:

"if you have two pools, you can decommission a smaller pool into 
a larger pool. Let's say you have five or 10 nodes, and you're 
bringing another 5-10 nodes, you could decommission the smaller 
pool into a larger pool with a thin provision[ed] device."

단계:
1. Pool 1: 5 노드 (기존)
2. Pool 2: 5 노드 추가
3. 목표: 10 노드 단일 Pool

방법:

"That's how you can actually make the second pool merge into the 
third pool."

7.2 Decommissioning (해체)

작동 원리

고객사 질문:

"when they want to... [merge] the pool, yeah, first, [decommission] 
the existing pool one... which means that, like data in [pool one] 
should be migrated to... pool zero or something, right?"

벤더 답변:

"if you didn't have thin provision... the CSI will say there's no 
space. So that's why you need thin provision, so that you can 
actually schedule both on the same place and then do the in place 
replacement"

메커니즘:
1. Thin Provision: 동일 물리 드라이브에 2개 볼륨 할당
2. Decommission: Pool 1 데이터를 Pool 0으로 이동
3. 물리 공간: 점진적으로 Pool 0 볼륨으로 재할당

예시:

"let's say [the pool] is 60%, [there's] some space. So you 
[thin] provision another seven terabyte drive [on the] same 
physical drive... So you start doing the decommission... and 
the data will start [filling]"
  • Pool 0: 60% 사용 (4.2TB / 7TB)
  • Thin Provision: 추가 7TB 볼륨 할당 (동일 물리 드라이브)
  • Decommission: Pool 1 데이터 이동
  • 실제 공간: 점진적 할당

7.3 Thin Provisioning 필요성

질문

고객사:

"What if... thin provisioning is [disabled]... decommission is 
still available... they're adding the existing pool one and then 
add to the additional node, and the creation pool... is still 
available, right?"

벤더 답변:

"Yes, that they can do that, yes... it's not a requirement, but 
still, it makes it convenient for the... dynamic [volume allocation]... 
and also... The management of the PVC"

결론:

  • Thin Provisioning 없어도: 병합 가능
  • 있으면: 훨씬 편리
  • In-place 병합: Thin Provisioning 필요

8. 확장 체크리스트

8.1 사전 준비

인프라 점검

고객사 요청:

"what are the checklists? Like, what is a driver? [What] should 
they check... what is difference in... production cluster?"

항목:
1. 네트워크

  • Spine 용량 확인 (108 노드 수용 가능?)
  • BGP 설정 완료
  • 대역폭 모니터링
  1. 하드웨어

    • 노드 수령 및 검수
    • NVMe 드라이브 설치
    • PCIe Gen 4+ 확인
  2. 소프트웨어

    • Kubernetes 버전 호환성
    • DirectPV 준비
    • AIStore 버전 확인

8.2 Pool 추가 절차

단계별 가이드

1단계: Pool 생성

  • AIStore/Helm Chart 설정
  • EC 설정 확인 (EC 8:3 유지)
  • Pool 이름 지정 (pool-1)

2단계: 검증

  • 노드 온라인 확인
  • 드라이브 인식 확인
  • 네트워크 연결 확인

3단계: 성능 테스트

  • Warp 벤치마크
  • Single Pool vs Multi-Pool 비교
  • 예상: 1~2% 차이

4단계: 프로덕션 투입

  • 모니터링 활성화
  • Alert 설정
  • 사용자 공지 (필요시)

8.3 데이터 마이그레이션

Hadoop 마이그레이션

계획:

  • 시기: Pool 1 추가 후
  • 데이터: Hadoop 데이터 (~3PB 예상)
  • 방법: Cloudera → MinIO 직접 복사

자동 분산:

"new files coming in will automatically start getting load balanced"
  • Pool 1이 비어있음 → 높은 가중치
  • Hadoop 데이터 자동으로 Pool 1 우선 배치
  • 점진적으로 Pool 0에도 분산

리밸런싱 불필요

벤더 확인:

"So the question, Can I split this and do ETL Hadoop? Yes, you can. 
Not a problem. So it's just a warning that we don't want you to 
think that rebalancing a petabyte of data is good idea."
  • ETL/Hadoop 분리 가능
  • 리밸런싱 권장 안 함
  • 자동 균등화로 충분

9. 용량 계획

9.1 확장 임계값

Golden Ratio

고객사 질문:

"what is the like, the golden ratio... when [do they] know they 
have [to] expand... like, 60%? when pool zero is at 60% or 70%?"

벤더 권장:

"80% [is the expansion trigger]"

임계값:

  • 50%: 여유 있음
  • 70%: 확장 계획 시작
  • 80%: 확장 실행
  • 85%: Critical

예비 용량:

"we reserve a certain amount of capacity... 50% [margin]... 
80% [trigger]"

9.2 하드웨어 공급 문제

시장 상황

벤더 경험:

"we have a customer that wants to add Four exabytes, they cannot 
get The servers in time."

현실:

"we also have a bank customer that wants to add a pool and they 
can't get memory for their servers."

공급사 선택:

"Customers... are selecting their hardware vendors based on who 
actually can deliver... Lenovo and Cisco were hoarding all the 
eight and 16 terabyte... drives"

권장:

  • Cisco, Lenovo: 재고 확보
  • 사전 주문 필수
  • 대안 벤더 확보

10. 성능 테스트 계획

10.1 Warp 벤치마크

테스트 환경

가용 자원:

"they currently have 20 nodes available for any type of testings 
right now, until March."
  • 노드: 20개
  • 드라이브: 노드당 20개
  • 기간: 3월까지

테스트 시나리오

벤더 제안:

"we do two warps, one small objects and one large objects on the 
large 20 node pool, single pool... and then we destroy it. And 
then we can do [multiple pools]"

방법론:
1. Single Pool (20 노드)

  • Small objects Warp
  • Large objects Warp
  1. Multiple Pools

    • 4 노드 × 5 Pool
    • 또는 5 노드 × 4 Pool
    • 동일 테스트 반복
  2. 비교 분석

    • Throughput
    • Request/sec
    • Latency

예상 결과:

"single pool should be like this, and like multiple should be 
slightly behind, but like pretty close, like one 2% difference, 
mainly on the request per second. Throughput should be the same"
  • Throughput: 동일
  • Request/sec: 1~2% 차이
  • Latency: 수 밀리초 차이

10.2 JP Morgan Chase 사례 재현

목표:

  • 빈 Pool vs 부분 찬 Pool
  • 성능 차이 검증

방법:

"test four node pool, add to a four node pool, add to a four node 
pool, versus just doing a 12 node pool."
  • 4 + 4 + 4 노드 (3 Pools)
  • vs 12 노드 (1 Pool)
  • 성능 비교

JPMC 결과:

"we did two fours, and then we did an eight... the performance 
it was, it was basically a nil. I mean, it was 1%"

11. 액션 아이템

11.1 벤더 측 (MinIO/AIStore)

즉시 실행 (1주일 내)

  1. 확장 가이드 문서

    • Pool 추가 체크리스트
    • 무중단 확장 절차
    • 검증 방법
  2. 리밸런싱 가이드

    • 불필요한 이유 설명
    • 예외 상황 (Rakuten 사례)
    • GET + PUT 방법 (Disney+ 사례)
  3. Warp 테스트 지원

    • 테스트 스크립트 제공
    • 결과 분석 템플릿
    • 성능 비교 보고서

단기 (2주일 내)

  1. Pool 병합 가이드

    • Thin Provisioning 설정
    • Decommission 절차
    • In-place 병합 방법
  2. 용량 계획 도구

    • 확장 임계값 계산기
    • 하드웨어 사이징 가이드
    • 비용 분석 템플릿

11.2 고객사 측 (SK Hynix)

즉시 실행

  1. Pool 1 준비

    • 하드웨어 주문 (Cisco/Lenovo)
    • 네트워크 용량 확인
    • Spine 확장 필요성 평가
  2. 테스트 환경

    • 20 노드 Warp 테스트
    • Single vs Multi-Pool 벤치마크
    • 결과 문서화
  3. Hadoop 마이그레이션 계획

    • CDP 라이센스 EOL 확인 (8월 31일)
    • 마이그레이션 일정 수립
    • 테스트 마이그레이션

단기

  1. Pool 1 투입 (3월)

    • 54 노드 설치
    • Pool 생성 및 검증
    • 성능 테스트
  2. 모니터링 강화

    • Pool별 용량 모니터링
    • 자동 밸런싱 확인
    • Alert 설정 (80% 임계값)

중기

  1. Pool 2 계획
    • 20 노드 규모 검토
    • 시기 결정
    • 예산 확보

11.3 공동 작업

성능 검증 (3월)

  1. Pool 1 추가 후 벤치마크

    • 일정: Pool 1 투입 직후
    • 내용: 성능 변화 측정
    • 목표: 1~2% 차이 확인
  2. 자동 밸런싱 모니터링

    • 일정: 3~6개월
    • 내용: Hadoop 데이터 분산 추적
    • 목표: 자연스러운 균등화 검증

12. 핵심 결정 사항

12.1 리밸런싱 불필요

결정: 리밸런싱 수행하지 않음

이유:
1. 자동 Weighted Placement Algorithm
2. 신규 데이터 자동 균등 분산
3. 높은 비용 (6~12일, 네트워크 부하)
4. 실익 없음 (성능 차이 미미)

예외:

  • 모든 데이터 균등 액세스 (Rakuten 사례만)
  • 현재 SK Hynix는 해당 없음

12.2 핫스팟 대응

결정: GET + PUT 방식 사용

방법:
1. 핫스팟 파일 식별
2. GET으로 읽기
3. PUT으로 다시 쓰기
4. 자동으로 용량 여유 있는 Pool로 이동

효과:

  • Disney+ 검증
  • 리밸런싱 불필요
  • 즉시 효과

12.3 확장 임계값

결정: 80% 용량 도달 시 확장

임계값:

  • 70%: 계획 시작
  • 80%: 확장 실행
  • 85%: Critical

12.4 Pool 크기

결정: 다양한 크기 Pool 허용

근거:

  • Erasure Set 크기만 중요
  • Pool 크기는 성능과 무관
  • 유연한 확장 가능

예시:

  • Pool 0: 54 노드
  • Pool 1: 54 노드
  • Pool 2: 20 노드 ✅ 가능

13. 위험 요소 및 완화 방안

13.1 하드웨어 공급 지연

위험

  • Pool 1 하드웨어 미도착
  • 확장 일정 차질

완화

  • ✅ Cisco/Lenovo 사전 주문
  • ✅ 대안 벤더 확보
  • ✅ 여유 시간 확보

13.2 네트워크 용량

위험

  • 108 노드 → Spine 부족
  • Oversubscription 발생

완화

  • ✅ Spine 확장 계획
  • ✅ 포트 사용률 모니터링
  • ✅ 사전 용량 확보

13.3 성능 저하

위험

  • Multi-Pool로 인한 성능 저하
  • 사용자 불만

완화

  • ✅ Warp 테스트로 사전 검증
  • ✅ 1~2% 차이 예상 (무시 가능)
  • ✅ JPMC/TI 사례 참고

14. 모범 사례

14.1 확장 시

  1. 리밸런싱 하지 말 것

    • 자동 균등화 신뢰
    • 비용 대비 효과 없음
  2. 80% 용량 도달 시 확장

    • 여유 있게 계획
    • 하드웨어 공급 시간 고려
  3. 다양한 크기 Pool 허용

    • 예산에 맞춰 유연하게
    • Erasure Set 크기만 일관되게

14.2 핫스팟 발생 시

  1. GET + PUT 사용

    • Disney+ 검증된 방법
    • 즉시 효과
  2. 리밸런싱 회피

    • 시간/비용 낭비
    • 동일 파일은 동일 Set

14.3 성능 관리

  1. Warp 정기 벤치마크

    • 확장 전후 비교
    • 기준선 유지
  2. 모니터링 강화

    • Pool별 용량
    • 자동 밸런싱 확인

15. 결론 및 다음 단계

15.1 핵심 합의사항

  1. Pool 1 추가 (3월): 54 노드
  2. 리밸런싱 불필요: 자동 균등화 신뢰
  3. 핫스팟 대응: GET + PUT 방식
  4. 확장 임계값: 80% 용량
  5. 다양한 Pool 크기: 허용

15.2 즉시 실행 항목

  • Warp 테스트 (20 노드)
  • Pool 1 하드웨어 주문
  • Hadoop 마이그레이션 계획

15.3 성공 기준

  • ✅ Pool 1 추가 후 성능 저하 < 2%
  • ✅ Hadoop 데이터 자동 분산 확인
  • ✅ 리밸런싱 없이 안정적 운영

문서 버전: 1.0
최종 수정일: 2026년 2월 6일
다음 리뷰: Pool 1 추가 후 (2026년 3월)


부록: 주요 인용문

Disney+

"they found out that there's a hot file, hot content, they can 
just get put the file in place, and MinIO automatically will 
move it to wherever there's capacity."

Rakuten

"Rakuten was adding pools like every quarter. At some point 
they had 10 pools... the [final] rebalance... took like three 
weeks."

JP Morgan Chase

"There was [like] 1% performance difference between... an empty 
cluster and a half cluster... they were really, really angry 
when we told them you don't need to rebalance"

자동 밸런싱

"new files coming in will automatically start getting load balanced... 
the weighted placement algorithm will start putting the data wherever 
there's [most] space."

0개의 댓글