26F12F

Young-Kyoo Kim·2026년 2월 11일

용량 및 성능 제한, Thin Provisioning 회의록


1. 회의 배경 및 현황

1.1 용량 제한 논의 배경

고객사 질문

기본 질문:

"Is it a billion? Is it 10 billion? Is it 100 billion?"

맥락:

  • 클러스터가 저장 가능한 객체 수는?
  • 확장 시 고려해야 할 한계는?
  • 파일시스템 제약사항은?

1.2 현재 클러스터 사양

Pool 0

  • 노드: 54개
  • 드라이브: 노드당 20개 NVMe
  • 드라이브 용량: 7.6TB 각
  • 총 Raw 용량: 54 × 20 × 7.6TB = 8,208TB ≈ 8.2PB
  • EC 8:3 효율: 63%
  • Usable 용량: ~5.2PB

2. Linux/XFS 파일시스템 제한

2.1 Inode 제한

기본 원리

벤더 설명:

"XFS is the file system... for every terabyte, you have 10 million 
objects where you can think of it."

공식:

1 TB = 10 Million inodes (files)

이론 vs 실제:

"theoretically it's 100 million, but Linux won't let you store that"
  • 이론적: 1TB당 100M 파일
  • 실제: 1TB당 10M 파일 (Linux 제한)
  • 제약: Linux 커널 레벨

2.2 현재 클러스터 용량 계산

단계별 계산

1단계: 드라이브당 객체 수

"So with a seven terabyte drive, you have approximately 70 million files."

계산:

  • 7.6TB × 10M/TB = 76M 파일/드라이브

2단계: 노드당 객체 수

"So if you have 20 per node, 20 times 70 equals 1400 [million]"

계산:

  • 20 드라이브 × 76M = 1,520M 파일/노드 = 1.52B 파일/노드

3단계: 클러스터 전체

"If you have 50 notes, 54... approximately 700... 70 billion"

계산:

  • 54 노드 × 1.52B = 82.08B 파일 (raw)

4단계: EC 8:3 적용

"what's the erasure? Erasure coding? 8:3, okay, so divide by eight"

최종:

"10, 10 billion objects I got 9.4"

결과:

  • 82B ÷ 8 (EC 오버헤드) ≈ 10.26B 객체
  • Pool 0: 약 100억 개 객체 저장 가능

2.3 NVMe vs HDD 차이

중요한 구분

벤더 경고:

"this is only for NVMe that you will actually get that. If this 
was hard drive, no way... hard drive will not... you will not 
get 10 million files on a hard drive that drive will stop working."

이유:

  • NVMe: Inode 처리 성능 우수
  • HDD: Inode 처리 시 극심한 성능 저하
  • 권장: NVMe 사용 필수

2.4 Inode 제한 근본 원인

벤더 설명:

"that has to do with the number of inodes that take up the files. 
So you'll run out of inode before you run [out of capacity]... 
The Linux File System is a bigger limitation."

핵심:

  • 용량보다 Inode 먼저 고갈
  • XFS 메타데이터 구조
  • Linux 커널 제약

3. XFS 메타데이터 예약

3.1 25% Inode 예약

XFS 기본 동작

벤더 설명:

"25% of the drives when you format XFS, one quarter is reserved 
for inodes, it doesn't have to have that, but that's what's 
allocated up front."

메커니즘:

  • 포맷 시: 25% 공간 Inode용 예약
  • 동적 할당: 실제로는 필요 시 할당
  • 보장: 충분한 Inode 공간 확보

3.2 작은 파일의 영향

극단적 시나리오

벤더 경고:

"When you, if you're writing a lot of small files, you'll only 
see 60% capacity available. That's just how XFS works, because 
you're going to run out of [inode] space."

상황:

  • 매우 작은 파일 대량 저장
  • Inode 25% 예약 공간 소진
  • 실제 사용 가능: 60%만

계산 예시:

  • Raw 용량: 100TB
  • EC 8:3: 63TB (63%)
  • 작은 파일 시: 60TB (60%)

3.3 예약 공간의 유연성

동적 할당

고객사 질문:

"100 terabyte raw... 25 terabyte is reserved... can be consumed 
by inode... Then what? How much can MinIO [use]?"

벤더 답변:

"it's just reserving that... it's reserved, but it's not [hard coded]... 
it's not taken."

설명:

  • 예약: 공간 확보만
  • 실제 사용: 필요 시 동적 할당
  • 유연성: 데이터도 사용 가능

최종 계산:

"75 terabyte, but that should include parity as well, right? 
So... 63% efficient"
  • 100TB raw
  • EC 8:3: 63TB usable
  • Inode 예약은 이미 고려됨

4. Pool 추가 시 용량 확장

4.1 Pool 1 추가 효과

추가 객체 수

벤더 설명:

"In their case, as soon as they have another Pool, they get 
another 4.5 billion [objects]. Exactly."

계산:

  • Pool 0: 10B 객체
  • Pool 1 (동일 사양): 10B 객체
  • : 20B 객체

4.2 제한 완화

벤더 확인:

"the limitation of a cluster is not based on the number... given 
the capacity of the cluster or the number of erasure sets... if 
you maxed it out, expanding the cluster will just open up the 
capacity"

핵심:

  • Pool 추가 = 객체 수 증가
  • Erasure Set 증가 = 용량 증가
  • 제한 해소: 클러스터 확장으로

5. Thin Provisioning

5.1 개념 및 필요성

정의

벤더 설명:

"on the AI store directly, there's thin provisioning of volumes. 
So even if the drive doesn't have capacity for a new volume, you 
can allocate a volume that's not really provisioned"

개념:

  • 물리 공간 없이 볼륨 할당
  • 사용 시 실제 할당
  • Over-commitment 가능

사용 목적

Pool 병합:

"you can actually leverage those to if you have two pools, you 
can decommission a smaller pool into a larger pool... with a 
thin provision[ed] device."
  • Pool 병합 시 In-place 작업
  • 물리 드라이브 교체 시 유연성
  • Multi-tenancy 격리

5.2 작동 원리

볼륨 할당

벤더 설명:

"if you have a seven terabyte drive, and you request a seven 
terabyte volume, so now we'll just say, Oh, you can use up to 
seven terabytes within this volume."

메커니즘:
1. 요청: 7TB 볼륨 생성
2. XFS 포맷: "최대 7TB 사용 가능" 설정
3. 실제 할당: 데이터 쓸 때만

Over-commitment

벤더 설명:

"let's say you put one terabyte of data. So now you have six 
terabytes left. So you can de provision another volume of seven 
terabytes, and then we'll assign it to the same drive."

시나리오:

  • 물리 드라이브: 7TB
  • PV 1: 7TB 할당, 1TB 사용
  • PV 2: 7TB 추가 할당 (같은 물리 드라이브)
  • 총 할당: 14TB
  • 실제 사용: 1TB
  • 가용: 6TB

제약:

"if both are filling and one reaches five terabytes, and [the other] 
two terabytes, they'll stop [writing], because they're [at] the 
real physical device."
  • PV 1 + PV 2 합산 사용량 ≤ 7TB (물리 제한)

5.3 비유 및 이해

항공사 오버부킹

벤더 비유:

"it's almost like... looking for a flight. Yes, exactly, yeah. 
It's like network bandwidth."

비유 1: 항공권

  • 좌석 100개
  • 예약 120명 (overbooking)
  • 실제 탑승: 95명 (no-show 고려)

비유 2: 네트워크 대역폭

"I can have as many clients connect in as I want, until it 
really happens... you only get eight gigs in the entire 
bandwidth we can share"
  • 8Gbps 회선
  • 여러 사용자 공유
  • 동시 사용 시 대역폭 분할

비유 3: 아파트 인터넷

"at home... kind of like a broadband... in house, like, apartment 
[I'm] watching Netflix. You're watching Disney... but you only 
get eight gigs in the entire bandwidth"

Guaranteed Not to Exceed

벤더 표현:

"guaranteed not to exceed... I guarantee not to exceed the capacity. 
So guaranteed I [can] have the whole capacity... but [they're] 
guaranteed to not have more than [the physical limit]."

핵심:

  • 보장: 물리 용량 초과 불가
  • 유연성: Over-commit 가능
  • 공정: 선착순 사용

5.4 XFS 기능

파일시스템 레벨

벤더 확인:

"it's a thin provision at the file system... it's a feature of 
XFS... You're not creating a hard partition. It's basically a 
soft [partition]"

특징:

  • XFS 네이티브: 파일시스템 기능
  • Soft Partition: 논리적 분할
  • 동적: 사용량에 따라 조정

Bare Metal 제약

고객사 질문:

"you can't do this... on bare metal, [there's] no way to do 
thin provisioning with bare [metal]"

벤더 답변:

  • DirectPV (Kubernetes CSI) 활용
  • XFS quota 기능 사용
  • Bare metal 단독으로는 불가

6. Thin Provisioning 활용 사례

6.1 Pool 병합 (In-place Decommission)

시나리오

목표:

  • Pool 1 (5 노드) + Pool 2 (5 노드) → Pool 3 (10 노드)

문제:

  • Pool 1 데이터를 Pool 3로 이동
  • 물리 공간 부족

해결책:

"if the [pool] is 60%, [there's] some space. So you [thin] 
provision another seven terabyte drive [on the] same physical 
drive... start doing the decommission... the data will start 
[filling] and [Pool 1 is] going down."

단계:
1. Pool 3 생성: Thin provisioning으로 볼륨 할당
2. 물리 공간: Pool 1과 공유 (60% 사용 중)
3. Decommission: Pool 1 → Pool 3 데이터 이동
4. 점진적 전환: Pool 1 감소, Pool 3 증가
5. 완료: Pool 1 제거, Pool 3만 남음

결과:

"at some point, okay, two seven terabyte [volumes] on the same 
physical drive"
  • 동일 물리 드라이브에 2개 볼륨
  • In-place 병합 완료

6.2 드라이브 교체 시 용량 증대

시나리오

벤더 설명:

"what if you replace a physical disc with something newer, that's 
larger... This is a great way to... start taking advantage."

문제:

  • 7.6TB 드라이브 → 15TB 드라이브 교체
  • 기존 PV는 7.6TB로 고정

Thin Provisioning 없이:

  • 15TB 중 7.6TB만 사용
  • 7.4TB 낭비

Thin Provisioning으로:
1. 기존 PV: 7.6TB (thin provisioned)
2. 새 PV: 7.4TB 추가 할당 (동일 물리 드라이브)
3. 활용: 15TB 전체 사용

6.3 Multi-tenancy 격리

목적

벤더 설명:

"the value this brings is I can have now multiple tenants 
[isolated]... I'm basically isolating them at the [volume level]"

사용 케이스:

  • 팀 A: PV 1 (7TB 할당)
  • 팀 B: PV 2 (7TB 할당)
  • 물리: 동일 14TB 드라이브

효과:

  • 논리적 격리
  • 쿼터 관리
  • 공정 분배

7. Thin Provisioning 제약사항

7.1 물리 용량 한계

명확한 제한:

"if both are filling and one reaches five terabytes, and [the other] 
two terabytes, they'll stop [writing]"

시나리오:

  • 물리: 7TB
  • PV 1: 5TB 사용
  • PV 2: 2TB 사용
  • 총 7TB 도달 → 쓰기 중단

7.2 Decommission 시 공간 필요

벤더 확인:

"if you didn't have thin provision... the CSI will say there's 
no space. So that's why you need thin provision"

필요성:

  • In-place 병합 시 필수
  • 물리 공간 부족해도 할당 가능
  • 점진적 데이터 이동 지원

7.3 선택 사항

고객사 질문:

"What if they [disable] thin provisioning... decommission is 
still available... they can do it... it's not a requirement, right?"

벤더 답변:

"Yes... it's not a requirement, but still, it makes it convenient 
for the dynamic [allocation]... flexibility... management of the PVC"

결론:

  • 필수 아님: Decommission 자체는 가능
  • 권장: 편의성 및 유연성 향상
  • In-place 병합: Thin provisioning 필요

8. 성능 관련 고려사항

8.1 Prefix 및 Listing

Flat Prefix 성능

벤더 경험:

"We have seen customers with, like, a billion objects on a single 
prefix. [Performance is] fine. Of course there's the problem of 
listing."

권장:

  • 단일 Prefix: 10억 객체까지 OK
  • Listing: 메타데이터 별도 관리 시 문제없음
  • 계층 구조: 가능하면 사용 (Iceberg 자동 분할)

8.2 Partitioning 효과

Inode 압력 완화:

"Lowers metadata and inode pressure: keeping files organized by 
partition avoids exploding file counts in a single namespace/location"

장점:

  • 파일 수 분산
  • 메타데이터 처리 효율
  • Inode 제약 완화

9. 권장사항

9.1 용량 계획

객체 수 기준

공식:

Max Objects = (Nodes × Drives/Node × TB/Drive × 10M/TB) ÷ EC_Divisor

Pool 0 예시:

= (54 × 20 × 7.6 × 10M) ÷ 8
= 82.08B ÷ 8
≈ 10.26B 객체

확장 임계값

권장:

  • 70억 객체 도달: Pool 추가 계획 시작
  • 80억 객체 도달: Pool 추가 실행
  • 여유: 20% 버퍼 유지

9.2 Thin Provisioning 활용

적용 시나리오

권장 사용:
1. ✅ Pool 병합 (In-place decommission)
2. ✅ 드라이브 교체 시 용량 증대
3. ✅ Multi-tenancy 격리

비권장 (대안 사용):
1. ❌ 단순 스토리지 추가 → 노드 추가
2. ❌ 성능 향상 목적 → NVMe 업그레이드

설정 권장

기본값:

  • DirectPV에서 Thin provisioning 활성화
  • 편의성 및 유연성 향상
  • 부작용 없음

9.3 파일 크기 최적화

작은 파일 회피

권장:

  • 최소 크기: 10MB 이상 (Iceberg 파일)
  • Compaction: 주기적으로 소형 파일 병합
  • Partitioning: 파일 분산

이유:

  • Inode 고갈 방지
  • XFS 60% 제약 회피
  • 성능 최적화

10. 액션 아이템

10.1 벤더 측 (MinIO/AIStore)

즉시 실행 (1주일 내)

  1. 용량 계획 가이드

    • 객체 수 계산 공식
    • Inode 제한 설명
    • Pool별 용량 계산기
  2. Thin Provisioning 문서

    • 설정 방법
    • Pool 병합 절차
    • 트러블슈팅 가이드
  3. 모니터링 가이드

    • Inode 사용률 체크
    • 객체 수 추적
    • 확장 임계값 Alert

단기 (2주일 내)

  1. 최적화 가이드
    • 파일 크기 권장사항
    • Partitioning 전략
    • Compaction 정책

10.2 고객사 측 (SK Hynix)

즉시 실행

  1. 현재 상태 확인

    • Pool 0 객체 수 측정
    • Inode 사용률 확인
    • 파일 크기 분포 분석
  2. Thin Provisioning 테스트

    • DirectPV 설정 확인
    • 테스트 환경 구축
    • Pool 병합 시뮬레이션
  3. 모니터링 설정

    • 객체 수 Dashboard
    • 80억 Alert (80%)
    • Inode 사용률 Graph

중기

  1. 확장 계획 수립

    • Pool 1 추가 (3월)
    • 총 20억 객체 확보
    • Pool 2 계획 (필요 시)
  2. 최적화 적용

    • Iceberg Compaction 활성화
    • 작은 파일 정리
    • Partitioning 검토

11. 핵심 결정 사항

11.1 용량 제한

Pool 0: 약 100억 개 객체

근거:

  • 54 노드 × 20 드라이브 × 7.6TB
  • EC 8:3 오버헤드
  • XFS Inode 제한

11.2 확장 전략

Pool 추가로 용량 확장

방법:

  • Pool 1 (3월): +100억 객체
  • Pool 2 (미정): +40~50억 객체 (20 노드)
  • 총 목표: 200~250억 객체

11.3 Thin Provisioning

활성화 권장

이유:
1. Pool 병합 편의성
2. 드라이브 교체 유연성
3. Multi-tenancy 지원
4. 부작용 없음

11.4 NVMe 필수

NVMe 드라이브 사용

이유:

"hard drive... you will not get 10 million files... that drive 
will stop working"
  • HDD: Inode 처리 성능 극히 저하
  • NVMe: 대량 객체 처리 가능

12. 위험 요소 및 완화 방안

12.1 Inode 고갈

위험

  • 작은 파일 대량 저장
  • XFS 25% 예약 소진
  • 용량 60%만 사용 가능

완화

  • ✅ Iceberg Compaction 활성화
  • ✅ 최소 파일 크기 10MB 권장
  • ✅ Partitioning으로 분산

12.2 객체 수 한계 도달

위험

  • Pool 0: 100억 객체 도달
  • 추가 저장 불가

완화

  • ✅ 모니터링 및 Alert (80%)
  • ✅ Pool 1 추가 (3월)
  • ✅ 확장 계획 수립

12.3 Thin Provisioning 오버커밋

위험

  • 물리 공간 초과
  • 쓰기 실패

완화

  • ✅ 사용량 모니터링
  • ✅ 적절한 오버커밋 비율 (1.5~2배)
  • ✅ Alert 설정

13. 기술적 세부사항

13.1 XFS Quota 메커니즘

Thin Provisioning 구현:

# XFS quota로 soft limit 설정
xfs_quota -x -c "limit -p bsoft=7T bhard=7T" /mount/point

특징:

  • Soft limit: 경고
  • Hard limit: 강제
  • Dynamic: 실시간 조정

13.2 DirectPV 통합

Kubernetes CSI:

  • PVC 요청 → DirectPV CSI
  • Thin provisioning 자동 적용
  • XFS quota 자동 설정

13.3 모니터링 메트릭

중요 지표:

1. Total Objects: mc admin prometheus metrics | grep minio_cluster_objects_total
2. Inode Usage: df -i
3. PV Usage: kubectl get pv
4. Physical Space: du -sh /mnt/drive

14. FAQ

Q1. 100억 객체 제한은 절대적인가?

A: 아니요. Pool 추가로 확장 가능. Pool 1 추가 시 총 200억 객체.

Q2. Thin provisioning 부작용은?

A: 물리 공간 고갈 시 쓰기 실패. 모니터링으로 예방 가능.

Q3. HDD를 사용하면 안 되는 이유는?

A: Inode 처리 성능 극히 저하. 수백만 객체도 어려움. NVMe 필수.

Q4. XFS 25% 예약은 낭비 아닌가?

A: 동적 할당. 실제로는 필요 시만 사용. 대부분 데이터 저장 가능.

Q5. Pool 병합 시 Thin provisioning 필수인가?

A: In-place 병합은 필수. 일반 Decommission은 선택 사항.


15. 결론 및 다음 단계

15.1 핵심 합의사항

  1. Pool 0: 100억 객체 (현재 확보)
  2. Pool 1 추가: +100억 객체 (3월)
  3. Thin Provisioning: 활성화 권장
  4. NVMe 필수: HDD 사용 불가
  5. 작은 파일 회피: 10MB 이상 권장

15.2 즉시 실행 항목

  • 객체 수 모니터링 Dashboard 구축
  • Thin provisioning 설정 확인
  • Iceberg Compaction 활성화

15.3 성공 기준

  • ✅ 100억 객체 안정적 저장
  • ✅ Inode 사용률 < 80%
  • ✅ Pool 병합 성공 (Thin provisioning)
  • ✅ 성능 저하 없음

문서 버전: 1.0
최종 수정일: 2026년 2월 6일
다음 리뷰: Pool 1 추가 후 (2026년 3월)


부록: 계산 공식 요약

객체 수 계산

Max Objects = (Nodes × Drives/Node × TB/Drive × 10M/TB) ÷ EC_Divisor

예시 (Pool 0):
= (54 × 20 × 7.6 × 10M) ÷ 8
= 8,208M ÷ 8
≈ 10.26B 객체

XFS 용량 계산

Raw Capacity = Nodes × Drives × TB/Drive
Usable (EC 8:3) = Raw × 0.63
Worst Case (Small Files) = Raw × 0.60

예시:
Raw = 54 × 20 × 7.6 = 8,208TB
Usable = 8,208 × 0.63 ≈ 5,171TB
Small Files = 8,208 × 0.60 ≈ 4,925TB

Thin Provisioning Over-commit

Physical Drive = 7TB
PV 1 Allocated = 7TB
PV 2 Allocated = 7TB
Total Allocated = 14TB (2× over-commit)

Constraint: PV1_Used + PV2_Used ≤ 7TB

부록: 주요 인용문

Inode 제한

"for every terabyte, you have 10 million objects... theoretically 
it's 100 million, but Linux won't let you store that"

NVMe 필수

"this is only for NVMe... If this was hard drive... you will not 
get 10 million files on a hard drive that drive will stop working"

XFS 25% 예약

"25% of the drives when you format XFS, one quarter is reserved 
for inodes... it's reserved, but it's not taken"

Thin Provisioning

"even if the drive doesn't have capacity for a new volume, you 
can allocate a volume that's not really provisioned... guaranteed 
not to exceed the capacity"

Pool 확장

"as soon as they have another Pool, they get another 4.5 billion 
[objects]... expanding the cluster will just open up the capacity"

0개의 댓글