DB가 느려짐
-> 함수가 마감 없이 기다리며 DB 커넥션을 붙잡음 <- 고리 4
-> 14분 고착
-> 클라이언트는 2.5초 뒤 "접수 완료"로 넘어감 <- 고리 5
-> 주문 70건 유실
이번 구조 변경하며 변경된 것들 정리
Retry-After: 5/api/auth/me, 장애 시 | 500 | 200 { isLoggedIn: false, degraded: true }client_request_id 멱등 키)사슬 끊기는 고리 6개를 수정 예정이였고, 해당 6개중 세개의 고리 먼저 끊는 것을 목표로 했다.
SQL 먼저 적용 후 코드 수정
client_request_id가 라이브에 없는데 코드가 먼저 나가면 주문 insert가 전부 42703 column does not exist로 실패
// src/shared/lib/supabase/timeouts.ts
export const isUpstreamError = (e: unknown): boolean -> {
if (e instanceof UpstreamError) return true
if (e instanceof Error && (e.name === 'AuthRetryableFetchError' || e.name === 'AbortError' || /fetch failed|ECONNRESET|ETIMEOUT|\b522\b/.test(e.message))) return true
if (e && typeof e === 'object' && (e as any).name === 'AuthRetryableFetchError') return true
if (...) {
if (/^(UpstreamError|AbortError|TypeError|FetchError):/.test(msg)) return true
}
if (code in ['PGRST000', 'PGRST001','PGRST002','PGRST003','PGRST004','57014']) return true
return false
}
주문 POST 만 예외로. Auth 가 죽어도 주문은 저장돼야 하므로 비로그인 (profile_id null)으로 진행. 주문은 전화번호로 나중에 계정에 연결 가능하다.
postgrest-js 2.104.0 은 읽기 요청(GET/HEAD/OPTIONS)을 기본으로 최대 3번 자동 재시도한다.
// postgrest-js 2.104.0 types/common/common.ts
export const DEFAULT_MAX_RETRIES = 3
export const RETRYABLE_STATUS_CODES = [520, 503] as const
export const RETRYABLE_METHODS = ['GET', 'HEAD', 'OPTIONS'] as const
// PostgrestBuilder.ts:90
protected retryEnabled: boolean = true
// PostgrestBuilder.ts:325-343 (요약) — fetch 가 던지면
if (fetchError?.name === 'AbortError' || fetchError?.code === 'ABORT_ERR') throw fetchError // 재시도 안 함
if (this.retryEnabled && attemptCount < DEFAULT_MAX_RETRIES) {
await sleep(getRetryDelay(attemptCount)) // 1s, 2s, 4s
continue
}
현재 fetch 래퍼는 5초가 지나면 UpstreamError 를 던진다. 이름은 AbortError가 아니니 postgrest-js는 재시도한다.
읽기 한 번 = 5s + (1s) + 5s + (2s) + 5s + (4s) + 5s = 27초
5초로 묶었다고 생각한 호출이 현재 27초 걸리고있었다.
결과는 세가지
1. Node 읽기 라우트는 503을 못 만든다: maxDuration이 10초인데 호출 하나가 27초. Vercel 이 10초에 함수를 죽이고 CORS 헤더 없는 504를 준다. 브라우저에게는 "네트워크 오류"
2. edge 라우트는 25초 한도를 넘는다. /api/auth/me 는 인증 4초 + profiles 조회 27초 = 31초.
3. 장애 중에 DB 부하가 4배가 된다.: 실패한 읽기 하나가 Supabase 에 네 번 간다. 503 + Retry-After 로 트래픽을 빼려던 설계와 정 반대다.
supabase-js 2.104.0 에는 재시도를 끄는 클라이언트 전역 설정이 없기때문에 쿼리마다 .retry(false)를 붙이는 방법뿐인데, 호출 지점이 수백 곳이다...
수정한 것은 우리가 던지는 에러에 code: 'ABORT_ERR'를 다는것.
// src/shared/lib/supabase/timeouts.ts:29-42
export class UpstreamError extends Error {
readonly retryable = true
readonly code = 'ABORT_ERR' // postgrest-js 재시도를 끄는 표식
constructor(message: string, readonly cause?: unknown) {
super(message)
this.name = 'UpstreamError'
}
}
// timeouts.ts:45, 73-76
const GATEWAY_STATUSES = new Set([502, 503, 504, 520, 521, 522, 523, 524, 525, 526, 527])
...
if (GATEWAY_STATUSES.has(res.status)) {
void res.body?.cancel().catch(() => {})
throw new UpstreamError(`upstream ${res.status}`)
}
부수 효과로 결함 하나가 같이 닫혔다. 게이트웨이가 JSON 본문으로 502 를 주면 예전 코드는 "JSON 이니 PostgREST 오류겠지" 하고 통과시켜 500 이 났다. 이제 상태 코드만 본다.
증명 - 실제 클라이언트로 호출 횟수를 센다.
// timeouts.test.ts:138-176 (요약)
it('응답이 없으면 한 번만 부르고 UpstreamError 로 끝난다', async () => {
const stub = abortAwareNeverResolvingFetch()
const supabase = createClient(URL, KEY, { global: { fetch: withTimeoutFetch(30) } })
const { error } = await supabase.from('t').select()
expect(stub).toHaveBeenCalledTimes(1)
expect(error.message).toMatch(/^UpstreamError:/)
})
수정 전에는 셋 다 4번 호출, 약 7초가 걸렸다. 수정 후에는 1번.
타임아웃은 우리 코드 한 층이 아니라 스택 전체의 성질.
우리가 fetch에 5초를 걸어도, 그 위에 있는 라이브러리가 재시도하면 5초가 아니다.
우리가 catch를 잘 짜더라도, 라이브러리가 던지지 않으면 catch가 오지 않는다.
타임아웃과 에러 경로를 설계할 때 쓰는 라이브러리가 해당 버전 소스를 열어 봐야한다.
┌ maxDuration Node 라우트: 읽기 10s / 쓰기 15s / 리뷰 30s Vercel 이 함수를 죽이는 마지막 보루
│ ┌ withDeadline 인증 4s, 라우트 읽기 8s auth-js 의 refresh 재시도 루프를 바깥에서 자른다
│ │ ┌ withTimeoutFetch 읽기 5s / 쓰기 10s Supabase 왕복 한 번
│ │ └ 넘으면 UpstreamError(code ABORT_ERR) → 재시도 없음
│ └ 넘으면 UpstreamError
└ 라우트 catch: isUpstreamError → 503 + Retry-After: 5 / 그 밖 → 500
안쪽을 먼저 끊어야 하는 이유. 안쪽을 끊게되면 우리 코드가 이유를 알고 503 을 만들게 된다.
maxDuration 먼저 끊으면 코드는 실행 중에 죽고, 고객은 CORS 헤더 없는 504를 받게 되고 로그에는 아무것도 없다.
withDeadline 이 따로 필요한 이유: Promise.race 는 진 쪽 프로미스를 취소하지 않는다.
그러므로 데드라인만 있으면 응답은 빨라지지만 안쪽 fetch 는 계속 DB 커넥션을 잡는다.
안쪽 fetch 타임아웃이 같이 있어야 커넥션이 실제로 반납된다.
환경변수 SUPABASE_TIMEOUT_MS_<이름> 으로 덮어쓸 수 있게 수정
생성부는
createSupabaseServerClient(kind): read 5s / write 10s, 쿠키 세션
createSupabasePublicClient(): 5s, 쿠키 없음, 세션 저장.갱신 없음
createSupabaseAdminClient(kind: read 5s / write 10s, 서비스 롤
supabaseAdmin 싱글턴: 10s (import 27곳 타임아웃만)
/api/auth/me — 503 대신 200 degraded./api/order/receipt-reminder — 장애가 아닌 오류는 여전히 빈 목록./api/auth/callback — 기존 에러 리다이렉트 유지. | 라우트 | 런타임 | maxDuration | 로그인 확인 503 |
|---|---|---|---|
auth/me | edge | — | degraded 200 |
my/orders | edge | — | GET 만(POST 는 비로그인 저장) |
order/lookup | edge | — | — |
alerts/subscribe, my/viewed, my/wishlist | node | 15 | 핸들러 2개 모두 |
kakao/sync-channel, my/phone, auth/withdraw | node | 15 | 있음 |
my/consultations, my/preorder-alarms, my/referrals, my/restock, my/summary | node | 10 | 있음 |
alimtalk/unsubscribe, auth/callback, consultations/phone-change, consultations/quick-quote, order/receipt-reminder | node | 15 | — |
auth/kakao, auth/logout, compare/deals, compare/devices, compare/price-trend, compare/prices, consultations/quick-quote/preview, freebies, order-counts, preorder-alarm/lookup | node | 10 | — |
reviews | node | 30 (기존, sharp) | — |
| 라우트 | 계산 | 합 | 상한 |
|---|---|---|---|
auth/me | 인증 4s + profiles 5s | 9s | edge 25s |
my/orders POST 보통 실패 | 인증 4s + 중복 조회 5s + insert 10s | 19s | edge 25s |
my/orders POST 23505 경합 | 위 + 재조회 5s | 24s | edge 25s |
order/lookup | 호출 제한 확인 + 조회 데드라인 8s + 썸네일·리뷰 5s | 약 23s | edge 25s |
const fetchPromise = fetch(`${API_URL}/api/my/orders`, { method: 'POST', keepalive: true, ... })
fetchPromise.catch(() => {})
const res = await Promise.race([
fetchPromise,
new Promise<null>((resolve) => setTimeout(() => resolve(null), 2500)),
])
if (res && !res.ok) { failSubmit(...); return }
// res === null (2.5초 안에 응답 없음)이면 여기로 온다
startTransition() // → "접수 완료"
2.5초 안에 응답이 없으면 res 는 null 이고, 코드는 성공 경로로 간다.
fetch는 취소되지 않고 계속 돈다.
서버가 나중에 504를 줘도 fetchPromise.catch(() => {})가 삼킨다.
keepalive: true 는 페이지를 떠나도 요청을 보내는 것을 보장할 뿐, 서버가 저장한 것은 보장 하지않고, 그러므로 70건의 유실이 해당 경로로 사라지게 됐다...
alter table public.online_order add column if not exists client_request_id text;
create unique index if not exists online_order_client_request_id_key
on public.online_order (client_request_id) where client_request_id is not null;
부분 인덱스라 기존행은 인덱스에 들어가지 않는다. 키를 안 보내는 옛 클라이언트도 unique에 안 걸린다.
CONCURRENTLY를 뺀 이유는 인덱스 적용기로 적어두겠다.
POST /api/my/orders { ..., clientRequestId }
│
├─ 인증 확인 (4s 데드라인)
│ 장애여도 막지 않는다 → 비로그인(profile_id null)으로 진행
│
├─ 사전 조회: 같은 client_request_id 행이 있나? (읽기 5s)
│ 있음 → { ok, receiptNo, deduped: true } ← 재시도가 여기서 끝난다
│ 장애 → 무시하고 insert 로 (정합성은 unique 인덱스가 지킨다)
│
├─ insert (쓰기 10s)
│ 성공 → { ok: true, receiptNo }
│ 23505(같은 키 동시 재시도) → 재조회 → { ok, receiptNo, deduped: true }
│ 장애 → 503
│
└─ after(): 같은 번호의 이전 신청 자동 취소 — 응답을 보낸 뒤에 돈다
deduped 응답에는 돌지 않는다(첫 저장 때 이미 돌았다)
PostgrestBuilder.ts 40줄 안에 있었다.maxDuration 과 곱해진다. 층마다 계산표를 만들고 가장 바깥 한도와 비교한다.'read' 로 둔 순간, 인자를 안 넘긴 쓰기 호출 전부가 5초 타임아웃을 물려받았다. 호출 지점마다 명시하게 했다.