Opus 4.6 early access 기간의 실제 모습. Harvey, bolt.new, Shopify, Lovable 4개 회사가 출시 전 모델 테스트한 내부 이야기. "war rooms, late nights, Slack lighting up" — 마케팅 X 진솔한 backstage. "polished output 뒤에 messier process" 의 공개.
본문 인용:
"They work with pre-production research models, test them against real workloads to figure out what the model is great at, where it breaks, and whether it's ready to ship to their own users the moment Anthropic launches it publicly."
(pre-production 모델 사용, 실제 워크로드 테스트, 강점·약점·출시 가능성 평가)
"Their honest assessments — what works and what doesn't — directly shape the version of the model Anthropic ultimately ships."
(솔직한 평가가 최종 출시 버전 형성)
리뷰 윈도우:
테스트 영역:
결과:
본문 강조:
"When your structured evals and your subject matter experts are both saying the same thing, that's a strong signal."
(구조화 eval + 도메인 전문가 = 같은 말 → 강한 시그널)
테스트 방법:
구체 사례:
Garrett Serviss (bolt.new VP of Marketing):
"Opus 4.6 diagnosed bugs on the first try that we'd failed to fix across five-plus attempts with previous models. The jump in reasoning depth is real."
Ben Lafferty (Shopify Assistants Staff Engineer) 사례:
방법:
"Instruction following is significantly improved. This was one of the first early access periods where I haven't had substantial feedback to give."
(이전 early access보다 피드백 거의 없음 — 그만큼 좋다)
"For me, Opus 4.6 is the first model from Anthropic that feels like a true collaborator in my day-to-day work. The time horizon of tasks that I can hand off to the model continues to grow."
테스트 두 트랙:
1. Design benchmarks + complex task evals (구조)
2. "Vibe checks" — 엔지니어가 새 모델로 앱 빌드 (직관)
Fabian Hedin (Lovable co-founder):
"Claude Opus 4.6 is an uplift in design quality. It's more autonomous, which is core to Lovable's values. People should be creating things that matter, not micromanaging AI."
Alexandre Pesant:
"It's always a bit of a race to discover the new rough edges."
(새 거친 부분 찾기 = 항상 경주)
본문 강조:
"Of course not all of the feedback was glowing, and that's the point."
피드백 = 솔직, 비판 포함:
이게 "early access의 진짜 목적" . 마케팅 X, 제품 개선.
이 글이 "마케팅 글이지만 마케팅 안 함" 의 정석이다.
전통 마케팅:
이 글:
이 "vulnerability marketing" 이 enterprise 신뢰의 자산이다:
비교 — 다른 AI 회사:
Anthropic:
이 글의 가장 중요한 사실 — Anthropic이 출시 전 모델을 외부 회사와 공유.
이게 의미하는 것:
리스크:
그러나 Anthropic의 결정:
이게 "win-win 모델" 의 정석이다. 단기 보호보다 장기 생태계.
Harvey = AI legal 가장 큰 startup ($1B+ 밸류).
BigLaw Bench = legal AI 표준 평가.
90.2% = 새 표준.
이 결과의 산업 임팩트:
비교:
본문이 시사하는 답:
이게 LawSites의 "foundation model vs legal-tech" 동학 (#84 글).
bolt.new의 waterfall graph 버그 사례:
이 사례가 "AI = Senior Engineer" 시그널이다:
비교 — 일반 디버깅:
Opus 4.6:
이게 "AI ≥ 시니어" 의 발견 시그널이다.
Lovable의 두 트랙 평가:
1. Structured benchmarks (객관)
2. Vibe checks (주관)
이 "vibe check" 이 새 평가 카테고리다:
전통 평가:
Vibe check:
이게 "AI 평가의 새 dimension" 이다. 정량 + 정성 = 전체.
Ben Lafferty 인용:
"The time horizon of tasks that I can hand off to the model continues to grow."
(모델에 위임 가능한 task 시간 horizon 계속 증가)
이게 "AI 자율성" 의 정확한 측정이다.
데이터 (METR Wikipedia 인용):
이 의미:
비교 — 1년 전:
이 진화가 "AI = 진짜 동료" 의 시그널이다.
이 글의 customer 인용:
이게 mutual marketing의 효과:
비교 — 일반 사례:
Anthropic 모델:
본문 강조:
"Their honest assessments — what works and what doesn't — directly shape the version of the model Anthropic ultimately ships."
이게 "고객이 모델 형성" 의 시그널이다:
비교 — 전통 model 출시:
새 패턴:
이 "피드백 루프" 가 enterprise B2B의 정석이다. 그러나 AI 모델에 적용은 새로움.
이 글은 "behind the launch" 같지만, 실제로는 AI 시대 product launch의 새 표준 정의다.
2026년 2월 5일 시점은 "AI 모델 = 회사 단독 출시" 시대가 끝난 시점이다. 고객 + 회사 = 공동 개발의 정석.
흥미로운 건 이 글이 "Opus 4.6 출시 동시 출시" 라는 점이다. 이 동시 출시의 의미:
이 "기능 + 신뢰" 의 동시 마케팅이 enterprise 결정 가속:
다음 글 (#87): Claude Enterprise self-serve — 이 신뢰 기반의 다음 단계, 고객이 영업 없이 enterprise 직접 구매. 신뢰 + 능력 + 자가 도입 = AI 시대 SaaS의 정석. 점차 "Anthropic 정복 청사진" 의 모든 layer가 보인다.