[논문 리뷰] Generative Agent Simulations of 1,000 People

2한나·2026년 2월 6일

Generative Agent Simulations of 1,000 People
Park, J. S., Zou, C. Q., Shaw, A., Hill, B. M., Cai, C., Morris, M. R., ... & Bernstein, M. S. (2024). Generative agent simulations of 1,000 people. arXiv preprint arXiv:2411.10109.

Abstract

  • 연구 목적
    • 다양한 domains에 걸쳐 인간의 행동을 복제하는 general-purpose computational agents를 만드는 것
  • 연구 방법
    • 참가자 1,052명의 삶에 대한 qualitative interviews에 large language models를 적용한 뒤, 이 agents들이 자신이 대변하는 개인의 태도와 행동을 얼마나 잘 복제하는지 측정
  • 연구 결과
    • generative agents는 General Social Survey(GSS) 응답에서 실제 인간이 2주 뒤 자신의 답변을 재현하는 수준의 85% 정확도를 보여줌
    • personality traits를 파악하거나 다양한 experimental replications의 결과를 예측하는 데 있어 인간과 대등한 성능을 기록함
    • 단순히 demographic descriptions(나이, 성별 등)만 입력한 agents보다, 이번 연구의 architecture가 racial and ideological groups의 accuracy biases 를줄임

Main Text

  • 연구의 의의
    • 다양한 social, political, or informational contexts를 넘나드는 general-purpose simulation을 통해 연구자들이 다양한 interventions 및 theories를 테스트할 수 있는 실험실 역할을 수행함
    • 개별 에이전트들이 collectives로 결합될 때, 경제·사회·정치학 등 다양한 분야의 institutions and networks 구조에 대한 이해를 확장하고 복잡한 causal and contextual interactions를 포착할 수 있음.
    • 실제 응용 사례로 새로운 보건 정책, 신제품 출시 반응, 혹은 major shocks에 대한 사회적 반응 등을 미리 시뮬레이션해 볼 수 있음
  • 기존 모델의 한계 -> LLMs
    • Traditional agent architectures: interpretability는 높지만 수동으로 지정된 행동(manually specified behaviors)에 의존하기 때문에, 컨텍스트가 제한적이고 실제 인간 행동을 지나치게 단순화함
    • LLMs: 인간 행동에 대한 방대한 지식을 갖춘 large language models는 다양한 컨텍스트에서 행동을 정확히 시뮬레이션할 수 있는 새로운 가능성을 제공함
    • 해결해야하는 것: 에이전트가 단순한 demographic stereotypes로 고착되는 것을 방지하고, 평균 이상의 정교한 측정이 필요함
  • 연구 접근 방식
    • 2시간 분량의 qualitative interviews를 large language model과 결합하여 1,000명 이상의 실제 개인을 복제하는 generative agent architecture를 구축함
    • General Social Survey(GSS), Big Five Personality Inventory, 5가지 behavioral economic games, 5개의 social science experiments를 통해 성능을 검증함

Creating 1,000 Generative Agents of Real People

  • 데이터 수집 방법론 (Data Collection Methodology)
    - 개인의 태도와 행동에 영향을 미치는 idiosyncratic한 요인들을 포착하기 위해 in-depth interviews 방식을 채택함
    - 미국 인구 통계를 반영하여 age, gender, race, region, education, and political ideology 등을 기준으로 stratified sample된 1,052명의 참가자를 모집함
    - 연구진이 개발한 AI interviewer가 semi-structured interview protocol에 따라 2시간 동안 음성 인터뷰를 진행했으며, 인당 평균 6,491 단어 분량의 transcripts를 생성함
    - 특정 지표에 치우치지 않도록 사회학 프로젝트인 American Voices Project의 기존 인터뷰 프로토콜을 활용하여 인생 이야기부터 사회적 이슈까지 폭넓게 다룸
    - e.g., “Tell me the story of your life—from your childhood, to education, to family and relationships, and to any major life events you may have had”
    - e.g., “How have you responded to the increased focus on race and/or racism and
    policing?

  • Agent Architecture

    • 인터뷰 전체 transcript를 model prompt에 주입하여, 해당 인물의 데이터를 기반으로 모델이 그 사람을 모사(imitate)하도록 설계함
    • 여러 단계의 의사결정이 필요한 실험의 경우, 에이전트에게 이전 질문-응답 대한 memory를 텍스트 형태로 제공함
    • 결과적으로 에이전트는 forced-choice prompts, surveys, surveys, and multi-stage interactional settings 기반 stimulus에 응답할 수 있음.
  • Evaluation and Validation

    • 평가 구성
      • the core module of the General Social Survey
      • the 44-item Big Five Inventory (BFI-44)
      • five well-known behavioral economic games (including the dictator game, trust game, public goods game, and prisoner’s dilemma)
      • five social science experiments with control and treatment conditions
    • Self-consistency 측정
      • 사람의 응답 가변성을 고려하여, 참가자들이 2주 간격으로 두 번 설문에 참여하게 함으로써 개인별 응답 일관성을 측정함
    • Normalized accuracy
      • 에이전트의 예측 정확도를 참가자의 재현 정확도(replication accuracy)로 나누어 계산함
      • normalized accuracy가 1.0에 도달하면, 에이전트가 그 사람의 응답을 예측하는 능력이 본인이 2주 후 자신의 답변을 다시 맞히는 것만큼 정확하다는 것을 뜻함

Predicting Individuals’ Attitudes and Behaviors

  • Comparison Baselines 설정

    • 목적: 인터뷰가 generative agents의 예측 정확도에 기여하는 정도를 평가하기 위해 두 가지 방식과 비교함
    • demographic-based agents: GSS 응답을 활용하여 age, gender, race, and political ideology 등 demographic attributes 정보를 에이전트에 제공함
    • persona-based agents: 인터뷰 후 참가자들이 직접 작성한 personal background, personality, demographic details 정보를 포함한 짧은 요약 문단을 사용함
  • GSS (General Social Survey) 결과

    • 평가 지표: 177개의 핵심 질문을 벤치마크로 사용 (random 확률 27.03%).
    • 결과
      • interview-based: normalized accuracy 0.85
      • baseline 비교
        • demographic-based: 0.71
        • persona-based: 0.70
          -> generative agents가 약 14~15 포인트 더 높음
  • Big Five Personality Traits (BFI-44) 결과

    • 평가 방식: 5가지 성격 차원(openness, conscientiousness, extraversion, agreeableness, neuroticism)에 대한 응답 예측
    • 결과
      • interview-based: normalized correlation 0.80
      • participants: 0.85
      • baseline 비교
        • demographic-based: 0.55
        • persona-based: 0.75
  • Economic Games 결과

    • 평가 항목: Dictator Game, Trust Games, Public Goods Game, Prisoner’s Dilemma 등 5가지 게임
    • interview-based: normalized correlation 0.66
  • Composite Agents와의 비교

    • Composite agents: 인터뷰 데이터 대신 참가자의 GSS, Big Five, Economic Games 응답 데이터를 조합하여 만든 베이스라인
    • 성능 비교
      • GSS: Composite(0.76) vs Interview-based(0.85)
      • Big Five: Composite(0.64) vs Interview-based(0.80)
      • Economic Games: Composite(0.31) vs Interview-based(0.66)
    • 결론: 평가 지표와 의미적으로 유사한 설문 데이터를 직접 제공한 composite agents보다도 interview-based agents의 예측력이 더 높게 나타남. 이는 인터뷰가 개인의 정체성을 포착하는 데 독보적인 효과가 있음을 시사함
  • Impact of Content Volume - Ablation Test

    • 인터뷰 transcript의 80%(전체 120분 중 약 96분 분량)를 무작위로 제거하는 ablating 실험을 진행함
    • 정보량이 대폭 감소했음에도 불구하고, interview-based generative agents는 GSS에서 평균 normalized accuracy 0.79를 기록함
  • Linguistic Cues vs. Knowledge

    • 인터뷰의 예측력이 linguistic cues(특유의 말투 등 언어적 스타일)에서 오는지, 아니면 지식의 깊이에서 오는지 확인하기 위해 "interview-summary" agents를 생성함
    • GPT-4o를 사용하여 인터뷰 원문을 핵심 응답 쌍 위주의 bullet-pointed summaries로 변환함으로써, factual content(사실 관계)는 유지하되 본래의 언어적 특징은 제거함
    • 결과
      • GSS에서 normalized accuracy 0.83 달성
        -> 이는 composite agents의 성능을 능가하며, 말투보다는 인터뷰를 통해 얻은 구체적인 정보 자체가 중요함을 시사함
  • 결론
    when informing language models about human behavior, interviews are more effective and efficient than survey - based methods.

0개의 댓글