
생성형 대규모 언어 모델(LLM)은 현재 실생활에서 널리 사용되고 있으며 특히 글쓰기 작업에서 협업 도구로 활용되고 있다.지금까지는 주로 인간-LLM 협업에 초점이 맞춰져 있었다.vLLM, LangChain, HuggingFace와 같은 unifying framewor

LLMs pose potential risks like generating unethical contentAssessing LLMs' values helps expose misalignmentReference-free evaluators (fine-tuned LLMs
LLM 기반 의사결정 에이전트가 다양한 분야에 배포되고 있음일반적인 정렬(alignment) 방법은 RLHF나 DPO와 같이 인간 선호 데이터에 기반함기존 방법은 암묵적으로 가치를 표현하며 데이터 수집 비용이 크고 편향될 수 있음본 연구는 핵심 인간 가치를 명시적으로

Webis-ArgValues-22를 확장Touche-ValueEval4780개 arguments → 9324 argumentssix diverse source : 종교 텍스트, 커뮤니티 토론, 자유 텍스트 주장, 신문 사설, 정치적 토론54개의 인간 가치로 3명의 크라

LLMs의 general language tasks는 많이 성능이 올라왔지만, complex reasoning tasks에서는 아직 한계를 보인다.Self-reflection (자기 반성) 과 같은 방법론들이 제시되었지만, 여전히 한계가 명확히 존재한다. 이 한계를 저

value LLM 태초격 논문인간처럼 가치 기반 대화를 하는 에이전트의 필요성대규모 인간 가치 데이터셋인 VALUENET21,374개의 텍스트 시나리오에 대한 인간의 태도를 포함intercultural research에서의 기본 인간 가치 이론에 부합하는 10-dim으

인간과 인공지능의 간극을 좁힌다LLM 내에서 성격, 기질, 감정 등의 특성이 나타날 가능성에 대한 질문을 제기한다.PsychoBenchpersonality traitsinterpersonal relationshipsmotivational testsemotional ab

인간 지능은 cognitive synergy라는 것을 활용함. 서로 다른 마음이 협업할 때 superior outcomes를 낸다.이에 착안하여 단일 LLM을 cognitive synergist로 변환하는 Solo Performance Prompting (SPP) 제안

LLM이 일부 financial tasks에서 잘하는 모습을 보이고 있지만, sequential financial decision making 에서는 volatile environment와 intelligent risk management의 필요성 때문에 한계가 있다.

Harmless AI assistant through self-improvement, without any human labels identifying harmful outputsThe only human oversight is provided through a lis

preference alignment가 중요해졌지만, human-annotated preference dataset이 비싸다.seed preference dataset을 활용하여 iterative 하게 response를 생성하고 preference annotation을

AI safety의 중요한 문제중 하나는 human moral mind 에 있는 flexibility를 capture하는 문제이다.인간은 단순히 정해진 규칙만을 무조건 strict하게 따르는 게 아니라 예외를 두어 규칙을 지키지 않기도 한다.moral exception

The paper introduces DAILYDILEMMAS, a dataset of 1,360 moral dilemmas from everyday lifeUsers increasingly seek guidance from LLMs for daily life deci

Relies on group-level distributional inferences rather than individual-level variationExisting methods pre-specify diversity through coarse categories
Problem: Standard LLM evaluations use minimal contexts, but deployed LLMs face diverse contexts that can drastically change their behavior and express