Learning on the Fly: Rapid Policy Adaptation via Differentiable Simulation 간단 리뷰

신희준·2026년 2월 2일

[Learning on the Fly: Rapid Policy Adaptation via Differentiable Simulation](Learning on the Fly: Rapid Policy Adaptation via Differentiable Simulation) (Jiahe Pan, Jiaxu Xing, Rudolf Reiter, Yifan Zhai, Elie Aljalbout, Davide Scaramuzza / RAL, 2025)

Visual feature based policy + differential simulation + sim2real fast adaptation (residual learning)

Problem

  1. sim2real gap 문제
  2. Domain randomization은 ood disturbance에 약함
  3. real2sim2real : 실세계 데이터 활용해서 시뮬레이터를 현실처럼 맞게 고치고, 그 시뮬레이터에서 정책 다시 학습→ 너무 오래 걸림

Core Idea

미리 모든 환경을 커버하려 하지 말고 상황에 맞게 즉석에서 fitting해서 쓰자

method

  • Policy Pretraining + Online Adaptation

    • Policy pretraining에서는 간단한 dynamics 를 통해 base policy 학습
  • Online Adaptation에서는 아래 세 루프를 동시에 돌림
    1. Real world rollout
    - 실제 드론 비행에서 trajectory 얻음
    2. Residual Dynamic Learning
    - 간단한 dynamic model이 맞추지 못한 가속도 residual 학습
    → 이건 supervised learning
    - state, action 주어졌을 때 실제로 드론의 가속도와 계산되는 가속도를 비교
    3. Policy Training
    - Residual network가 주는 residual을 통해 dynamics가 정확해지고, 이를 기반으로 하는 differentiable simulator에서 policy 학습

➡️ 5초 정도면 adaptation이 끝난다고 함.


💡Low-fidelity vs. High-fidelity Dynamics model

  • 동역학이 정확해질수록 gradient 계산이 복잡해짐
  • low-fidelity 모델도 gradient 방향은 대충 맞음 → 성능 차이 거의 없음.
  • low-fidelity 모델에 residual 보정 해주면 충분히 쓸만한 forward model 설계 가능
  • 즉, 여기서 differentiable simulator에서 활용되는 dynamic model은 policy를 미분하기 위한 surrogate model이지, 실제 물리를 그대로 재현하기 위한 simulator가 아님.
profile
공부하고 싶은 사람

0개의 댓글