-
Efficient RLHF: Reducing the Memory Usage of PPO
Paper • 2309.00754 • Published • 16 -
Statistical Rejection Sampling Improves Preference Optimization
Paper • 2309.06657 • Published • 15 -
Aligning Large Multimodal Models with Factually Augmented RLHF
Paper • 2309.14525 • Published • 32 -
Stabilizing RLHF through Advantage Model and Selective Rehearsal
Paper • 2309.10202 • Published • 11
Dong Li
dongleecsu
·
AI & ML interests
None yet
Recent Activity
liked a Space 9 days ago
agent-memory-leaderboard/leaderboard updated a collection almost 3 years ago
RLHF papers updated a collection almost 3 years ago
RLHF papersOrganizations
None yet