Resources for "Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs"
Xin Lai
xinlai
AI & ML interests
Multimodal LLM, LLM Reasoning, Point Cloud Segmentation, Image Segmentation
Recent Activity
upvoted a paper about 1 month ago
Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models liked a model 3 months ago
tencent/Hy3 upvoted a paper 3 months ago
Training Open Models for Agentic Phone UseOrganizations
None yet