Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
Kaixiang Zhao
kzhao5
2
6
1
Follow
YT6556's profile picture
1 follower
·
5 following
https://kzhao5.github.io/
kzhao5
AI & ML interests
None yet
Recent Activity
authored
a paper
4 days ago
Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It
upvoted
a
paper
4 days ago
An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning
upvoted
a
paper
4 days ago
Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It
View all activity
Organizations
None yet
kzhao5
's models
8
Sort: Recently updated
kzhao5/opd-baselines-pt06b
Updated
8 days ago
kzhao5/mea-soda-defenses
Updated
12 days ago
kzhao5/mea-qedks-defenses
Updated
12 days ago
kzhao5/mea-seqkd-defenses
Updated
12 days ago
kzhao5/opd-rethink-k2-horizon-7b-semantic-class-step200
9B
•
Updated
24 days ago
•
56
kzhao5/opd-rethink-gemma3-4b-semantic-class-step200
5B
•
Updated
26 days ago
•
31
kzhao5/opd-rethink-llama3.2-3b-semantic-class-step200
4B
•
Updated
26 days ago
•
34
kzhao5/opd-rethink-qwen3-1.7b-semantic-class-step200
2B
•
Updated
26 days ago
•
36