Theory-Grounded and Culture-Aware Multilingual Moral Reasoning
AI & ML interests
Factuality, reasoning, alignment, LLM applications
Recent Activity
View all activity
Papers
MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning
Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation
spaces 7
Running
LudoBench
🎲
Multimodal Game Reasoning Benchmark [ICLR 2026]
Sleeping
Agents
Answer Convergence Early Stopping
🛑
Demo for EMNLP Paper "Answer Convergence as a Signal..."
Runtime error
FactRBench
🏆
View and analyze long-form factuality leaderboard
Running
3
ExpertLongBench
🚀
Leaderboard for ExpertLongBench
Sleeping
1
ManyICLBench
🚀
Leaderboard for ManyICLBench
Running
MLRC-BENCH
📊
Display model performance rankings
models 15
launch/MET-D-Gemma3-4B-en-only
Text Generation • 4B • Updated • 1.02k
launch/MET-D-Gemma3-4B
Text Generation • 4B • Updated • 1.03k
launch/MET-D-Qwen3-8B-en-only
Text Generation • 8B • Updated • 1.02k
launch/MET-D-Qwen3-8B
Text Generation • 8B • Updated • 1.03k
launch/MET-D-Qwen3-4B-zh-only
Text Generation • 4B • Updated • 1.01k
launch/MET-D-Qwen3-4B-ms-only
Text Generation • 4B • Updated • 990
launch/MET-D-Qwen3-4B-ko-only
Text Generation • 4B • Updated • 981
launch/MET-D-Qwen3-4B-hi-only
Text Generation • 4B • Updated • 959
launch/MET-D-Qwen3-4B-es-only
Text Generation • 4B • Updated • 967
launch/MET-D-Qwen3-4B-en-only
Text Generation • 4B • Updated • 952
datasets 14
launch/MCLASH
Viewer • Updated • 2.61k • 125
launch/CLASH
Viewer • Updated • 345 • 73 • 3
launch/thinkprm-1K-verification-cots
Viewer • Updated • 1k • 125 • 8
launch/LudoBench
Viewer • Updated • 638 • 37
launch/ExpertLongBench
Preview • Updated • 127 • 10
launch/ManyICLBench
Viewer • Updated • 66 • 458 • 1
launch/CMV
Viewer • Updated • 133 • 22
launch/FactRBench
Viewer • Updated • 1.06k • 25 • 2
launch/FactBench
Viewer • Updated • 1k • 60 • 3
launch/gov_report
Viewer • Updated • 58.4k • 656 • 14