arxiv:2510.25110
Ruixuan Tu
TURX
AI & ML interests
LLMs & NLP
Recent Activity
liked a model about 10 hours ago
moonshotai/Kimi-K3 authored a paper 7 months ago
DEBATE: A Large-Scale Benchmark for Role-Playing LLM Agents in Multi-Agent, Long-Form Debates authored a paper about 1 year ago
FaithBench: A Diverse Hallucination Benchmark for Summarization by
Modern LLMs