OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue Paper β’ 2609.21465 β’ Published 7 days ago β’ 143
Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation Paper β’ 2609.11638 β’ Published 15 days ago β’ 699
StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling Paper β’ 2608.15089 β’ Published Aug 15 β’ 448
Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent Paper β’ 2608.03979 β’ Published Aug 4 β’ 54