VIEScore2: Unified Image Evaluation with Spatially Grounded Explanations Paper • 2610.00994 • Published 11 days ago • 26
World Editing: Intervening on Executable Worlds at Increasing Depth Paper • 2610.02331 • Published 11 days ago • 30
World Editing: Intervening on Executable Worlds at Increasing Depth Paper • 2610.02331 • Published 11 days ago • 30
TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining Paper • 2609.33419 • Published 15 days ago • 19
Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents Paper • 2609.39982 • Published 12 days ago • 120
SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning Paper • 2606.13673 • Published Jun 11 • 115 • 5
TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining Paper • 2609.33419 • Published 15 days ago • 19
AI for Media Collection NVIDIA AI for Media is a collection that enhance audio, video, and augmented reality effects for media and entertainment workflows • 2 items • Updated 24 days ago • 8
ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding Paper • 2609.07941 • Published Sep 7 • 19
Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction Paper • 2609.04201 • Published Sep 3 • 49
ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding Paper • 2609.07941 • Published Sep 7 • 19
ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding Paper • 2609.07941 • Published Sep 7 • 19
Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction Paper • 2609.04201 • Published Sep 3 • 49
PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration Paper • 2608.21031 • Published Aug 21 • 4
PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration Paper • 2608.21031 • Published Aug 21 • 4
PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration Paper • 2608.21031 • Published Aug 21 • 4