RobotWorld: Benchmarking Multimodal Agents for Robot Use Across Diverse Tasks and Embodiments Paper • 2610.10409 • Published 1 day ago • 23
Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents Paper • 2610.01892 • Published 7 days ago • 15
DuoMatching: Joint-Marginal Distribution Matching for Few-Step Video Generation Paper • 2610.03543 • Published 6 days ago • 76
Prism: Dynamic Sparse Attention for Native 2K Joint Video-Audio Generation Model Training Paper • 2610.05416 • Published 4 days ago • 10
MotionInsight: Diagnosing Object Motion Deficiencies in Generated Videos Paper • 2609.37030 • Published 9 days ago • 1
OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction Paper • 2610.01762 • Published 7 days ago • 233
Agentic Visual Generation: From Generative Models to Agentic Control Paper • 2609.06758 • Published Sep 6 • 34
VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation Paper • 2608.18607 • Published Aug 19 • 12
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Paper • 2608.14391 • Published Aug 14 • 287