OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining Paper • 2609.07398 • Published 3 days ago • 48
Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds Paper • 2608.23383 • Published 17 days ago • 19
EchoWM: Open and Enterable Omnimodal World Models Paper • 2608.23189 • Published 17 days ago • 79
view post Post 118 Omni-Rewriter Replay: feed it a local clip, get a formatted H3 prompt, then generate if you want.Gallery: Wayne-King/omni-rewriter-replayCode: https://github.com/WayneJin0918/Omni-Rewriter See translation 👍 1 1 + Reply
SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs Paper • 2603.20253 • Published Mar 11 • 2
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents Paper • 2604.23781 • Published Apr 26 • 34
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos Paper • 2603.22529 • Published Mar 23 • 7
VisionCoach: Reinforcing Grounded Video Reasoning via Visual-Perception Prompting Paper • 2603.14659 • Published Mar 15 • 6
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning Paper • 2602.08236 • Published Feb 9 • 9
WildRayZer: Self-supervised Large View Synthesis in Dynamic Environments Paper • 2601.10716 • Published Jan 15 • 5
WildRayZer: Self-supervised Large View Synthesis in Dynamic Environments Paper • 2601.10716 • Published Jan 15 • 5
SimWorld: An Open-ended Realistic Simulator for Autonomous Agents in Physical and Social Worlds Paper • 2512.01078 • Published Nov 30, 2025 • 34
Next-Embedding Prediction Makes Strong Vision Learners Paper • 2512.16922 • Published Dec 18, 2025 • 91
SixAILab/nepa-large-patch14-224-sft Image Classification • 0.3B • Updated Dec 21, 2025 • 16 • 1
SixAILab/nepa-base-patch14-224-sft Image Classification • 86.3M • Updated Dec 21, 2025 • 130 • 4