Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge Paper • 2608.28478 • Published 15 days ago • 20
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model Paper • 2607.17977 • Published Jul 20 • 199
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Paper • 2607.17423 • Published Jul 19 • 167
AsySplat: Efficient Asymmetric 3D Gaussian Splatting for Long-Sequence Scene Modeling Paper • 2607.10995 • Published Jul 13 • 10
Phone Segmentation and Recognition through Phonological Activation Mapping Paper • 2607.09020 • Published Jul 10 • 9
When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing Errors Paper • 2606.32029 • Published Jun 30 • 15
MemSlides: A Hierarchical Memory Driven Agent Framework for Personalized Slide Generation with Multi-turn Local Revision Paper • 2606.17162 • Published Jun 15 • 178
RhymeFlow: Training-Free Acceleration for Video Generation with Asynchronous Denoising Flow Scheduling Paper • 2606.06309 • Published Jun 4 • 12
Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models Paper • 2606.03988 • Published Jun 3 • 127
StreamChar: Long-Horizon Streaming Character Audio-Video Generation with Decoupled Orchestration Paper • 2605.25659 • Published May 25 • 19
RoboSemanticBench: Diagnosing Semantic Grounding in Action Prediction for VLA Models Paper • 2606.02277 • Published Jun 1 • 8