Gaze as Evidence for Common Grounding: A Cross-Corpus Analysis of MapTask and MUNDEX Paper • 2609.18011 • Published 3 days ago • 22
EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents Paper • 2609.05903 • Published 14 days ago • 63
FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow Estimation Paper • 2609.11486 • Published 9 days ago • 34
Procedural Graphs: Self-Evolving Execution Structures for LLM Agents Paper • 2609.09153 • Published 11 days ago • 40
RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests Paper • 2608.27831 • Published 19 days ago • 33
Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training Paper • 2608.26730 • Published 23 days ago • 154
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills Paper • 2609.02749 • Published 17 days ago • 399
Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving Paper • 2609.00111 • Published 19 days ago • 311
Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See Paper • 2608.17744 • Published Aug 18 • 16
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Paper • 2608.17426 • Published Aug 18 • 159
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Paper • 2608.14391 • Published Aug 14 • 283