Teaching LLMs to Recommend and Defer in Underrepresented Epilepsy Care Paper • 2606.31036 • Published 23 days ago • 9
HarnessBridge: Learnable Bidirectional Controller for LLM Agent Harness Paper • 2606.12882 • Published Jun 11 • 14
Guidance Contrastive Token Credit Assignment for Discrete Policy Optimization Paper • 2605.29198 • Published May 29 • 2
Less is More: Early Stopping Rollout for On-Policy Distillation Paper • 2605.27028 • Published May 26 • 15
ARLArena: A Unified Framework for Stable Agentic Reinforcement Learning Paper • 2602.21534 • Published Feb 25 • 26
Adam Improves Muon: Adaptive Moment Estimation with Orthogonalized Momentum Paper • 2602.17080 • Published Feb 19 • 3
TIC-VLA: A Think-in-Control Vision-Language-Action Model for Robot Navigation in Dynamic Environments Paper • 2602.02459 • Published Feb 2 • 4