Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation Paper • 2609.35347 • Published 3 days ago • 143
Post-Training Leaves Behavioral Shadows on Unrelated Decisions Paper • 2609.29233 • Published 7 days ago • 262
TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces Paper • 2609.33295 • Published 4 days ago • 62
view article Article tokenizers v1: encode, decode and scaling, measured +2 ArthurZ, sbrandeis, mcpotato, lysandre • 10 days ago • 80
view article Article Training a coding model to paint watercolours with TRL and OpenEnv sergiopaniego • 28 days ago • 76
view article Article State of Open Models: Summer 2026 Observations +1 AdinaY, multimodalart, irenesolaiman • Aug 14 • 217
view article Article Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident +2 hlarcher, XciD, raphael-gl, chris-rannou • Jul 27 • 507
Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization Paper • 2607.10169 • Published Jul 11 • 14
Environment-free Synthetic Data Generation for API-Calling Agents Paper • 2607.16900 • Published Jul 18 • 21
The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence Paper • 2605.26494 • Published May 26 • 38
view article Article Harness, Scaffold, and the AI Agent Terms Worth Getting Right sergiopaniego, ariG23498 • May 25 • 148
From Context to Skills: Can Language Models Learn from Context Skillfully? Paper • 2604.27660 • Published May 3 • 74