SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring Paper • 2608.09802 • Published 1 day ago • 75
The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows Paper • 2608.06714 • Published 4 days ago • 4
HelloWorld: Enabling Socially Interactive Characters in Video World Models Paper • 2608.05070 • Published 6 days ago • 35
OPD-V: Visual On-Policy Self-Distillation with Modality Balance Paper • 2608.05131 • Published 5 days ago • 10
view article Article Training a coding agent using the OpenCode harness in remote HF sandboxes with TRL and OpenEnv sergiopaniego • 6 days ago • 19
AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling Paper • 2608.02602 • Published 8 days ago • 79
UniWorld-Design: From Pixel Generation to Layer-Native Design Paper • 2608.03971 • Published 7 days ago • 21
Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging Paper • 2608.03316 • Published 7 days ago • 26
JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion Paper • 2608.03974 • Published 7 days ago • 90
Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing Paper • 2608.02711 • Published 8 days ago • 88
SkillJack: Persistent Skill Backdoors in Self-Evolving Agents Paper • 2608.03509 • Published 7 days ago • 23
Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures Paper • 2607.28802 • Published 12 days ago • 10
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks Paper • 2608.01964 • Published 8 days ago • 165
Roomer: Reflective Object-Grounded Model Editing and Repair for 3D Indoor Layout Synthesis Paper • 2608.01973 • Published 8 days ago • 16