X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras Paper • 2607.12993 • Published 6 days ago • 131
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable Paper • 2607.13285 • Published 6 days ago • 201
ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation Paper • 2607.08741 • Published 11 days ago • 11
RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation Paper • 2607.06559 • Published 13 days ago • 94
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Paper • 2606.29526 • Published 22 days ago • 168
SpheRoPE: Zero-Shot Optimization-Free 360 Panorama Generation with Spherical RoPE Paper • 2606.32033 • Published 20 days ago • 7
Focusing on What Matters: Saliency-Harnessing Accurate Routing for Diffusion MoE Paper • 2606.26938 • Published 25 days ago • 6
OpenRath: Session-Centered Runtime State for Agent Systems Paper • 2606.19409 • Published Jun 17 • 78
One Click per Cell Type Suffices: Training-free Group Interaction for Cell Instance Segmentation Paper • 2605.29429 • Published May 28 • 8
On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters Paper • 2606.02437 • Published Jun 1 • 239
How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning Paper • 2605.27310 • Published May 26 • 20
SQuTR: A Robustness Benchmark for Spoken Query to Text Retrieval under Acoustic Noise Paper • 2602.12783 • Published Feb 13 • 246
ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling Paper • 2603.25746 • Published Mar 26 • 155
SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models Paper • 2603.16859 • Published Mar 17 • 248