SimWAM: A Simple World Action Model for End-to-End Autonomous Driving Paper • 2608.07468 • Published 3 days ago • 20
Modular TTT: Rethinking Test-Time Training as Composable Modules Paper • 2608.07110 • Published 3 days ago • 3
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Paper • 2608.05987 • Published 4 days ago • 88
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Paper • 2608.05000 • Published 5 days ago • 57
ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation Paper • 2608.04436 • Published 5 days ago • 55
ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment Paper • 2608.05102 • Published 5 days ago • 64
Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs Paper • 2608.01755 • Published 7 days ago • 141
Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging Paper • 2608.03316 • Published 6 days ago • 26
Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing Paper • 2608.02711 • Published 7 days ago • 86
UniWorld-Design: From Pixel Generation to Layer-Native Design Paper • 2608.03971 • Published 6 days ago • 21
AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling Paper • 2608.02602 • Published 7 days ago • 79
JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion Paper • 2608.03974 • Published 6 days ago • 90
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks Paper • 2608.01964 • Published 7 days ago • 164
WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Paper • 2607.29613 • Published 10 days ago • 27
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Paper • 2608.02023 • Published 7 days ago • 155