From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models Paper • 2608.06020 • Published 22 days ago • 35
Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements Paper • 2607.28661 • Published Jul 22 • 17
ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog Paper • 2607.04438 • Published Jul 5 • 64
ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes Paper • 2607.04439 • Published Jul 5 • 63
MonkeyOCRv2: A Visual-Text Foundation Model for Document AI Paper • 2607.11562 • Published Jul 13 • 80
Uncertainty-Aware World Model for Aerial Image-Goal Navigation Paper • 2608.05597 • Published 22 days ago • 12
When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles Paper • 2607.23379 • Published Jul 25 • 18
Small Foundation Models of Human Cognition and Behaviour Paper • 2608.05224 • Published 19 days ago • 23
Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events Paper • 2608.06485 • Published 22 days ago • 4
MBA: Multimodal Benchmark and Agents for Real-World Business Ideation Paper • 2608.11616 • Published 16 days ago • 6
Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop Paper • 2608.11215 • Published Jul 19 • 6
AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale Paper • 2608.20634 • Published 7 days ago • 11
Hydra-0: Action Flow for Generalist World Modeling and Control Paper • 2608.18077 • Published 10 days ago • 8
Human-Centric Intelligence in the Era of Foundation Models: A Survey Paper • 2608.18184 • Published 10 days ago • 9
Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs Paper • 2608.12781 • Published 11 days ago • 33
Hadith computational science in the age of large language models: a critical narrative review Paper • 2608.20364 • Published Jun 18 • 2
Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence Paper • 2608.21156 • Published 7 days ago • 54
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 4 days ago • 197
MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks Paper • 2608.23035 • Published 4 days ago • 41
Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization Paper • 2608.23311 • Published 4 days ago • 16
One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows Paper • 2608.19741 • Published 8 days ago • 12
Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection Paper • 2608.20169 • Published 4 days ago • 10
LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks Paper • 2608.23200 • Published 4 days ago • 6
Industrial-Instruction: An End-to-End Framework for Building Instruction-Tuning and Benchmark Datasets from Industrial Technical Reports Paper • 2608.22817 • Published 4 days ago • 5
One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders Paper • 2606.13610 • Published 4 days ago • 5
Same Agent, Different Answers: A Repeat-Aware Audit of Corpus-Induced Answer Churn in Retrieval-Augmented QA Paper • 2608.22856 • Published 4 days ago • 4