-
FlowRL: Matching Reward Distributions for LLM Reasoning
Paper • 2509.15207 • Published • 118 -
Kwaipilot/KAT-Dev-72B-Exp
Text Generation • 73B • Updated • 24 • 157 -
Agentic Entropy-Balanced Policy Optimization
Paper • 2510.14545 • Published • 107 -
Multi-Agent Deep Research: Training Multi-Agent Systems with M-GRPO
Paper • 2511.13288 • Published • 19
Malkesh Dalia
malkesh2911
·
AI & ML interests
None yet
Recent Activity
upvoted a paper about 5 hours ago
Emergent Social Intelligence Risks in Generative Multi-Agent Systems upvoted a paper 1 day ago
Gen-Searcher: Reinforcing Agentic Search for Image Generation upvoted a paper 4 days ago
Vega: Learning to Drive with Natural Language InstructionsOrganizations
None yet