FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth
Paper • 2608.20574 • Published • 3
None defined yet.
When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models
AURA: Action-Gated Memory for Robot Policies at Constant VRAM