๐ง Remember-R1: Our fix for MLLMs forgetting the image during long reasoning
We noticed a frustrating problem: when multimodal models reason over long chains, they gradually stop looking at the imageโand start hallucinating based on their own text.
So we built RememberโR1, a simple RL framework that directly supervises visual attention on the original reasoning trajectoryโno inference overhead, no proxy tasks.
We use three complementary rewards: coverage, persistence, and focus. They encourage the model to keep attending to relevant visual evidence even in later reasoning steps.
Results across 7 benchmarks and 2 model sizes: better reasoning, andโmore importantlyโvisual attention decays much more slowly during generation.
No extra cost at inference, just cleaner supervision where it counts.
So far I've been pointing it at Markdown Minimap, an Obsidian plugin that adds a scrollable IDE-style minimap to your notes. This week I've been clearing a backlog of user-reported issues on it, with Claude often handling them end to end.
We're exited to announce BananaMind OS, our OS specically for running BananaMind models! Its able to run BananaMind 2 Nano at 4 bit on only 7-8MB of ram, the 2 bit on 6MB of ram and the 8 bit version on 14MB of RAM! It runs on a 486 or newer! Check out this video and image running BananaMind 2 Nano 4 Bit on 9 MB of RAM and a emulated 486 in QEMU at ~1TPS! We asked it: "What is the first letter of the alphabet?" The response is: "The first letter of the alphabet is: - A. " And if you're asking because of the video, yes I am a arch btw. Comment and like this post for a GitHub link and comment for adding other models!