The dataset will focus on refusals and redirects surrounding dangerous or extreme sexual content, designed to be shaped sized appropriately for SLMs, without significantly lowering benchmark performance.
More info will be out soon!
NVIDIA has openly supported open weights models. I am interested to see how the community evolves.
I think it's due to people asking questions that Bananamind is also good at. As I use the SLM Arena I'm finding there's some things that SmolLM can do that the others cannot do - and if I ask that question in repeated battles, SmolLM goes up in points consistently while others go down.
I'm literally checking the leaderboards on this every day. I'm surprised to see SmolLM has fallen to fourth.
I'd suggest saving up more data... this is the breadth of human knowledge and probably 700 isn't enough to teach classifying correctly. I'd suggest waiting until you have more like 10,000 or 50,000 records.
Oh, classifier would be the much better option - but a lot more work. The good news with the classifier is you can always back-fill the existing records. Or - if I might suggest - use a small LM to do the classification. I would pick qwen3.5-0.8B or 2B. You could also offer to let users pick the category AND record your classifier's classification separately.
If you aren't already, I invite you to join the SLM Consortium. There's a discord channel - the plan is to pursue group endeavors that all can benefit from.
https://huggingface.co/slmconsortium
There's a discord channel too that you should join.
Another suggestion - and this is just a suggestion. Allow the user to select an optional category field. Default to no category. I suggest these:
Then in the leader boards, optionally allow to pick a category, defaulting to none which would render a graph including all battles.
Do you think it's possible to include twice the output from each model? just when it's getting juicy for some, it's truncated.
It's pretty amazing, I hope you add more SLMs to it!
Where can I find them?
I would love to be part of this consortium.
I have an rtx pro 6000 blackwell with 96gb vram. much of the journey is on my blog. https://wayne.theworkmans.us/
I've been working on it since November of 2025. I'm publishing my datasets that I'm creating. Have made a lot of iterations on the model but I feel it's just not ready yet. Making an enormous amount of data corrections. So at this rate, probably a month or more. Maybe longer. So, I don't know when.