CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes Paper • 2608.27455 • Published 14 days ago • 12
One Example Is Enough to Pass Fairness Benchmarks Collection EMNLP 2026. One-shot GRPO LoRA adapters: a single BBQ example saturates fairness benchmarks without making models fairer. • 12 items • Updated 18 days ago