Learning Self-Correction in Vision-Language Models via Rollout Augmentation Paper • 2602.08503 • Published 3 days ago • 2
Learning Self-Correction in Vision-Language Models via Rollout Augmentation Paper • 2602.08503 • Published 3 days ago • 2
Octopus Collection RL checkpoints of Octopus-8B and baselines of paper: Learning Self-Correction in Vision–Language Models via Rollout Augmentation • 6 items • Updated 3 days ago