From the source
June 18th, 2026 Teaching VLMs to think visually What are good reasoning priors for visual problem solving in VLMs?
June 18th, 2026 Teaching VLMs to think visually What are good reasoning priors for visual problem solving in VLMs?
Abstract The success of reinforcement learning on verifiable rewards (RLVR) is constrained by the reasoning behaviors already present in the base model, yet it remains poorly understood which initial reasoning strategies, which we call "reasoning priors", make an effective starting point, particularly for vision-language models whose priors are inherited from a text-only pre-training phase.
We study this question directly: we teach a 9B VLM to play the card game "Set!"
by supervised fine-tuning on programmatically constructed chains of thought that isolate individual reasoning strategies (textual vs. visually grounded reasoning, with and without backtracking), then apply RL on outcome-level rewards to each variant.
…





