Anchors align with reasoning decisions.
High-entropy anchors concentrate around operation-selecting tokens such as examine, reconsider, and identify, and respond more strongly to visual perturbations.
1 Shanghai Jiao Tong University2 Lanzhou University
† Corresponding author
Keep visual evidence in the loop.
Train sparse reflection anchors to carry visual influence
through long reasoning chains.
The idea at a glance
RAPO selects uncertain branching points in a reasoning trajectory and trains the policy to preserve visual dependence in the continuation.
High-entropy positions selected
as training anchors
Reasoning and general-domain
evaluation benchmarks
Standard decoding at inference,
without visual re-injection
The challenge
As a multimodal chain of thought grows, later steps can rely increasingly on the generated text and lose their connection to the image.
RAPO asks: where can a small policy update make visual evidence matter in future reasoning?
Conceptual illustration · reasoning progresses from left to right
Reflection-anchor policy optimization
An information-theoretic analysis motivates two ingredients: local branching room and downstream visual relevance. RAPO turns them into a practical policy optimization method.
Sample reasoning rollouts with the image and prompt, recording next-token distributions.
Keep the top 20% of positions by token entropy, where multiple continuations remain plausible.
Compare normal continuations with a chain-masked reference over a finite downstream window.
Combine group-based task rewards with a visual-dependence objective at the selected anchors.
Why chain masking?
The chain mask blocks visual access and evidence from preceding anchors while retaining the rest of the textual context. A finite-window contrastive KL supplies a tractable training signal for downstream visual dependence.
All anchor selection and masking happen during training. The learned policy decodes normally at inference.
Main results
Compare methods within the same backbone across six multimodal benchmarks.
| Method | Reasoning-intensive | General-domain | Avg. | ||||
|---|---|---|---|---|---|---|---|
| MathVision | MathVerse | EMMA | LogicVista | MMMU-Pro | RealWorldQA | ||
| Base | 27.30 | 36.75 | 26.87 | 47.13 | 24.44 | 60.47 | 37.16 |
| GRPO | 28.29 | 44.35 | 28.50 | 48.71 | 30.61 | 64.28 | 40.79 |
| PAPOG | 28.29 | 45.18 | 27.25 | 46.43 | 31.45 | 64.05 | 40.44 |
| VPPOG | 30.59 | 45.43 | 28.75 | 45.76 | 31.73 | 63.14 | 40.90 |
| RAPOG | 33.55 | 46.13 | 31.25 | 49.55 | 32.20 | 65.36 | 43.01 |
Qwen3-VL baselines are retrained under controlled settings. Scores report mean@8 accuracy, not pass@8. G and D denote GRPO- and DAPO-based variants.
Evaluation uses temperature 1.0, a maximum generation length of 8,192 tokens, and eight responses per input. The reported metric is mean@8 accuracy.
Models train on ViRL. Qwen3-VL models use 100 RL steps; the Qwen2.5-VL and Qwen2-VL groups use 200. RAPO selects the top 20% of token positions by entropy, with a visual-dependence coefficient of 0.01. The KL window is 3 for Qwen3-VL-2B-Instruct and 1 for the other main-result backbones.
See the paper for full settings, compute-matched comparisons, ablations, and additional backbone evaluations.
Mechanistic insights
Mechanism analyses connect sparse reflection anchors with stronger visual-dependence signals later in the reasoning trajectory.
Correct trajectories retain stronger late-stage contrastive KL than incorrect ones. RAPO raises the token-wise KL profile beyond GRPO, including later stages of generation.
High-entropy anchors concentrate around operation-selecting tokens such as examine, reconsider, and identify, and respond more strongly to visual perturbations.
Masking visual access at anchors produces larger downstream KL reductions after RAPO training, supporting the persistence of anchor-mediated visual influence.
RAPO shows faster reward growth and higher reward values than GRPO in the reported Qwen3-VL training runs.
Mechanism studies use 500 randomly sampled ViRL validation instances and Qwen3-VL-2B-Instruct. Contrastive KL is a visual-dependence diagnostic; these plots do not directly measure mutual information.
A closer look
A real case from the paper compares the base model with the RAPO-trained policy.

Geometry · MMMath
In rectangle ABCD, diagonals AC and BD intersect at O. AE is the perpendicular bisector of BO, and AE = √3 cm.
The trajectory treats triangle AOB as right-angled, introducing a geometric constraint that does not follow from the figure.
The trajectory uses coordinates and the perpendicularity of AE and BO to recover the rectangle’s geometric constraints.
Original trajectories and token-wise contrastive KL visualizations are available in the paper’s case study. This example illustrates one observed outcome.
The paper
Long chain-of-thought (CoT) reasoning improves large vision–language models, but visual information often fades during generation, limiting long-horizon multimodal reasoning. Existing methods either re-inject vision at inference or train policies for stronger grounding, but where to intervene relies on perception heuristics rather than principled gain analysis, and how local visual influence propagates remains implicit. We study this problem from an information-theoretic standpoint and derive a lower bound on the downstream visual gain of a one-step intervention, which suggests two factors: local branching room (token entropy) and downstream visual propagation potential (suffix divergence from a vision-marginalized reference). Guided by this analysis, we propose reflection-anchor policy optimization (RAPO), a GRPO-based policy optimization method that selects high-entropy reflection anchors and optimizes a chain-masked finite-window KL surrogate for downstream visual dependence. Experiments on reasoning-intensive and general-domain benchmarks show that RAPO delivers substantial gains over strong baselines across multiple LVLM backbones. Mechanism analyses further indicate that reflection anchors are enriched for visually sensitive decision points and that RAPO increases contrastive visual-dependence signals along generated trajectories.
Build on this work
@article{gong2026rapo,
title={Reflection Anchors for Propagation-Aware Visual Retention
in Long-Chain Multimodal Reasoning},
author={Gong, Xuan and Huang, Hanbo and Zheng, Hao and Zhang, Yiran
and Dai, Wenbin and Zhao, Weishu and Liang, Shiyu},
journal={arXiv preprint arXiv:2605.09614},
year={2026},
doi={10.48550/arXiv.2605.09614},
url={https://arxiv.org/abs/2605.09614}
}