NeurIPS 2026

Reflection Anchors for Propagation-Aware Visual Retention in Long-Chain Multimodal Reasoning

Xuan Gong1 Hanbo Huang1Hao Zheng1 Yiran Zhang1Wenbin Dai1,2 Weishu Zhao1Shiyu Liang1,†

1 Shanghai Jiao Tong University2 Lanzhou University

† Corresponding author

Keep visual evidence in the loop.
Train sparse reflection anchors to carry visual influence
through long reasoning chains.

Explore RAPO

The idea at a glance

Small updates. Lasting visual influence.

RAPO selects uncertain branching points in a reasoning trajectory and trains the policy to preserve visual dependence in the continuation.

Figure 1 Entropy selects where to update. Chain-masked contrastive KL encourages visual influence to persist downstream.
20%

High-entropy positions selected
as training anchors

6

Reasoning and general-domain
evaluation benchmarks

Native

Standard decoding at inference,
without visual re-injection

The challenge

Longer reasoning.
Fading visual evidence.

As a multimodal chain of thought grows, later steps can rely increasingly on the generated text and lose their connection to the image.

RAPO asks: where can a small policy update make visual evidence matter in future reasoning?

Reflection-anchor policy optimization

Select by uncertainty.
Train for propagation.

An information-theoretic analysis motivates two ingredients: local branching room and downstream visual relevance. RAPO turns them into a practical policy optimization method.

  1. 01

    Generate trajectories

    Sample reasoning rollouts with the image and prompt, recording next-token distributions.

  2. 02

    Select reflection anchors

    Keep the top 20% of positions by token entropy, where multiple continuations remain plausible.

  3. 03

    Measure visual dependence

    Compare normal continuations with a chain-masked reference over a finite downstream window.

  4. 04

    Optimize the policy

    Combine group-based task rewards with a visual-dependence objective at the selected anchors.

Why chain masking?

Follow the influence
beyond one token.

The chain mask blocks visual access and evidence from preceding anchors while retaining the rest of the textual context. A finite-window contrastive KL supplies a tractable training signal for downstream visual dependence.

All anchor selection and masking happen during training. The learned policy decodes normally at inference.

Figure 3 Structured masking isolates a local visual-dependence signal.

Main results

Stronger visual reasoning
across model scales.

Compare methods within the same backbone across six multimodal benchmarks.

Average across six benchmarks

Qwen3-VL-2B-Instruct

+5.85percentage points vs. Base
Table 1 · Qwen3-VL-2B-Instruct · mean@8 accuracy (%)
MethodReasoning-intensiveGeneral-domainAvg.
MathVisionMathVerseEMMALogicVistaMMMU-ProRealWorldQA
Base27.3036.7526.8747.1324.4460.4737.16
GRPO28.2944.3528.5048.7130.6164.2840.79
PAPOG28.2945.1827.2546.4331.4564.0540.44
VPPOG30.5945.4328.7545.7631.7363.1440.90
RAPOG33.5546.1331.2549.5532.2065.3643.01

Qwen3-VL baselines are retrained under controlled settings. Scores report mean@8 accuracy, not pass@8. G and D denote GRPO- and DAPO-based variants.

Evaluation and training details

Evaluation uses temperature 1.0, a maximum generation length of 8,192 tokens, and eight responses per input. The reported metric is mean@8 accuracy.

Models train on ViRL. Qwen3-VL models use 100 RL steps; the Qwen2.5-VL and Qwen2-VL groups use 200. RAPO selects the top 20% of token positions by entropy, with a visual-dependence coefficient of 0.01. The KL window is 3 for Qwen3-VL-2B-Instruct and 1 for the other main-result backbones.

See the paper for full settings, compute-matched comparisons, ablations, and additional backbone evaluations.

Mechanistic insights

Visual influence that
outlasts the anchor.

Mechanism analyses connect sparse reflection anchors with stronger visual-dependence signals later in the reasoning trajectory.

01 / Retention

Keep the continuation connected to the image.

Correct trajectories retain stronger late-stage contrastive KL than incorrect ones. RAPO raises the token-wise KL profile beyond GRPO, including later stages of generation.

Figure 5 Contrastive KL measures divergence between vision-conditioned and vision-masked next-token distributions.
02 / Selection

Anchors align with reasoning decisions.

High-entropy anchors concentrate around operation-selecting tokens such as examine, reconsider, and identify, and respond more strongly to visual perturbations.

Figure 6a Fraction of each token type’s occurrences selected as anchors.
03 / Propagation

A local change carries downstream.

Masking visual access at anchors produces larger downstream KL reductions after RAPO training, supporting the persistence of anchor-mediated visual influence.

Figure 7a–b Downstream KL change under visual-token masking at selected anchors.
Training, too.

RAPO shows faster reward growth and higher reward values than GRPO in the reported Qwen3-VL training runs.

Mechanism studies use 500 randomly sampled ViRL validation instances and Qwen3-VL-2B-Instruct. Contrastive KL is a visual-dependence diagnostic; these plots do not directly measure mutual information.

A closer look

From visual evidence
to the right geometric constraints.

A real case from the paper compares the base model with the RAPO-trained policy.

Rectangle ABCD with diagonals AC and BD intersecting at O. Point E lies on BO, and segment AE is perpendicular to BO.

Geometry · MMMath

What is the length of OD?

In rectangle ABCD, diagonals AC and BD intersect at O. AE is the perpendicular bisector of BO, and AE = √3 cm.

Base model · incorrect

OD = √(12/7) cm

The trajectory treats triangle AOB as right-angled, introducing a geometric constraint that does not follow from the figure.

RAPO · correct

OD = 2 cm

The trajectory uses coordinates and the perpendicularity of AE and BO to recover the rectangle’s geometric constraints.

Original trajectories and token-wise contrastive KL visualizations are available in the paper’s case study. This example illustrates one observed outcome.

The paper

Abstract

Long chain-of-thought (CoT) reasoning improves large vision–language models, but visual information often fades during generation, limiting long-horizon multimodal reasoning. Existing methods either re-inject vision at inference or train policies for stronger grounding, but where to intervene relies on perception heuristics rather than principled gain analysis, and how local visual influence propagates remains implicit. We study this problem from an information-theoretic standpoint and derive a lower bound on the downstream visual gain of a one-step intervention, which suggests two factors: local branching room (token entropy) and downstream visual propagation potential (suffix divergence from a vision-marginalized reference). Guided by this analysis, we propose reflection-anchor policy optimization (RAPO), a GRPO-based policy optimization method that selects high-entropy reflection anchors and optimizes a chain-masked finite-window KL surrogate for downstream visual dependence. Experiments on reasoning-intensive and general-domain benchmarks show that RAPO delivers substantial gains over strong baselines across multiple LVLM backbones. Mechanism analyses further indicate that reflection anchors are enriched for visually sensitive decision points and that RAPO increases contrastive visual-dependence signals along generated trajectories.

Build on this work

Citation

@article{gong2026rapo,
  title={Reflection Anchors for Propagation-Aware Visual Retention
         in Long-Chain Multimodal Reasoning},
  author={Gong, Xuan and Huang, Hanbo and Zheng, Hao and Zhang, Yiran
          and Dai, Wenbin and Zhao, Weishu and Liang, Shiyu},
  journal={arXiv preprint arXiv:2605.09614},
  year={2026},
  doi={10.48550/arXiv.2605.09614},
  url={https://arxiv.org/abs/2605.09614}
}

Figure

Full resolution ↗