REACT Rolling Denoising and Dual Decoupling for
Reactive Robot Control with VLA Models

Houlong Xiong1,2* Zhenqi Qiu2,3* Zechen Wang2* Suohang Zhang4 Yiyu Ren2 Wanting Xu3 Hongfei Niu2 Chengyang He5 Ge Sun5 Ran Cheng2 Qian Zhu2†

1Shanghai Jiao Tong University 2PrimeBot 3ShanghaiTech University 4Zhejiang University 5National University of Singapore

*Equal contribution  ·  †Corresponding author

CoRL 2026 · Spotlight

The target bowl is moved mid-grasp

π0.5 · standard horizon

Keeps executing a stale chunk after the bowl moves.

π0.5 · short execution horizon

Replans often, but moves in a fragmented stop-and-replan pattern.

REACT (ours)

Adapts to the new target while keeping a smooth approach.

64.8%real-world success
best π0.5 baseline 58.7%, Async+RTC 17.7%
~2×faster reaction than full-chunk π0.5
734 ms vs. 1505 ms (human teleop. 705 ms)
63%Pour Rice success
vs. ≤33% for every baseline
Smootherthan Async+RTC on all 3 platforms
>3× lower jerk on ARX X5

TL;DR REACT turns flow-based VLA action chunking into rolling denoising: future action blocks are refined under several fresh observations before they are executed, so the robot reacts quickly without giving up smooth, long-horizon motion. Dual Decoupling then streams actions at camera rate on a single GPU.

Abstract

Flow-based vision-language-action (VLA) models generate action chunks for temporally coherent robot motion, but chunked control creates a fundamental closed-loop trade-off: long chunks provide smooth execution, whereas frequent replanning improves reactivity at the cost of action discontinuities. We introduce REACT, a rolling-denoising framework that makes flow-based VLAs more reactive while preserving long-horizon context. Instead of regenerating entire action chunks from scratch, REACT maintains a persistent action buffer with staggered flow timesteps. At each control step, the full horizon is denoised using the latest observation, the cleanest action block is executed, partially refined future blocks are shifted forward, and fresh noise is appended to the tail. As a result, each executed action block is refined across multiple recent observations before deployment. To support real-time control, we further introduce dual decoupling, which separates sensing, VLM encoding, DiT denoising, and action execution, enabling high-frequency observation updates and action streaming under practical compute constraints. Across the RoboTwin 2.0 simulation benchmark and real-world tasks spanning bimanual manipulation and dynamic control on multiple robot platforms, REACT improves task success and reduces reaction latency while producing smoother trajectories than frequent-replanning and asynchronous baselines.

The coherence–freshness trade-off

A closed-loop VLA must both incorporate fresh observations often (input throughput) and keep producing executable actions (output throughput). Faster inference and shorter chunks each address only part of this.

Faster inference is not enough

Executing a 50-action chunk dominates the control cycle, so even 2× faster inference only lifts fresh-observation throughput from 0.58 Hz to 0.59 Hz.

SchemeΦin (Hz)Φout (Hz)
Synchronous, He=500.5828.8
Sync + 2× acceleration, He=500.5929.4
Synchronous, He=102.5025.0
Asynchronous, He=103.0030.0
REACT (S=10)3.0030.0

REACT matches asynchronous throughput while keeping the full 50-action rolling buffer instead of independent 10-action chunks. Analytical estimate for π0.5 on one RTX 4090 (TVLM=37 ms, TDiT=3 ms, 10 denoising steps, 30 Hz camera and control).

Shorter chunks break action coherence

Replanning every few actions refreshes observations, but every chunk boundary is a chance to switch to a different action mode. Asynchronous execution hides latency, yet still stitches chunks predicted from observations that no longer match the scene.

  • Long chunks go stale. π0.5 H50E50 keeps pouring after the rice passes the mark: 3% success on Pour Rice.
  • Short chunks retry and switch modes. π0.5 H50E10 keeps returning to the cap instead of twisting it: 0% on Bottle Cap Unscrewing on both robots.
  • Async chunks are stitched. Async+RTC has the highest simulation jerk of all methods (2022.2) and only 17.7% real-world success.

Overview

REACT overview: a rolling action buffer is refined under consecutive observations; throughput and performance comparison.
REACT refines the whole active horizon under each new observation, executes the cleanest block, shifts the partially refined blocks forward, and appends fresh noise. Colored borders track the same action block as it is refined across observations. Right: REACT reaches the high-input, high-output throughput regime. Click to enlarge.

Method

1 · Rolling denoising

REACT keeps a persistent buffer of H = K·S = 50 actions split into K = 5 blocks of S = 10, with staircase flow times τj = (⌊j/S⌋ + 1)/K: the front block is nearly clean, the tail is pure noise.

Each control update applies one DiT Euler step to the whole buffer under the latest observation, executes the now-clean front block, shifts the remaining blocks forward, and appends fresh noise. Every executed block has therefore been refined across multiple recent observations — while the policy still reasons over the full 50-action horizon.

observation ot
← executed nextfar future →
pure noise clean action
Each tick: one Euler step under the newest observation → execute the clean front block → shift → append noise at the tail. Border colors follow one block over time, as in the Overview figure.

2 · Dual Decoupling

A serialized implementation would wait for VLM encoding, DiT denoising and execution before using the next observation. Dual Decoupling is an inference-time scheduler that leaves the weights unchanged: sensing selects one frame every S camera periods, a pool of M VLM workers encodes frames into a latest-ready cache, and each DiT update consumes the freshest embedding together with the rolling buffer.

With S = 10 and M = 2 workers on one RTX 4090, REACT consumes observations at 3 Hz and streams actions at 30 Hz. The embedding is on average 233 ms old (p95 383 ms, max 416 ms) against a 333 ms observation period, so conditioning is never more than one selected observation behind.

Dual Decoupling timing diagram: parallel VLM workers feed a latest-ready cache consumed by rolling DiT denoising.

3 · Staircase training

Standard flow matching samples a single random timestep for the whole chunk. REACT instead trains the action expert on the same per-block noise levels {1/K, 2/K, …, 1} that rolling inference queries, supervising all K denoising stages in one forward pass.

Training therefore matches deployment, and each demonstration covers every stage the scheduler needs; empirically this also speeds up training (see below).

B0
τ=0.2
B1
τ=0.4
B2
τ=0.6
B3
τ=0.8
B4
τ=1.0
Training samples exactly these per-block noise levels — the same ones rolling inference sees (K = 5).

Simulation: RoboTwin 2.0

Seven bimanual tasks, 50 demonstrations per task, 100 evaluation rollouts per task under Clean and Randomized conditions.

Average success rate (%) over the seven tasks. Best in bold, second best underlined.

Settingπ0.5
H50E50
π0.5
H50E10
π0.5
H10E10
Async
+RTC
REACT
w/o DD
REACT
Clean37.4333.7110.8633.2949.7144.86
Randomized11.438.712.4310.5720.8619.86

Real world: ARX X5 and Franka Research 3

Two bimanual platforms, six platform–task pairs spanning size-based reasoning, bimanual coordination and deformable or articulated objects; 30 trials per task with matched randomized placements.

ARX X5
Bowl Stacking by Size
Bottle Cap Unscrewing
Cable Hanging
Franka R3
Bowl Stacking by Size
Bottle Cap Unscrewing
Keyring Hanging

Click any video to pause or resume.

Real-world success rate (%). Best in bold, second best underlined.

RobotTaskπ0.5
H50E50
π0.5
H50E10
π0.5
H10E10
Async
+RTC
REACT
w/o DD
REACT
ARX X5Bowl Stacking by Size736670368676
Bottle Cap Unscrewing360005647
Cable Hanging56665636380
Average55.044.042.013.068.367.7
FrankaBowl Stacking by Size60500476663
Bottle Cap Unscrewing570036370
Keyring Hanging70330176053
Average62.327.80.022.363.062.0
Overall58.735.921.017.765.764.8
Mean jerk norm per method in simulation, ARX X5 and Franka.
Smoother trajectories. Mean jerk over successful episodes (lower is smoother); Demos (GT) applies the same metric to the demonstrations. REACT reduces jerk relative to Async+RTC on every platform (simulation 1219.0 vs. 2022.2, ARX X5 90.1 vs. 284.1, Franka 28.6 vs. 31.6).

Dynamic tasks

Two specialized tasks isolate the two requirements of dynamic control: reacting to continuously changing physical states, and reacting to sudden visual events.

Pour Rice · ARX X5

The robot must stop pouring once the rice reaches the target mark. Long open-loop chunks react late and over-pour; REACT's closed-loop refinement stops at the mark far more often (63% vs. 3% for π0.5 H50E50).

π0.5 baseline

Keeps pouring past the target mark.

REACT (ours)

Tracks the rising level and stops near the mark.

Consecutive wrist-camera frames of the baseline continuing to pour after the rice passes the target mark.
Failure of long open-loop chunks: consecutive frames from π0.5 H50E50, which keeps pouring after the rice level passes the target mark.
Pour Rice success rate (%)
π0.5 H50E50
3
π0.5 H50E10
33
π0.5 H10E10
0
Async+RTC
0
REACT w/o DD
57
REACT
63

Reaction Game · Franka R3

The robot must press a key as soon as the on-screen signal turns green, directly measuring visual-to-motor latency. A single timing script drives both the stimulus and the readout; the human reference teleoperates the same arm through the same timing path.

Screen shows red: wait.
Wait (red)
Screen turns green.
Go (green)
Robot presses the key; the measured latency is shown on screen.
Press — latency of this single trial shown on screen
Reaction latency (ms, mean ± std; lower is better)
π0.5 H50E50
1505 ± 502
π0.5 H50E10
768 ± 113
π0.5 H10E10
772 ± 119
Async+RTC
747 ± 123
REACT w/o DD
754 ± 127
REACT
734 ± 94
Human teleop.
705 ± 73

Full REACT has the lowest mean robot latency, 29 ms above human teleoperation of the same arm and about 2× faster than full-chunk π0.5.

Training efficiency

Staircase training supervises the noise levels used at inference, and REACT also learns faster: on Click Bell, Move Playing Card Away and Press Stapler it reaches 52.7% average success after 3k steps (π0.5: 25.0%) and matches π0.5's 30k-step performance with only 15k steps. On the out-of-distribution demo_randomized data, average open-loop MSE drops from 0.272 to 0.048.

Training loss over optimization steps for three tasks.
Training loss
Average success rate as a function of training steps.
Success vs. training steps
Open-loop MSE on out-of-distribution data; Ours denotes REACT.
OOD open-loop MSE ("Ours" = REACT)

BibTeX

@inproceedings{xiong2026react,
  title     = {{REACT}: Rolling Denoising and Dual Decoupling for Reactive Robot Control with {VLA} Models},
  author    = {Xiong, Houlong and Qiu, Zhenqi and Wang, Zechen and Zhang, Suohang and Ren, Yiyu and Xu, Wanting and Niu, Hongfei and He, Chengyang and Sun, Ge and Cheng, Ran and Zhu, Qian},
  booktitle = {Proceedings of the 10th Conference on Robot Learning},
  series    = {Proceedings of Machine Learning Research},
  year      = {2026},
  eprint    = {2610.12007},
  archivePrefix = {arXiv},
  primaryClass = {cs.RO}
}