Zero-Shot VLA Adaptation
with Retrieved Guidance

RAGu: Retrieval-Augmented Guidance
Anonymous Authors
Under review at ICRA 2027
RAGu enables zero-shot adaptation across domains and embodiments.

A single retrieval-augmented guidance mechanism transfers manipulation knowledge across four settings: Real→Real (DROID demos to novel real-world Franka scenes), Real→Sim (DROID to RoboLab, and to a 29-DoF GR-1 bimanual humanoid in RoboCasa), and Sim→Real (Meta-World Sawyer demos to the real-world DROID setup) — all with the base VLA's weights completely frozen.

Supplementary Video

Abstract

Pre-training large parametric “Vision-Language-Action” (VLA) models on large demonstration datasets has recently yielded promising results for general-purpose manipulation. However, these models still struggle under even small domain shifts relative to the training data, such as changes in camera viewpoint, background, object identity, or robot morphology.

Motivated by the success of Retrieval-Augmented Generation (RAG) in language models, we pursue retrieval to boost VLA performance. Specifically, we seek to leverage retrieval from domain-mismatched, “in-the-wild” demonstration datasets like DROID to guide a VLA’s behavior in a new target domain. To achieve this, we introduce Retrieval-Augmented Guidance (RAGu), which overcomes robotics-specific challenges for RAG: (1) It incorporates simple geometric calibration to transform source demonstrations into the target setup and reduce cross-domain differences. (2) It uses a coarse-to-fine retrieval strategy that selects, from a large pool of pre-warped candidates, the trajectory most compatible with the base VLA. (3) It guides the VLA’s flow-matching inference procedure rather than seeking to retrain the base model or expand its limited context window.

We evaluate RAGu across real-world, simulation, sim-to-real transfer, and cross-embodiment humanoid manipulation. We find it improves real-world VLA success rates from 20% to 80% and outperforms the closest-related baselines in sim.

Method

Applying retrieval to a frozen VLA must overcome three robotics-specific challenges: (1) Cross-domain generalization — unlike the universal token space of LLMs, robotic data suffers from severe distribution shifts in camera viewpoint, layout, and embodiment, so even the best-matched trajectories in large-scale datasets are still domain-shifted; prior robotic retrieval methods therefore restrict their pools to narrow, identical-embodiment datasets or rely on task-specific fine-tuning. (2) Policy-aware retrieval — geometric proximity does not ensure that a policy can actually follow a trajectory, and close-but-incompatible trajectories can cause unstable motion, so enlarging the retrieval pool requires policy-aware selection. (3) Context-efficient adaptation — VLAs have limited context windows that do not permit directly concatenating retrieved demonstrations to the current observation. RAGu is a zero-shot, training-free adaptation framework that handles all three. It first filters semantically relevant demonstrations and warps them into the target scene via semantic correspondences, turning cross-domain demonstrations into usable spatial priors. It then retrieves the guiding trajectory via coarse offline DTW pruning followed by policy-compatible selection at inference time, avoiding control-loop overhead. Finally, it steers the VLA’s flow-matching denoising toward the chosen trajectory through an energy-based guidance field, leaving the policy’s weights and context untouched.

Overview of the RAGu pipeline: semantic task filtering, warp, retrieve, steer.

Overview of the RAGu pipeline. (a) Semantic task filtering. We filter real-world datasets using semantic keywords for broad task categories with similar kinematic profiles, such as pick-and-place, via language instruction labels in the datasets. (b) Warp. Task-relevant objects in the target scene are identified using semantic bounding box detection to extract target waypoints W′ (W1→W1′, W2→W2′), allowing the source trajectories to be smoothly adapted despite visual shifts. (c) Retrieve. Retrieval candidates i are coarsely pruned to the top-n trajectories, N, using Dynamic Time Warping (DTW)-based matching. We then select the optimal trajectory i* from the pruned trajectories i′ ∈ [N] that minimizes the guidance gradient norm (‖∇A0 Ci*(Xt)‖2), retrieving the demonstration best aligned with the VLA’s expected Cartesian trajectory Xt. (d) Steer. At test time, the chosen reference i* defines an energy field (Lwarp ≈ ‖Xt − i*‖2), steering denoising steps for zero-shot adaptation.

1
Semantic task-category filtering

Large unstructured datasets (e.g., DROID) or simulation benchmarks (e.g., Meta-World) are first filtered into a ~100-trajectory pool using semantic keywords for broad task dynamics such as pick-and-place, wiping, or closing — ensuring structural compatibility with the target task before any expensive computation.

2
Warp — cross-domain trajectories into the target scene

Because cross-domain scenes rarely share object instances, strict in-domain visual matching is replaced with semantic bounding boxes for task-relevant entities (via Grounding DINO), localizing target waypoints in the new scene. Each source trajectory’s 3D waypoints are then warped into the target geometry. Crucially, warped paths are used only as spatial priors — the VLA’s innate generalization manages the closed-loop manipulation, avoiding the unreliability of directly executing warped paths with inverse kinematics.

3
Retrieve — coarse-to-fine, policy-aligned

Coarse (static): for each task-dynamics group, a reference trajectory that is geometrically relevant and compatible with the VLA prior anchors a Dynamic Time Warping (DTW) pruning step, keeping only the top-K warped candidates (e.g., K=8 for real-robot, K=16 for RoboLab, K=3 for the humanoid). Fine (inference-time): in the control loop, the VLA’s predicted action chunk is mapped to a Cartesian end-effector path via forward kinematics, and RAGu selects the candidate whose tracking-cost gradient has the smallest norm — the trajectory requiring the least change to the VLA’s current dynamics, hence closest to the policy’s own intent. Only this fine stage runs in the loop, once per inference at the first denoising step.

4
Steer — test-time guidance of the action denoising

The selected warped trajectory instantiates an energy function whose gradient is subtracted from the flow-matching velocity field (vguided = v − γ∇Lwarp), steering action generation toward the retrieved trajectory while the VLA’s weights stay frozen and its motor priors stay intact.

Warp → Retrieve → Steer, in action

The clips below show the RAGu pipeline on the real robot. RAGu draws on many cross-domain source demonstrations from DROID (left; the rainbow curve traces each demo’s 3D end-effector trajectory, yellow marks the gripper) — different tables, objects, cameras, and backgrounds. All of their trajectories are warped into the target scene via semantic correspondences (right; three warped candidates, each shown from two camera views, all sharing the same bowl and pot keypoints). At run time, RAGu retrieves the warped candidate best aligned with the VLA’s current intent (middle; the full candidate pool overlaid, retrieved one highlighted) and uses it to steer the frozen π0.5 policy in the real world (bottom).

Step 1 Source demos (cross-domain, DROID)
Step 2 Warped into the target scene
selected
selected
selected
discarded
discarded
⋯
~100 demos
warped
→
Camera view 1
Camera view 2
warped candidate 1
warped candidate 1
warped candidate 2
warped candidate 2
warped candidate 3
warped candidate 3

Six of the ~100 pick-and-place source demonstrations used by RAGu. The reference trajectory anchors DTW pruning: selected candidates are retained in the top-K pool and warped into the target scene (right), while discarded ones are pruned. The warped candidates on the right are drawn from the full ~100-demo pool, not only the six tiles shown.

↓
Step 3 Retrieve (best policy-aligned candidate from the warped pool)
Camera view 1
Camera view 2

The warped candidates share the same bowl and pot keypoints but differ in approach. At the first denoising step of each inference, RAGu retrieves the candidate whose guidance gradient is smallest — the one closest to the policy’s own intent — tracked by the gripper.

↓
Step 4 Steer — frozen π0.5 guided rollout

The retrieved trajectory defines the guidance energy that steers the denoising process. The strip shows the actual robot execution from multiple cameras.

One Mechanism, Four Transfer Settings

Because warped 3D trajectories abstract away appearance, viewpoint, and embodiment, the same retrieval-augmented guidance transfers across radically different source–target pairs.

Real → Sim: DROID demonstrations guide simulated robots

DROID source demos warped into nine RoboLab tasks.

RoboLab environment overview. 100 real-world source trajectories are queried from DROID (left) and warped into nine simulated RoboLab tasks. The transparent streams visualize the spatial diversity of the warped candidate trajectories; RAGu’s in-the-loop retrieval picks whichever candidate best matches the policy’s intent at each moment.

Successful RoboLab rollouts (π0.5 + RAGu)

Sauce Bottle in Crate
Take Measuring Spoon Out

Successful rollouts of the frozen π0.5 policy guided by retrieved warped DROID trajectories, shown from all three camera views. Longer episodes are sped up for brevity.

Sim → Real: Meta-World demonstrations guide a real Franka

Cross-embodiment warping from Meta-World Sawyer to real-world DROID.

Cross-embodiment sim-to-real warping. 100 simulated expert trajectories from a Meta-World Sawyer robot (bin-picking-v3) are warped into the physical DROID setup for the multi-stage Banana-to-Plate task across three origin–destination layouts, exploiting free simulation data for robust real-world adaptation.

MetaWorld Source demo trajectory (Sawyer, simulation)
Camera view 1
Camera view 2
↓
Warped Warped into the real scene (Shelf → Table layout, camera view 2)
warped MetaWorld demo
↓
Real world Guided rollout — Put Banana on Plate (success)

Simulated Sawyer demonstrations from MetaWorld are warped into the real DROID scene (middle: the rainbow curve is the warped end-effector trajectory in the same Shelf → Table layout as the first rollout, 00/01 mark the grasp and release keypoints) and steer the frozen π0.5 policy to successful real-world executions across different scene layouts. Longer episodes are sped up for brevity.

Cross-morphology: a 7-DoF arm guides a 29-DoF humanoid

Franka DROID demonstrations guiding the Fourier GR-1 humanoid.

Franka → humanoid guidance. End-effector trajectories from the Franka-based DROID dataset guide GR00T-N1.5 policies on the Fourier GR-1 humanoid tabletop tasks — despite entirely different kinematics, grippers, and camera setups.

Successful GR-1 humanoid rollouts (GR00T-N1.5 + RAGu, zero-shot)

Cuttingboard to Pot — ego-centric view (left), third-person view (right)
Placemat to Plate — ego-centric view (left), third-person view (right)
Plate to Bowl — ego-centric view (left), third-person view (right)

Successful zero-shot rollouts on the 29-DoF Fourier GR-1 humanoid in RoboCasa: the frozen GR00T-N1.5 policy is guided by K=3 warped Franka demonstrations from DROID. Shown at 3× speed.

Results

80%
avg. real-world success (5 tasks)
vs. 40% Steering+VLM, 20% π0.5
25%
RoboLab avg. with 100 DROID demos
vs. 1% RiCL with the same demos
+11.7
points over π0.5 on 9 RoboLab tasks
by scaling retrieval to 100 trajectories
+13.6
points on the GR-1 humanoid
(GR00T-N1.5: 33.6 → 47.2, zero-shot)

Real-world evaluation

Real-world success rates across five tasks, with error bars.

Real-world success rates (10 rollouts per task). Bars show the mean success rate over 10 rollouts per task and method; the Average column pools all 50 rollouts per method. Error bars are ±1 standard error of the binomial proportion (no whisker is drawn at 0% and 100%, where the standard error is zero). RAGu doubles the inference-time steering baseline (Steering + VLM) and quadruples the base π0.5 policy on average, reaching 100% on Wipe Table and 90% on Open Drawer and Close Microwave. The retrieval-augmented fine-tuning baseline RiCL fails to complete any task in this setting.

Simulation evaluation (RoboLab)

Task RiCL π0.5 GR00T N1.7 Ours Δ
BananaInBowl0[0.0, 13.3]4[0.7, 19.5]0[0.0, 13.3]36[20.2, 55.5]+32
MouseOnKeyboard0[0.0, 13.3]12[4.2, 30.0]0[0.0, 13.3]24[11.5, 43.4]+12
MustardAboveRaisin4[0.7, 19.5]0[0.0, 13.3]4[0.7, 19.5]16[6.4, 34.7]+16
Average1.3[0.2, 7.2]5.3[2.1, 12.9]1.3[0.2, 7.2]25.3[16.9, 36.2]+20

RoboLab evaluation (25 rollouts per task). Success rates (%) with 95% Wilson confidence intervals in brackets, comparing the base VLA policies (π0.5, GR00T N1.7) and the retrieval-augmented fine-tuning baseline RiCL with RAGu. Both RiCL and RAGu use the same pool of 100 demonstrations warped from DROID; Δ is the absolute improvement over π0.5, and the Average row pools all 75 rollouts per method. RAGu (25.3% avg.) substantially outperforms RiCL (1.3%), base π0.5 (5.3%), and GR00T (1.3%) — evidence that trajectory-aligned, prior-aware steering is a more context-efficient adaptation mechanism than concatenating demonstrations into the context or updating the VLA’s weights.

Scaling the number of retrieved trajectories (RoboLab)

Task GROOT π0.5 1 traj 10 traj 30 traj 50 traj 100 traj
SRΔ SRΔ SRΔ SRΔ SRΔ
BananaInBowl0[0.0, 13.3]4[0.7, 19.5]28[14.3, 47.6]+2428[14.3, 47.6]+2428[14.3, 47.6]+2424[11.5, 43.4]+2036[20.2, 55.5]+32
MouseOnKeyboard0[0.0, 13.3]12[4.2, 30.0]16[6.4, 34.7]+48[2.2, 25.0]−416[6.4, 34.7]+420[8.9, 39.1]+824[11.5, 43.4]+12
MustardAboveRaisin4[0.7, 19.5]0[0.0, 13.3]8[2.2, 25.0]+80[0.0, 13.3]+00[0.0, 13.3]+00[0.0, 13.3]+016[6.4, 34.7]+16
WoodSpatulaToBowl4[0.7, 19.5]0[0.0, 13.3]4[0.7, 19.5]+40[0.0, 13.3]+08[2.2, 25.0]+84[0.7, 19.5]+412[4.2, 30.0]+12
RubiksCubeOrBanana0[0.0, 13.3]44[26.7, 62.9]52[33.5, 70.0]+847[29.2, 65.6]+352[33.5, 70.0]+856[37.1, 73.3]+1268[48.4, 82.8]+24
SauceBottlesCrate0[0.0, 13.3]4[0.7, 19.5]12[4.2, 30.0]+816[6.4, 34.7]+129[2.7, 26.2]+524[11.5, 43.4]+2012[4.2, 30.0]+8
TakeMeasuringSpoonOut0[0.0, 13.3]36[20.2, 55.5]32[17.2, 51.6]−444[26.7, 62.9]+832[17.2, 51.6]−436[20.2, 55.5]+036[20.2, 55.5]+0
RubiksCube0[0.0, 13.3]52[33.5, 70.0]20[8.9, 39.1]−3268[48.4, 82.8]+1664[44.5, 79.8]+1264[44.5, 79.8]+1265[45.5, 80.5]+13
ToyInBin28[14.3, 47.6]60[40.7, 76.6]84[65.3, 93.6]+2475[55.5, 87.8]+1564[44.5, 79.8]+464[44.5, 79.8]+460[40.7, 76.6]+0
Mean3.6[1.8, 6.9]22.8[17.8, 28.7]26.4[21.1, 32.5]+3.629.5[23.9, 35.8]+6.728.1[22.6, 34.3]+5.329.6[24.0, 35.9]+6.834.5[28.6, 40.9]+11.7

Evaluation on RoboLab tasks with different numbers of trajectories. Task-wise success rates (%, 25 rollouts per task) with 95% Wilson confidence intervals in brackets, across GROOT, π0.5, and π0.5 + RAGu guided by 1, 10, 30, 50, or 100 warped DROID demonstrations. Each Δ column shows the absolute improvement over π0.5; the best setting per task is in bold; the Mean row pools all 225 rollouts of each column. Scaling the number of trajectories significantly enhances policy robustness, with the mean success rate increasing from 22.8% (π0.5) to a peak of 34.5% when utilizing 100 demonstrations. The largest gains appear on tasks requiring precise spatial alignment (BananaInBowl +32, RubiksCubeOrBanana +24). Performance is not strictly monotonic at intermediate pool sizes, but the 100-demo setting is the most robust overall.

Trajectory quality (RoboLab)

Task π0.5 π0.5 + Ours (1 traj) π0.5 + Ours (100 traj)
SPARCPath (m)Speed (cm/s) SPARCPath (m)Speed (cm/s) SPARCPath (m)Speed (cm/s)
BananaInBowl−9.404.147.4−8.623.637.3−10.346.067.3
MouseOnKeyboard−10.516.179.6−10.715.338.8−11.355.468.4
MustardAboveRaisin−8.554.7211.1−10.474.868.5−9.215.499.4
WoodSpatulaToBowl−9.878.2112.6−9.857.8811.8−9.737.5111.5
RubiksCubeOrBanana−7.012.217.9−6.972.097.9−6.401.798.0
SauceBottlesCrate−8.723.948.4−8.213.468.4−9.264.078.7
TakeMeasuringSpoonOut−9.865.578.8−10.816.988.0−10.116.008.3
RubiksCube−10.628.468.4−8.683.198.2−11.168.348.7
ToyInBin−9.769.9510.1−9.8110.879.1−6.312.579.7
Average−9.375.939.37−9.355.378.67−9.325.258.89

Trajectory metrics across RoboLab tasks for π0.5, Ours with 1 trajectory, and with 100 trajectories. EE SPARC is unitless; Path is measured in meters; Speed is measured in cm/s. Higher EE SPARC indicates smoother motion. Averaged across the nine tasks, RAGu with 100 trajectories produces the smoothest motions (EE SPARC −9.32), the shortest average end-effector path (5.25 m vs. 5.93 m for π0.5), and comparable speed (8.89 cm/s), indicating that larger retrieval sets improve efficiency through smoother, more direct motion rather than more aggressive execution.

Secondary Retrieval ablation: candidate selection rule (RoboLab)

Task RAGu (1 traj) RAGu (min dist) RAGu (min grad, default)
EE SPARC ↑Path (m) ↓Speed (cm/s) EE SPARC ↑Path (m) ↓Speed (cm/s) EE SPARC ↑Path (m) ↓Speed (cm/s)
BananaInBowl−8.623.637.30−10.013.696.93−10.346.067.30
MouseOnKeyboard−10.715.338.80−12.004.437.46−11.355.468.40
MustardAboveRaisin−10.474.868.50−13.209.256.90−9.215.499.40
Average−9.934.618.20−11.745.797.10−10.305.678.37

Ablation of Secondary Retrieval: candidate selection rule (Table IV in the paper). Trajectory metrics (25 rollouts per task) on RoboLab for three tasks of the RoboLab evaluation table above. We use the same pool of 100 warped DROID demonstrations, while only the rule that picks the guiding trajectory changes. 1 traj uses only a single trajectory, as in π0.5 + RAGu (1 traj) of the trajectory-quality table above; min dist selects arg mini Ci(Xt) (Eq. 3), i.e., standard nearest-neighbor retrieval in trajectory space; min grad selects arg mini ‖∇A0Ci(Xt)‖2 (Eq. 4, π0.5 + RAGu (100 traj) of the trajectory-quality table above).

Cross-morphology evaluation (GR-1 humanoid, RoboCasa)

Task GR00T-N1.5 (w/o FT) GR00T-N1.5 (w/ FT)
BaseOurs (1 traj)ΔOurs (100 traj)Δ BaseOurs (1 traj)ΔOurs (100 traj)Δ
CuttingboardToPot4[0.7, 19.5]12[4.2, 30.0]+812[4.2, 30.0]+84[0.7, 19.5]8[2.2, 25.0]+48[2.2, 25.0]+4
PlacematToPlate48[30.0, 66.5]60[40.7, 76.6]+1272[52.4, 85.7]+2448[30.0, 66.5]72[52.4, 85.7]+2472[52.4, 85.7]+24
PlateToBowl44[26.7, 62.9]44[26.7, 62.9]+056[37.1, 73.3]+1268[48.4, 82.8]72[52.4, 85.7]+480[60.9, 91.1]+12
PlateToPlate56[37.1, 73.3]44[26.7, 62.9]−1260[40.7, 76.6]+420[8.9, 39.1]40[23.4, 59.3]+2036[20.2, 55.5]+16
TrayToPlate16[6.4, 34.7]36[20.2, 55.5]+2036[20.2, 55.5]+2044[26.7, 62.9]40[23.4, 59.3]−448[30.0, 66.5]+4
Average33.6[25.9, 42.3]39.2[31.1, 48.0]+5.647.2[38.7, 55.9]+13.636.8[28.9, 45.5]46.4[37.9, 55.1]+9.648.8[40.2, 57.5]+12

Cross-embodiment humanoid manipulation with in-context trajectory guidance. Success rates (%, 25 rollouts per task) with 95% Wilson confidence intervals in brackets, on five RoboCasa tabletop tasks for the 29-DoF Fourier GR-1 humanoid, using the base GR00T-N1.5 model and NVIDIA’s task-specific fine-tuned variant, each guided by 1 or 100 DROID (Franka) demonstrations warped into the humanoid’s workspace. The Δ columns indicate absolute improvement over the corresponding unguided baseline; the Average row pools all 125 rollouts of each column. Retrieval guidance improves both checkpoints, from 33.6 to 47.2 average success without fine-tuning and from 36.8 to 48.8 with fine-tuning, and in both configurations the 100-trajectory interval lies entirely above the base-policy point estimate. Base-model robustness raises the ceiling while remaining complementary to retrieval-based guidance.

Inference wall-clock cost (GR-1 humanoid, RoboCasa)

Task GR00T-N1.5 (w/o FT) GR00T-N1.5 (w/ FT)
BaseOurs (1 traj)ΔOurs (100 traj)Δ BaseOurs (1 traj)ΔOurs (100 traj)Δ
CuttingboardToPot37.230.8−6.333.3−3.828.131.0+2.832.9+4.8
PlacematToPlate37.230.8−6.433.4−3.826.730.2+3.532.7+6.0
PlateToBowl36.332.0−4.434.3−2.027.331.4+4.133.5+6.2
PlateToPlate31.830.4−1.433.1+1.325.129.0+3.931.8+6.7
TrayToPlate28.230.9+2.733.7+5.527.030.0+3.132.5+5.6
Average34.131.0−3.233.6−0.626.830.3+3.532.7+5.8

Inference wall-clock cost of cross-embodiment trajectory guidance. We report mean wall-clock time per episode (seconds) on the five-task RoboCasa subset, for the base GR00T-N1.5 model and its task-specific fine-tuned variant. Every cell uses the same protocol — 25 episodes per task under a fixed 720-step horizon (10 Hz) timed on identical hardware — so values are directly comparable. The Δ columns give the change in seconds relative to the corresponding unguided Base (green: faster; red: slower).

Key findings

  • Scaling retrieval helps. On nine RoboLab tasks, mean success grows from 22.8% (π0.5) to 34.5% with 100 warped demos — large gains on tasks needing precise object-relative positioning (BananaInBowl +32, RubiksCubeOrBanana +24). Broad trajectory coverage acts as a dense structural prior rather than a template to copy.
  • Sim-to-real transfer survives multi-stage tasks. In the Banana-to-Plate task, RAGu guided by warped Meta-World trajectories reaches 60–70% final-stage completion across all three layouts. The gain is largest in Cabinet→Shelf, where both π0.5 and Steering+VLM fail completely (0%) while RAGu reaches 60%; in Shelf→Table and Table→Shelf the steering baseline already performs well, and the smaller gains there are directional rather than statistically conclusive.
  • Cross-morphology guidance works. Warped DROID (Franka) demos improve GR00T-N1.5 on the 29-DoF Fourier GR-1 humanoid from 33.6 to 47.2 average success without fine-tuning, and from 36.8 to 48.8 with fine-tuning.
  • Guided trajectories are also more efficient. With 100 demos, RAGu produces the smoothest motions (best EE-SPARC) and the shortest average end-effector paths (5.25 m vs. 5.93 m for π0.5) at comparable speed — the gains come from more direct motion, not from moving faster.
  • Gradient-norm selection beats nearest-neighbor selection. With the demonstration pool, the top-N candidates and γ held fixed on three RoboLab tasks, picking the guiding trajectory by nearest neighbor (min dist) gives less smooth (EE SPARC −11.74 vs. −10.30) and slower (7.10 vs. 8.37 cm/s) motion than RAGu’s gradient-norm rule (min grad) at nearly the same path length. A single curated reference is smoother still but drops success from 25.3% [16.9, 36.2] to 17.3% [10.4, 27.4] on the same tasks (Table IV in the paper).

How the confidence intervals are computed. Every success rate is a binomial proportion p̂ = k/n, with k successful rollouts out of n = 25 per simulated task (10 per real-world task). The tables above report each rate together with its 95% Wilson score interval in brackets, (p̂ + z2/2n) / (1 + z2/n) ± z √(p̂(1 − p̂)/n + z2/4n2) / (1 + z2/n) with z = 1.96. Mean and Average rows pool every rollout of that column (e.g., 9 tasks × 25 = 225 rollouts), so their intervals are narrower than the per-task ones. We use the Wilson interval rather than the normal (Wald) approximation p̂ ± 1.96 √(p̂(1 − p̂)/n) because it stays inside [0, 100]% and remains informative at 0% and 100% with small n (a 0% cell with n = 25 has the interval [0, 13.3]%). The error bars in the real-world bar plot show ±1 standard error, √(p̂(1 − p̂)/n), roughly half the width of a 95% interval. Intervals are given for success rates only, not for the Δ columns or the trajectory-quality metrics.

Systematic generalization on the real robot (STAR-Gen)

Evaluation protocol. We evaluate real-robot generalization using STAR-Gen, which constructs controlled task variations from a base task along four axes: visual-task changes, object-pose changes, action-verb changes, and new-object changes. Each task includes the original setup and four generated variants. For each task–variant pair, we conduct one trial per method using the same initial scene configuration and language instruction. Rather than reporting only binary success, we assign a task-progress score from 0 to 100, following sequential and visually verifiable task milestones. This protocol lets us distinguish failures to approach, interact with, and complete the task.

Per-trial task progress of RAGu and baselines across STAR-Gen generalization axes on the real Franka.

Per-trial task progress across STAR-Gen perturbations. Columns are the STAR-Gen generalization axes (original task, visual-task changes, a new object, shifted object poses, and rephrased action verbs); rows are the pick-and-place and articulated-object base tasks on the real Franka. Above each initial scene, a bar runs from white (0) to black (100) task progress and a triangle marks the score each method reached on that trial: Ours (red), π0.5 (blue), Omniguide (teal), and RiCL (gray). RAGu stays near the high end on every axis, while the baselines mostly stall at the approach or grasp milestones.

Task-progress criteria

Score Pick-and-place tasks
approach → grasp → transport → place
Articulated-object tasks (drawers, cabinets)
approach → handle interaction → opening completion
0 No task-directed motion toward the object. No task-directed motion.
25 Reaches or touches the object but does not securely grasp it. Reaches the drawer or cabinet and contacts its front, door, or an incorrect part.
50 Successful grasp, evidenced by the object moving with the gripper. Contacts or grasps the correct handle without a meaningful opening motion.
75 Transports the object toward the target but fails to release it correctly or leaves it short of the target. Partially opens the drawer or door but releases it, stalls, uses an incorrect interaction, or does not open it fully.
100 Object is stably placed on the target plate and released. The specified drawer or cabinet door is fully opened through the instructed interaction.

Scoring. The final score reflects the furthest milestone reached in a single trial, providing credit for meaningful partial progress while reserving full credit for complete and correct task execution.

Real-Robot Rollouts

Real-robot rollouts across five manipulation tasks. All 200 episodes (4 methods, 5 tasks, 10 episodes each) with action histories are uploaded to the referenced dataset — except one success episode of the Omniguide baseline on the Towel on Plate task, whose video was accidentally not saved. For each task we show representative success and failure episodes for our method (Ours) and two baselines (Omniguide, RiCL).

Each clip shows the first camera viewpoint (left) and the wrist camera (right), the only two views used by the policy. Select a task and a method below to view its episodes.

Success
Failure
Success
Failure
Success
N/A
Failure
Success
Failure
Success
Failure
Success
N/A
Failure
Success
Failure
Success
Failure
Success
N/A
Failure
Success
Failure
Success
Failure
Success
N/A
Failure
Success
Failure
N/A
Success
Failure
Success
N/A
Failure

Conclusion

RAGu is a zero-shot, training-free framework that adapts pretrained VLAs to out-of-distribution environments by semantically warping cross-domain trajectories into the target geometric space and using them as energy fields that guide the flow-matching denoising process. It bridges embodiment and sim-to-real gaps — enabling a Franka manipulator to execute complex multi-stage tasks from simulated Sawyer demonstrations — and scaling the retrieval pool to 100 trajectories significantly improves robustness. Limitations: RAGu relies on visual semantic correspondences and is therefore sensitive to camera placement and occlusion; severe occlusions, including self-occlusion by the robot, can degrade 3D warping and inject noise into the guidance. Future work will incorporate more robust multi-view 3D representations. Overall, RAGu offers a scalable and computationally efficient path to deploy foundation robot models in the wild without costly fine-tuning.