LIBERO-MAX · DYNAMIC ROBUSTNESS

Do robot policies adapt
when the world changes?

Yunbei Zhang, Zijian Jin, Yuanzhe Liu, Janet Wang, Xilun Zhang, Yuyou Zhang Zhenyu Zhang, Daoan Zhang, Shuaicheng Niu, Gen Li, Jianfei Yang, Jihun Hamm Ismini Lourentzou, Weirui Ye, Bo Liu, Peter Stone, Marco Pavone

LIBERO-MAX changes the environment after a robot has already begun acting. Every Dynamic rollout is matched to a Base control with the same task, reset, policy seed, and executed prefix. Base continues without an added event; Dynamic receives one change that persists for the rest of the rollout.

LIBERO-MAX introduces one event during execution and compares a no-event Base rollout with a Dynamic rollout sharing the same executed prefix
01

the evaluation gap

Same past. Changed present.

Base and Dynamic share the same executed prefix, isolating the outcome effect of adding one event.

STATIC ROBUSTNESS

The shift is visible before action begins.

shift→observe→act

LIBERO-Plus and LIBERO-PRO test generalization from a shifted reset.

DYNAMIC ROBUSTNESS

The world changes after the policy commits.

reset→act→change→continue

LIBERO-MAX measures whether task success is preserved after the event. It does not infer internal detection or deliberate replanning.

the frozen main track

Eight ways the world can change

1,000 matched cases per event type

Target relocation animation
GEOMETRY · 01

Target relocation

The manipulated object moves after approach.

Receptacle relocation animation
GEOMETRY · 02

Receptacle relocation

The destination moves during the approach.

Camera shift animation
OBSERVATION · 03

Camera shift

The viewpoint changes while the robot is acting.

Sensor corruption animation
OBSERVATION · 04

Sensor corruption

Noise and occlusion begin after execution starts.

Illumination switch animation
APPEARANCE · 05

Illumination switch

Scene lighting changes abruptly.

Visual theme switch animation
APPEARANCE · 06

Visual theme switch

The scene appearance changes online.

Distractor burst animation
CLUTTER · 07

Distractor burst

Five irrelevant objects enter the active support.

Obstacle insertion animation
PATH · 08

Obstacle insertion

A new obstacle blocks the approach region.

02

benchmark construction

From static source cases to changes during execution.

LIBERO-MAX keeps the task substrate fixed, preserves the executed action prefix, and introduces one controlled change after the robot begins acting.

Construction of LIBERO-MAX from LIBERO tasks and selected Plus-derived and PRO-derived source cases

8,000 source cases become 8,000 matched dynamic evaluations. The frozen track combines 5,600 Plus-derived and 2,400 PRO-derived cases, then applies one controlled event during policy execution.

03

main results

Online changes break successful episodes.

Fourteen policies are evaluated on the same 8,000 Base/Dynamic pairs, using each policy's released inference settings.

20.8–56.1%

of Base successes become failures after the event.
Across all assigned cases, every policy loses success, with an 11.0–25.7 percentage-point Base-to-Dynamic drop.

The matched comparison holds the task, initial state, policy seed, and executed prefix fixed within each pair. Every policy's paired 95% bootstrap interval for Dynamic minus Base success lies below zero.

Base and Dynamic success rates for all fourteen policies
One controlled mid-task change reduces success for every policy. Open circles mark Base and filled squares Dynamic. Values summarize all 8,000 pairs per policy.
Paired task outcomes and regression among Base-successful cases
Aggregate losses reflect many previously successful cases becoming failures. The paired outcomes distinguish preserved success, event-associated regression, change-associated success, and persistent failure.

where policies fail

Event type matters.

Geometry and observation changes generally cause larger losses than appearance, clutter, and path changes. Sensitivity also varies across policies, motivating evaluation by event rather than a single model-family ranking.

Paired success-rate changes across eight event types and fourteen policies
Event sensitivity varies across policies. Each policy-event cell contains 1,000 matched pairs. Values are Dynamic minus Base success in percentage points.

viewpoint and execution history

A changed view can be difficult from the start.

On the same 1,000 camera cases per policy (700 Plus-derived and 300 PRO-derived), we compare no added change (Base), the camera change before the first input (Reset), and the same change during execution (Mid-task). The three-policy control contains 9,000 complete rollouts.

Success rate (%) under the three camera conditions
PolicyBaseResetMid-task
X-VLA68.2[65.4, 71.0]1.3[0.6, 2.0]12.4[10.4, 14.5]
π0.579.5[77.1, 81.9]51.3[48.2, 54.4]55.3[52.2, 58.4]
HiMem-WAM72.5[69.7, 75.2]68.8[65.9, 71.6]72.6[69.8, 75.3]

Italic values are 95% source-stratified bootstrap intervals. This camera diagnostic uses π0.5 with H = 10 and Q = 5; its primary benchmark setting uses H = 50.

Paired differences between Base, Reset, and Mid-task camera conditions
Mid-task success exceeds Reset for all three policies. Points show paired differences with 95% bootstrap intervals. Reset versus Mid-task also changes visited states, exposure duration, and remaining budget, so the contrast does not isolate internal adaptation.

Trajectory analysis reveals failures hidden by aggregate rates: 84 model–case instances succeed in both Base and Reset but fail in Mid-task. All reach a new policy query, and 13 execute no stale actions after the event.

X-VLA success under quarter, half, and full camera-shift strength
Milder camera changes improve both timing conditions. X-VLA uses the same 1,000 cases at each strength. At quarter strength, Reset reaches 36.0% and Mid-task 51.7%, while the shared Base success rate is 68.2%.

feedback timing

More frequent queries do not close the gap.

Changing the number of actions executed per policy query changes performance, but the best cadence depends on the policy. Across all 18 valid settings, Dynamic remains 11.1–23.0 percentage points below Base.

Base and Dynamic success across policy query intervals
Dynamic remains below Base at every valid query cadence. Each setting uses the same 800-pair Lite pool. Base/Dynamic prefixes match within a setting. Only Q ≤ H is valid, so π0.5 stops at Q = 8; for X-VLA and GR00T N1.7, H changes with Q.

a targeted response

Restoration recovers part of the sensor-noise loss.

A simple image-restoration module improves X-VLA on 300 fixed sensor-noise cases without retraining. The quality gate uses only current images to trigger inpainting and denoising, with no event flag or clean reference image.

X-VLA success (%) and Dynamic gain over Native
ProcessingBaseDynamicGain (pp)
Native66.740.7—
Always-on68.347.3+6.7[2.3, 11.0]
Quality-gated67.047.7+7.0[3.0, 11.3]

Italic values are paired 95% bootstrap intervals for Dynamic gains. Each strategy uses 300 Base/Dynamic pairs, totaling 1,800 rollouts. All assigned cases remain in the denominator.

Quality-gated restoration narrows the Base/Dynamic gap by 6.7 points (95% CI: [2.3, 11.3]), with a 0.3-point Base change. Always-on restoration also helps; the gated-versus-always Dynamic difference is not statistically resolved. A 19.3-point gap remains.

Additional analyses

Explore source-category results, coverage diagnostics, and the rapid evaluation track. Every figure opens at full size.

Source benchmarks and all 17 source categories

The 5,600 Plus-derived and 2,400 PRO-derived cases expose different source difficulties. These descriptive slices retain the source categories and paired event outcomes.

Summary of paired performance across Plus and PRO source categories
Losses extend across source categories. Source-specific results complement the aggregate comparison.
Policy performance across seven LIBERO-Plus source categories
Seven Plus source categories. Each source category contributes 800 matched cases per policy.
Policy performance across ten LIBERO-PRO source categories
Ten PRO source categories. Each source category contributes 240 matched cases per policy. Low Base success can limit the observable loss in some categories.
Trigger coverage and post-event response diagnostics

Coverage distinguishes whether a rollout reaches the event and receives a subsequent policy query. Cases that miss either remain in the headline score with their observed task outcome.

Event coverage and post-event response diagnostics
Outcome measures and coverage answer complementary questions. These diagnostics locate event exposure and opportunities for feedback without attributing an internal detection or replanning mechanism.
LIBERO-MAX Lite: a fixed subset for rapid evaluation

Lite contains 800 pairs selected independently of policy outcomes, preserving Max’s event and source proportions. It requires 1,600 rollouts per policy.

Agreement between LIBERO-MAX and LIBERO-MAX Lite across fourteen policies
Lite approximates the full benchmark for rapid iteration. Every reported Base rate, Dynamic rate, and paired gap is within 2.4 percentage points of Max, and 88 of 91 pairwise Dynamic orderings are preserved.
04

open benchmark

Validate first.
Evaluate next.

The frozen manifest, source revisions, validation, evaluation launchers, and automated tests are released with the benchmark.

$ git clone https://github.com/liberomax/LIBERO-MAX.git
$ cd LIBERO-MAX
$ python3 -m venv .venv
$ source .venv/bin/activate
$ pip install -e .
$ make validate
05

citation

Cite LIBERO-MAX.

If you use LIBERO-MAX, please cite our arXiv paper. Download BibTeX.

BIBTEX · arXiv:2609.36518
@misc{zhang2026liberomaxrobotpoliciesadapt,
  title = {LIBERO-MAX: Do Robot Policies Adapt When the World Changes?},
  author = {
    Yunbei Zhang and
    Zijian Jin and
    Yuanzhe Liu and
    Janet Wang and
    Xilun Zhang and
    Yuyou Zhang and
    Zhenyu Zhang and
    Daoan Zhang and
    Shuaicheng Niu and
    Gen Li and
    Jianfei Yang and
    Jihun Hamm and
    Ismini Lourentzou and
    Weirui Ye and
    Bo Liu and
    Peter Stone and
    Marco Pavone
  },
  year = {2026},
  eprint = {2609.36518},
  archivePrefix = {arXiv},
  primaryClass = {cs.RO},
  url = {https://arxiv.org/abs/2609.36518}
}