Copper-Policy
arXiv · 2026.09

Copper-Policy

Focus on the Representation for Robust Robot Manipulation

Zexin Feng1 Yixu Feng2 Lingyu Xiao1 Shang Su3,4 Kexin Zheng1
Chang Xu2 Mengkai Shi4 Shuo Feng3 Xintao Yan1✉

1The University of Hong Kong 2The University of Sydney 3Tsinghua University 4DenseAI Inc.

Cheap Training, Strong Control.

0:00 / 0:00
0:00 / 0:00
0:00 / 0:00
0:00 / 0:00
Core idea: a compact World representation evolves under task intention and stays decodable into actions.
LIBERO
97.25%
Average over 4 suites
LIBERO-Plus
80.85%
10,030 perturbed tasks
RoboTwin
70.84%Clean 12.98%Random
Real robot
96.3%
3 real-robot tasks
Parameters
2B
No embodied pretraining
Training
9.67h
30K steps, 8× RTX 5090
Abstract

Learn the space, predict in it, act with detail.

Conclusion: Predicting a compact, policy-learned World representation improves control without generating pixels or futures at deployment. Current-frame features preserve the spatial detail needed for action.

Evidence: 80.85% on LIBERO-Plus, 41.91% average on RoboTwin, and 96.3% across three real-robot tasks. The 2B model trains in 9.67 hours on 8× RTX 5090. On 8× A100, it trains about 6× faster than Fast-WAM and 2× faster than π0.5.

Method
Full Copper-Policy architecture and its training and inference attention masks.

Select a variant to inspect its inputs and prediction path. Deployment: action decoding only; no future generation.

01

Learn a policy-shaped space

Joint policy training shapes the future targets.

02

Predict in that space

Future representations; no pixel reconstruction.

03

Act with visual detail

Decode actions from World representation and detailed contexts.

Results

LIBERO & LIBERO-Plus — full comparison

Method Pretrain WAM LIBERO LIBERO-Plus
LongGoalObjectSpatialAvg. CameraRobotLanguageLightBackgroundNoiseLayoutTotal
Copper-Policy (ours) ×✓‡ 95.295.699.498.897.2578.6773.7467.3495.2784.2094.4476.5280.85
π0.5 ✓× 92.498.098.298.896.975.477.585.696.994.689.785.785.7
Cosmos-Policy ×✓ 97.698.2100.098.198.575.863.381.796.588.992.782.282.2
ABot-M0 ✓× 96.699.099.898.898.660.467.986.496.291.686.482.680.5
VLA-JEPA ✓✓‡ 95.897.299.696.297.264.267.788.191.893.465.883.977.9
X-VLA ✓× 97.697.898.698.298.123.489.775.788.296.062.771.871.4
Fast-WAM ×✓‡ 95.297.0100.098.297.616.444.568.978.253.737.760.751.5

Success rate (%); Pretrain = embodied pretraining, WAM = future modeling, ‡ = no future generation at deployment, and bold/underline = best/second-best.

RoboTwin: clean versus randomized, per ablation

ConfigurationClean ↑Random ↑Avg. ↑
Copper-Policy 70.8412.9841.91
- w/o Joint-Embedding Prediction 48.169.3828.77
- w/o Current-Frame Visual Features 67.8210.2239.02
- w/o World Input & Future Expert 55.2610.2232.74
- w/o JE Pred. & Visual Features 51.669.4230.54

Training efficiency

Takeaway: 244 tokens per sample versus 392 for dense video targets. 9.67 hours for 30K steps on 8× RTX 5090; about 6× faster than Fast-WAM.

Method Trainable params Hardware Parallelism Peak mem / GPU ↓ Relative time ↓ Samples / s ↑ Est. time, 30K steps ↓ Inference ↓
Fast-WAM6B8× A100 80GBFSDP Fully Shard 62.42 GB6.13×14.3674.3 h~80 ms
π0.53.3B8× A100 80GBFSDP Fully Shard 38.65 GB1.96×44.9123.7 h~75 ms
Copper-Policy on 8× A100 80GB — sharding strategies
Copper-Policy2B8× A100 80GBDDP 31.83 GB1.00×88.012.1 h~85 ms
Copper-Policy2B8× A100 80GBDDP + Optimizer Shard 24.53 GB1.08×81.813.0 h~85 ms
Copper-Policy2B8× A100 80GBFSDP Fully Shard 21.24 GB1.50×58.5618.2 h~85 ms
Copper-Policy on 8× RTX 5090 32GB
Copper-Policy (ours)2B8× RTX 5090 32GBDDP 31.83 GB0.80×110.349.67 h~85 ms

Global batch 128. Relative time uses Copper-Policy DDP on 8× A100 (1.454 s/step) as 1.00×. Peak memory is reserved memory per GPU; estimated time assumes 30K steps; inference is end-to-end.

Real robot

96.3% across three real-robot tasks.

50 seeds per task · dual arm · three RGB views · 300 demonstrations per task.

0:00 / 0:00
Fold the towel · 93 — deformable object; 5 partial successes among 50 seeds.
0:00 / 0:00
Put banana · 96 — pick and place across seeded object placements.
0:00 / 0:00
Stack three bowls · 100 — 50 of 50 seeds succeed.

Real-robot results

ConfigurationFold the towelPut bananaStack three bowlsAvg. ↑
π0.5 95949494.3
Fast-WAM 98609484.0
Copper-Policy 939610096.3
- w/o World Input & Future Expert 84828884.7
- w/o Joint-Embedding Prediction 42610043.3
- w/o Current-Frame Visual Features 2461420.7
- w/o JE Pred. & Visual Features 0401.3
Why it works

Predictive training organizes the World representation.

Future prediction keeps the World state coherent over time

At horizon 16, removing joint-embedding prediction lowers cosine similarity in LIBERO (0.794→0.684), RoboTwin (0.913→0.729), and the real robot (0.902→0.756).

Configuration LIBERORoboTwinReal robot
cos ↑disp. ↓cos ↑disp. ↓cos ↑disp. ↓
Copper-Policy 0.7940.634 0.9130.413 0.9020.434
- w/o Joint-Embedding Prediction 0.6840.778 0.7290.724 0.7560.684
- w/o Current-Frame Visual Features 0.7780.660 0.8900.463 0.9040.429
- w/o JE Pred. & Visual Features 0.6650.809 0.5670.921 0.7730.658

40 LIBERO, 20 RoboTwin, and 18 real-robot episodes per variant; one fixed seed. RoboTwin samples every fourth frame, so cross-domain values are not directly comparable.

Episode-level temporal cosine similarity across five horizons for Copper-Policy and three ablations on LIBERO, RoboTwin, and the real robot.
Temporal coherence: joint-embedding prediction sustains higher similarity across horizons.
Perturbation separation and complementary features

Change separates. Perturbation–temporal subspace overlap falls from 0.798 to 0.187; perturbation projection into the temporal subspace falls from 0.807 to 0.140.

The state resists perturbations. Paired perturbations shift zt by 0.134 versus 0.344 without prediction; cosine similarity rises from 0.932 to 0.990.

Two channels complement each other. The action probe reaches R2 = 0.927 with zt and vt together, versus 0.802 and 0.910 alone.

Perturbation-temporal subspace overlap across ranks 1 to 64, with action-probe R-squared for four representations below.
Representation Geometry at rank 8Paired perturbationAction probe
overlap ↓pert. in temp. ↓norm. L2 ↓cosine ↑R2 ↑
Copper zt 0.1870.1400.1340.9900.802
- w/o Joint-Embedding Prediction 0.7980.8070.3440.9320.787
Current-frame tokens vt ——0.1710.9690.910
Dense V-JEPA features ——0.4280.9040.831
Token-matched V-JEPA features ——0.4620.8910.803
zt and vt together ————0.927

Attention: task object versus background

Same episode, two checkpoints. Joint-embedding prediction keeps attention on the manipulated object.

Copper-Policy
w/o Joint-Embedding Prediction
0:00 / 0:00

LIBERO — “put both moka pots on the stove”, episode 129.

Copper-Policy
w/o Joint-Embedding Prediction
0:00 / 0:00

RoboTwin — “move the red, green and blue blocks to the centre and stack them”, episode 27. Sampled every fourth frame.

Copper-Policy
w/o Joint-Embedding Prediction
0:00 / 0:00

Real robot — “fold the towel”, episode 66.

Citation

BibTeX

@misc{feng2026copperpolicyfocusrepresentationrobust,
      title={Copper-Policy: Focus on the Representation for Robust Robot Manipulation}, 
      author={Zexin Feng and Yixu Feng and Lingyu Xiao and Shang Su and Kexin Zheng and Chang Xu and Mengkai Shi and Shuo Feng and Xintao Yan},
      year={2026},
      eprint={2609.32779},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2609.32779}, 
}