Learn a policy-shaped space
Joint policy training shapes the future targets.
Copper-Policy
Focus on the Representation for Robust Robot Manipulation
1The University of Hong Kong 2The University of Sydney 3Tsinghua University 4DenseAI Inc.
Cheap Training, Strong Control.
Conclusion: Predicting a compact, policy-learned World representation improves control without generating pixels or futures at deployment. Current-frame features preserve the spatial detail needed for action.
Evidence: 80.85% on LIBERO-Plus, 41.91% average on RoboTwin, and 96.3% across three real-robot tasks. The 2B model trains in 9.67 hours on 8× RTX 5090. On 8× A100, it trains about 6× faster than Fast-WAM and 2× faster than π0.5.
Select a variant to inspect its inputs and prediction path. Deployment: action decoding only; no future generation.
Joint policy training shapes the future targets.
Future representations; no pixel reconstruction.
Decode actions from World representation and detailed contexts.
| Method | Pretrain | WAM | LIBERO | LIBERO-Plus | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Long | Goal | Object | Spatial | Avg. | Camera | Robot | Language | Light | Background | Noise | Layout | Total | |||
| Copper-Policy (ours) | × | ✓‡ | 95.2 | 95.6 | 99.4 | 98.8 | 97.25 | 78.67 | 73.74 | 67.34 | 95.27 | 84.20 | 94.44 | 76.52 | 80.85 |
| π0.5 | ✓ | × | 92.4 | 98.0 | 98.2 | 98.8 | 96.9 | 75.4 | 77.5 | 85.6 | 96.9 | 94.6 | 89.7 | 85.7 | 85.7 |
| Cosmos-Policy | × | ✓ | 97.6 | 98.2 | 100.0 | 98.1 | 98.5 | 75.8 | 63.3 | 81.7 | 96.5 | 88.9 | 92.7 | 82.2 | 82.2 |
| ABot-M0 | ✓ | × | 96.6 | 99.0 | 99.8 | 98.8 | 98.6 | 60.4 | 67.9 | 86.4 | 96.2 | 91.6 | 86.4 | 82.6 | 80.5 |
| VLA-JEPA | ✓ | ✓‡ | 95.8 | 97.2 | 99.6 | 96.2 | 97.2 | 64.2 | 67.7 | 88.1 | 91.8 | 93.4 | 65.8 | 83.9 | 77.9 |
| X-VLA | ✓ | × | 97.6 | 97.8 | 98.6 | 98.2 | 98.1 | 23.4 | 89.7 | 75.7 | 88.2 | 96.0 | 62.7 | 71.8 | 71.4 |
| Fast-WAM | × | ✓‡ | 95.2 | 97.0 | 100.0 | 98.2 | 97.6 | 16.4 | 44.5 | 68.9 | 78.2 | 53.7 | 37.7 | 60.7 | 51.5 |
Success rate (%); Pretrain = embodied pretraining, WAM = future modeling, ‡ = no future generation at deployment, and bold/underline = best/second-best.
| Configuration | Clean ↑ | Random ↑ | Avg. ↑ |
|---|---|---|---|
| Copper-Policy | 70.84 | 12.98 | 41.91 |
| - w/o Joint-Embedding Prediction | 48.16 | 9.38 | 28.77 |
| - w/o Current-Frame Visual Features | 67.82 | 10.22 | 39.02 |
| - w/o World Input & Future Expert | 55.26 | 10.22 | 32.74 |
| - w/o JE Pred. & Visual Features | 51.66 | 9.42 | 30.54 |
Takeaway: 244 tokens per sample versus 392 for dense video targets. 9.67 hours for 30K steps on 8× RTX 5090; about 6× faster than Fast-WAM.
| Method | Trainable params | Hardware | Parallelism | Peak mem / GPU ↓ | Relative time ↓ | Samples / s ↑ | Est. time, 30K steps ↓ | Inference ↓ |
|---|---|---|---|---|---|---|---|---|
| Fast-WAM | 6B | 8× A100 80GB | FSDP Fully Shard | 62.42 GB | 6.13× | 14.36 | 74.3 h | ~80 ms |
| π0.5 | 3.3B | 8× A100 80GB | FSDP Fully Shard | 38.65 GB | 1.96× | 44.91 | 23.7 h | ~75 ms |
| Copper-Policy on 8× A100 80GB — sharding strategies | ||||||||
| Copper-Policy | 2B | 8× A100 80GB | DDP | 31.83 GB | 1.00× | 88.0 | 12.1 h | ~85 ms |
| Copper-Policy | 2B | 8× A100 80GB | DDP + Optimizer Shard | 24.53 GB | 1.08× | 81.8 | 13.0 h | ~85 ms |
| Copper-Policy | 2B | 8× A100 80GB | FSDP Fully Shard | 21.24 GB | 1.50× | 58.56 | 18.2 h | ~85 ms |
| Copper-Policy on 8× RTX 5090 32GB | ||||||||
| Copper-Policy (ours) | 2B | 8× RTX 5090 32GB | DDP | 31.83 GB | 0.80× | 110.34 | 9.67 h | ~85 ms |
Global batch 128. Relative time uses Copper-Policy DDP on 8× A100 (1.454 s/step) as 1.00×. Peak memory is reserved memory per GPU; estimated time assumes 30K steps; inference is end-to-end.
50 seeds per task · dual arm · three RGB views · 300 demonstrations per task.
| Configuration | Fold the towel | Put banana | Stack three bowls | Avg. ↑ |
|---|---|---|---|---|
| π0.5 | 95 | 94 | 94 | 94.3 |
| Fast-WAM | 98 | 60 | 94 | 84.0 |
| Copper-Policy | 93 | 96 | 100 | 96.3 |
| - w/o World Input & Future Expert | 84 | 82 | 88 | 84.7 |
| - w/o Joint-Embedding Prediction | 4 | 26 | 100 | 43.3 |
| - w/o Current-Frame Visual Features | 2 | 46 | 14 | 20.7 |
| - w/o JE Pred. & Visual Features | 0 | 4 | 0 | 1.3 |
At horizon 16, removing joint-embedding prediction lowers cosine similarity in LIBERO (0.794→0.684), RoboTwin (0.913→0.729), and the real robot (0.902→0.756).
| Configuration | LIBERO | RoboTwin | Real robot | |||
|---|---|---|---|---|---|---|
| cos ↑ | disp. ↓ | cos ↑ | disp. ↓ | cos ↑ | disp. ↓ | |
| Copper-Policy | 0.794 | 0.634 | 0.913 | 0.413 | 0.902 | 0.434 |
| - w/o Joint-Embedding Prediction | 0.684 | 0.778 | 0.729 | 0.724 | 0.756 | 0.684 |
| - w/o Current-Frame Visual Features | 0.778 | 0.660 | 0.890 | 0.463 | 0.904 | 0.429 |
| - w/o JE Pred. & Visual Features | 0.665 | 0.809 | 0.567 | 0.921 | 0.773 | 0.658 |
40 LIBERO, 20 RoboTwin, and 18 real-robot episodes per variant; one fixed seed. RoboTwin samples every fourth frame, so cross-domain values are not directly comparable.
Change separates. Perturbation–temporal subspace overlap falls from 0.798 to 0.187; perturbation projection into the temporal subspace falls from 0.807 to 0.140.
The state resists perturbations. Paired perturbations shift zt by 0.134 versus 0.344 without prediction; cosine similarity rises from 0.932 to 0.990.
Two channels complement each other. The action probe reaches R2 = 0.927 with zt and vt together, versus 0.802 and 0.910 alone.
| Representation | Geometry at rank 8 | Paired perturbation | Action probe | ||
|---|---|---|---|---|---|
| overlap ↓ | pert. in temp. ↓ | norm. L2 ↓ | cosine ↑ | R2 ↑ | |
| Copper zt | 0.187 | 0.140 | 0.134 | 0.990 | 0.802 |
| - w/o Joint-Embedding Prediction | 0.798 | 0.807 | 0.344 | 0.932 | 0.787 |
| Current-frame tokens vt | — | — | 0.171 | 0.969 | 0.910 |
| Dense V-JEPA features | — | — | 0.428 | 0.904 | 0.831 |
| Token-matched V-JEPA features | — | — | 0.462 | 0.891 | 0.803 |
| zt and vt together | — | — | — | — | 0.927 |
Same episode, two checkpoints. Joint-embedding prediction keeps attention on the manipulated object.
LIBERO — “put both moka pots on the stove”, episode 129.
RoboTwin — “move the red, green and blue blocks to the centre and stack them”, episode 27. Sampled every fourth frame.
Real robot — “fold the towel”, episode 66.
@misc{feng2026copperpolicyfocusrepresentationrobust,
title={Copper-Policy: Focus on the Representation for Robust Robot Manipulation},
author={Zexin Feng and Yixu Feng and Lingyu Xiao and Shang Su and Kexin Zheng and Chang Xu and Mengkai Shi and Shuo Feng and Xintao Yan},
year={2026},
eprint={2609.32779},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2609.32779},
}