CoRP
Compact Robot Policies Need Fine-Grained Visual Representations

Compressed Representation Policy · 48.9M parameters · no vision-language model · no video-generative prior

Anonymous Authors

Under review at ICLR 2027
Contribution.
Simulation.

LIBERO

Four suites in one grid. The arrows step through the tasks in that suite.

LIBERO-Spatial

LIBERO-Object

LIBERO-Goal

LIBERO-Long

RoboTwin 2.0

One successful rollout per task. The arrows move to the next task. Left, head, and right play together in the same row.

Clean

Randomized

Same tasks, with clutter and changes to background, lighting, and table height.

Peg Insertion

An 8.0 mm peg into an 8.1 mm hole. In simulation, CoRP reaches 90% success. These are success clips from training variants of that task. Use the arrows to move between them.

Real-World.

The policy, after pretraining on more than 10,000 hours and 5,887 tasks. Each clip is left, head, and right. The policy runs at about 20 ms per chunk.

Pick cup

Place spoon

Plug insertion

Abstract.
CoRP architecture and a comparison of parameter count against larger policies

CoRP. (a) The representation extractor, built from DINOv2, T5, and an MLP, outputs the conditioning sequence zt. A one-block flow-matching action generator turns that sequence into an action chunk. (b) CoRP stays competitive with much larger systems at 0.049B parameters.

Multi-task manipulation policies differ in architecture, scale, and pretrained priors all at once, so published comparisons cannot attribute performance to any single component. We argue that most of it comes from the visual representation, and that parameter scale and generative priors are largely incidental.

To test this, we build CoRP (Compressed Representation Policy), a deliberately compact policy (48.9M parameters, no vision-language model and no video-generative prior) that factorizes into a representation extractor and a flow-matching action generator. It reaches 97.0% on LIBERO and 75.78% / 73.36% on RoboTwin 2.0 Clean / Randomized, matching systems 40.9–163.6× larger.

Holding the action generator fixed, we then vary one extractor property at a time. Pretrained initialization is decisive: a random ViT-S/14 drops to 78.1% and an ImageNet ResNet-34 to 74.5% on LIBERO. Pretraining alone is not enough, as freezing the encoder costs 19.8 points. Compression matters as much: resampling each view to 48 tokens beats passing all patch tokens (97.0% vs 83.2%), and a variational information bottleneck over those tokens is worse than a hard token budget, cutting LIBERO-Goal from 95.8% to 33.0% by suppressing the instruction-dependent token selection the policy relies on. Language conditioning contributes only where the observation leaves the goal ambiguous (LIBERO-Goal: 9.2% to 95.8%), while on RoboTwin 2.0, where observations are unambiguous, removing it slightly improves success.

A compact policy works when its representation is pretrained, task-adapted, and compressed.

Experiments.

Can a compact policy match much larger systems?

Yes. On LIBERO the default CoRP (0.049B) reaches 97.0%, 5.9 points above ProgVLA, the strongest compact policy in the comparison that also uses neither a VLM nor a video prior. A wider or deeper variant goes further: 93.9% at d = 180 (0.037B) and 98.4% at L = 3 (0.063B). All ablations below use the one-block default.

MethodParams (B)SpatialObjectGoalLongAvg.
With a vision-language model
SmolVLA2.2593.094.091.077.088.8
π03.396.898.895.885.294.1
π0.53.398.898.298.092.496.9
OpenVLA7.084.788.479.253.776.5
OpenVLA-OFT7.097.698.497.994.597.1
With a video generation model
Cosmos-Policy2.098.1100.098.297.698.5
Fast-WAM-Joint6.099.699.498.296.898.5
Motus8.096.899.896.697.697.7
Without either prior
ACT0.08482.078.866.144.067.7
Diffusion Policy0.15778.392.568.350.572.4
Octo0.09378.985.784.651.175.1
ProgVLA0.10987.696.092.088.691.1
CoRP (L=1, d=180)0.03797.699.493.285.493.9
CoRP (L=1, d=768, default)0.04998.299.495.894.497.0
CoRP (L=3, d=768)0.06398.299.698.897.098.4

Success rate (%) on LIBERO. 50 trials per task.

On RoboTwin 2.0, CoRP (75.8% / 73.4%) outperforms X-VLA and π0.5 and approaches Motus trained without its video pretraining (77.6% / 77.0). It stays below video-pretrained Motus and GigaWorld-Policy. Diffusion Policy and ACT, which learn a visual encoder without a large-scale pretrained vision model, collapse under randomization (0.6% and 1.7%). CoRP loses 2.4 points.

MethodParams (B)CleanRandomized
With a vision-language model
X-VLA0.972.972.8
π03.346.416.3
π0.53.343.043.8
With a video generation model
Motus8.088.787.0
Motus, without pretraining8.077.677.0
GigaWorld-Policy11.486.485.0
Without either prior
ACT0.08429.71.7
Diffusion Policy0.15728.00.6
CoRP0.04975.873.4

Success rate (%) on RoboTwin 2.0. 100 trials per task, averaged over 50 tasks.

Does a pretrained visual encoder help?

Yes. A DINOv2-initialized ViT-S/14 beats every other encoder variant by at least 14.9 points on average LIBERO success. Neither factor is enough on its own: with random initialization the ViT is no better than the ResNet (78.1% vs 82.1%), and ImageNet pretraining does not help the ResNet (74.5% vs 82.1%). All four variants identify the task at 100% accuracy. The pretrained ResNet still only reaches 37.0% on LIBERO-Goal: the policy knows the task and lacks the visual detail to carry it out.

EncoderSpatialObjectGoalLongAvg.NMAER²
Random ResNet-3491.297.872.866.682.10.120.54
ImageNet ResNet-3485.498.837.076.874.50.090.73
Random ViT-S/1484.495.686.046.478.10.110.59
DINOv2 ViT-S/1498.299.495.894.497.00.070.81

LIBERO success (%). NMAE and R² are from a frozen probe that predicts the action chunk from the visual prefix. Lower NMAE and higher R² mean a more action-informative representation.

Is pretraining enough without adaptation?

t-SNE of visual-prefix tokens with a frozen encoder and with fine-tuning

t-SNE of the visual prefix, frozen versus fine-tuned. A frozen DINOv2 encoder leaves a less useful representation for control.

No. Freezing the pretrained encoder drops average LIBERO success from 97.0% to 77.2%, no better than a randomly initialized ViT trained end to end (78.1%). The probe agrees: freezing cuts R² from 0.81 to 0.24 and roughly doubles NMAE, from 0.07 to 0.16.

Visual encoderAverageNMAER²
Frozen DINOv277.20.160.24
Fine-tuned97.00.070.81

When does language conditioning help?

t-SNE of visual-prefix tokens with and without language on LIBERO

t-SNE of the visual prefix with and without language. Without the instruction, Goal tasks collapse together. Spatial tasks stay mostly separable.

Language is a disambiguator, not a general source of task knowledge. On LIBERO-Goal, where one scene is shared across instructions, removing language drops success from 95.8% to 9.2%, near the 10% chance level of a 10-task suite. The other suites lose at most 11.2 points. On RoboTwin 2.0, where the scene identifies the task, removing language raises success from 75.8% to 78.5% (Clean) and from 73.4% to 76.1% (Randomized).

LIBERO Spatial Object Goal Long Average
Acc.S.R. Acc.S.R. Acc.S.R. Acc.S.R. Acc.S.R.
Without language90.487.0100.0100.08.69.289.083.472.069.9
With language100.098.2100.099.4100.095.8100.094.4100.097.0

Acc. is first-frame task-head accuracy. S.R. is success rate.

RoboTwin 2.0 Clean Randomized
Acc.S.R.Acc.S.R.
Without language99.9678.5099.9276.06
With language99.7875.7899.5873.36

Does the information bottleneck help?

A hard token budget does. Passing every patch token scores 83.2%. Resampling each view to 48 instruction-conditioned tokens scores 97.0%. The gain is concentrated on LIBERO-Goal (55.8% to 95.8%), the suite where the instruction alone specifies the goal. Swapping the instruction while holding the observation fixed changes the resampled visual prefix by 49.7%, and the all-token prefix by only 5.4%.

A variational bottleneck does the opposite. It compresses every token toward the same prior, regardless of what the current instruction needs. On the resampled prefix it cuts the instruction-induced change from 49.7% to 13.6% and Goal success from 95.8% to 33.0%. All-tokens plus VIB still names the Goal task at 99.9% accuracy and succeeds only 33.6% of the time.

Spatial Object Goal Long Average
Method Acc.S.R. Acc.S.R. Acc.S.R. Acc.S.R. Acc.S.R.
All tokens93.889.6100.099.691.455.8100.087.696.383.2
All tokens + VIB99.887.8100.099.499.933.6100.090.299.977.8
48 tokens100.098.2100.099.4100.095.8100.094.4100.097.0
48 tokens + VIB91.687.899.999.269.633.095.583.689.275.9
Conclusion.

CoRP is a 48.9M-parameter policy with an ordinary flow-matching action generator and no VLM or video prior. On these closed-set benchmarks it matches systems up to 163× larger, and the same compact design transfers to a real bimanual robot. Across the ablations, weaker variants still recognize the task and fail to perform it. What they lack is fine-grained, instruction-specific visual detail.

The study measures in-distribution control. It does not test the open-world semantic and physical generalization that VLAs and world-action models are built for. Whether a pretrained, task-adapted, compressed representation is enough in that regime is open.

BibTeX
@inproceedings{corp2027,
  title     = {Compact Robot Policies Need Fine-Grained Visual Representations},
  author    = {Anonymous},
  booktitle = {Under review at the International Conference on Learning Representations},
  year      = {2027},
  url       = {https://corp-policy.github.io/}
}