Anonymous submission · Vision–Language–Action

MoQ: Improving Vision-Language-Action Policies with Object Detection Auxiliary Learning for Robotic Manipulation

A sparse, decoder-level grounding signal that helps VLA policies identify the right object, target, and grasp-relevant part when generating actions.

Anonymous Authors

Under review

For the most reliable video playback, please use Google Chrome or Microsoft Edge. Safari may experience playback issues.

Read the story

01 · Motivation

Action-only supervision underconstrains task-relevant visual grounding.

Behavior cloning constrains the demonstrated action trajectory, but it does not explicitly supervise the semantic identity, spatial location, or manipulation-relevant part that supports that trajectory. This leaves room for positional shortcuts and unstable behavior under layout, lighting, and object-distribution shifts.

Paper figure illustrating semantic confusion and localization drift in action-only decoding
MoQ addresses semantic confusion and localization drift by coupling object-level grounding directly to action decoding.
Bad case 01
Bad case 02
Bad case 03
Bad case 04

02 · Method

Decoder-level query coupling for object grounding and action generation.

MoQ inserts a small set of role-specific detection queries into a DiT-style action decoder. Shared self- and cross-attention enables bidirectional information exchange; task-specific feed-forward pathways and prediction heads preserve the distinct output spaces of continuous control and object detection.

MoQ decoder architecture from the paper
Detection and action queries share self- and cross-attention, then branch into task-specific heads for object grounding and action chunks.

03 · Experimental evidence

Experimental evaluation across simulation and physical manipulation.

MoQ achieves strong absolute performance and consistent relative gains across both simulation and physical manipulation, including the normal and disturbed real-world settings.

Green: highest value in each rowGold: second-highest value in each row

+8.89pp

RoboTwin clean average over Base-scratch

+7.78pp

RoboTwin random average over Base-scratch

+19.2pp

Real-world normal-setting gain

+24.0pp

Real-world disturbed-setting gain

RoboTwin2.0 average success rate

Clean / Random
1007550250
PI0.5-FT70.56 / 27.11
XVLA-FT76.00 / 15.89
Base-scratch74.67 / 30.78
MoQ-scratch83.56 / 38.56

clean random / OOD

Real-world comparison

Base → MoQ
EvaluationBaseMoQ
Multi-position first success52.0879.17
Multi-position recovery success72.9291.67
Object selection accuracy70.3788.89
Geometric insertion accuracy68.7590.63
Mug task success91.6797.22

Physical experiments · clean setting

Complete comparison across the four real-world tasks

Success rate (%)
TaskMetricPI0.5-FTXVLA-scratchXVLA-FTBase-scratchMoQ-scratch
Multi-position pick & placeFirst success37.5025.0077.0852.0879.17
Recovery success54.1735.4291.6772.9291.67
Multi-object selection & transferTransfer success92.5988.8996.30100.00100.00
Selection accuracy33.3337.0451.8570.3788.89
Geometric block insertionTask success77.7841.6758.3361.1180.56
Selection accuracy100.0058.3394.4488.8988.89
Insertion accuracy77.7871.4361.7668.7590.63
Mug into holdersTask success94.4491.6788.8991.6797.22

Physical experiments · disturbed setting

Lighting and clutter robustness

Success rate (%)
Task / disturbanceMetricPI0.5-FTXVLA-scratchXVLA-FTBase-scratchMoQ-scratch
Multi-position pick & place
+ random light
First success33.3318.7575.0031.2575.00
Recovery success47.9220.8389.5837.5087.50
Geometric block insertion
+ clutter
Task success50.0019.4436.1155.5675.00
Selection accuracy97.2252.7886.1180.5688.89
Insertion accuracy51.4336.8441.9468.9784.38
Mug into holders
+ random light
Task success77.7888.8986.1188.8991.67

Six answers from the paper

Six practical advantages of MoQ

Q1

Spatial localization

Improves multi-position pick-and-place by +27.08 pp on the first attempt and +18.75 pp after recovery.

Q2

Language-specified selection

Raises correct object selection from 70.37% to 88.89% under randomized layouts.

Q3

Precise insertion

Raises insertion accuracy from 68.75% to 90.63% through better spatial conditioning.

Q4

Part-aware grasping

Improves mug-into-holder success from 91.67% to 97.22% by grounding the handle.

Q5

Illumination robustness

Under unseen colored lighting, improves first-attempt and recovery success by +43.75 pp and +50.00 pp.

Q6

Clutter robustness

In cluttered insertion, reaches 75.00% task success and 84.38% insertion accuracy.

04 · Physical setup & overview

Embodied platforms.

We evaluate MoQ on a dual-arm 5-DoF setup with a head camera and two wrist cameras, and a single-arm 6-DoF setup with a head and wrist camera. The following grid layout and platform views define the physical evaluation protocol.

Dual-arm 5-DoF platform from the paper
Dual-arm 5-DoF · head + two wrist cameras
Single-arm 6-DoF platform from the paper
Single-arm 6-DoF · head + wrist camera

Multi-position evaluation layout

Grid-based placement protocol.

The multi-position pick-and-place task places the object on a 5 cm grid. Each arm is evaluated over a 4 × 3 set of test positions, isolating spatial localization and recovery behavior from object identity changes.

Grid spacing
5 cm
Test positions
4 × 3 per arm
Views
Head + wrist
Third-person overview of the projected placement grid
Projected workspace overview
Detailed projected grid and multi-position manipulation workspace
Grid placement detail
MoQ highlight reelIndependent overview video

Third-view demonstrations

Third-view demonstrations.

These four clips, extracted from the provided presentation, show the physical scene from an independent third-person viewpoint before the camera-recorded experiment videos.

Third-view example 01
Third-view example 02
Third-view example 03
Third-view example 04

05 · Camera-aligned real-world demonstrations

Real-world experiment recordings.

Each clean task is recorded from the head and wrist cameras used by the policy. Tasks 1, 3, and 4 are shown as a single two-view row; task 2 retains all three available camera views. Below each recording, the original PPT detection frames are kept as individual images: head frames form one row and wrist frames form separate row(s). Hover, focus, or tap any frame to inspect that specific frame.

01

Multi-position pick and place

Spatial generalization · 2 views

Head cameraOverview
Right wrist cameraGrasp & placement

Head camera · 4 detection frames

Right wrist camera · 4 detection frames

Detection frames · task 1 cleanOriginal PPT frames · hover or focus to enlarge
03

Geometric block insertion

Shape-aware manipulation · 2 views

Head cameraTask overview
Left wrist cameraInsertion view

Head camera · 4 detection frames

Left wrist camera · 4 detection frames

Detection frames · task 3 cleanOriginal PPT frames · hover or focus to enlarge
04

Mug into holders

Target-aware placement · 2 views

Head cameraTask overview
Left wrist cameraApproach & placement

Head camera · 4 detection frames

Left wrist camera · 4 detection frames

Detection frames · task 4 cleanOriginal PPT frames · hover or focus to enlarge

Real-world disturbed settings

Lighting and clutter demo videos.

These demo videos correspond to the disturbed evaluations reported above: random colored lighting for tasks 1 and 4, and unseen tabletop clutter for task 3.

Task 1 · random lightPick & place
HeadDisturbed
WristDisturbed

Head camera

Right wrist camera

Detection frames · random lightOriginal PPT frames
Task 3 · clutterBlock insertion
HeadDisturbed
WristDisturbed

Head camera

Left wrist camera

Detection frames · clutterOriginal PPT frames
Task 4 · random lightMug placement
HeadDisturbed
WristDisturbed

Head camera

Left wrist camera

Detection frames · random lightOriginal PPT frames

07 · RoboTwin2.0 experiment recordings

Clean and random evaluation videos.

For each selected RoboTwin2.0 task, the clean and random clips are displayed together. The random setting is held out during training and probes robustness to cluttered textures and visual distractors.

Clean settingin-distribution recording
Beat block hammer
Detection frames · 5 original imagesHover or focus to enlarge
Place phone stand
Detection frames · 5 original imagesHover or focus to enlarge
Place object basket
Detection frames · 9 original imagesHover or focus to enlarge
Place container plate
Detection frames · 6 original imagesHover or focus to enlarge
Place object stand
Detection frames · 6 original imagesHover or focus to enlarge
Place object scale
Head detection frames · 5 imagesHover or focus to enlarge
Random settingheld-out visual perturbation
Beat block hammer · random
Detection frames · 5 original imagesHover or focus to enlarge
Place phone stand · random
Detection frames · 5 original imagesHover or focus to enlarge
Place object basket · random
Detection frames · 9 original imagesHover or focus to enlarge
Place container plate · random
Detection frames · 6 original imagesHover or focus to enlarge
Place object stand · random
Detection frames · 6 original imagesHover or focus to enlarge
Place object scale · random
Head detection frames · 5 imagesHover or focus to enlarge

MoQ

Object-centric auxiliary supervision
for robust VLA action decoding.