Poke &Wiggle

Reality Check Leaderboard

PAW-GEN-10

Data Scaling

Model scaling by training episodes per environment: SuccessMolmoAct 2: D10 4% (95% CI 3%–6%), D100 37% (95% CI 33%–40%), D300 44% (95% CI 40%–48%). Pi 0.5: D10 5% (95% CI 3%–6%), D100 26% (95% CI 22%–29%), D300 33% (95% CI 29%–37%). DiT-Flow: D10 2% (95% CI 1%–4%), D100 24% (95% CI 21%–28%), D300 29% (95% CI 26%–33%). GR00T N1.7: D10 4% (95% CI 2%–5%), D100 14% (95% CI 11%–17%), D300 19% (95% CI 16%–22%).Success rate (%)03060MolmoAct 2: 4% (95% CI 3%–6%) on D10MolmoAct 2: 37% (95% CI 33%–40%) on D100MolmoAct 2: 44% (95% CI 40%–48%) on D300Pi 0.5: 5% (95% CI 3%–6%) on D10Pi 0.5: 26% (95% CI 22%–29%) on D100Pi 0.5: 33% (95% CI 29%–37%) on D300DiT-Flow: 2% (95% CI 1%–4%) on D10DiT-Flow: 24% (95% CI 21%–28%) on D100DiT-Flow: 29% (95% CI 26%–33%) on D300GR00T N1.7: 4% (95% CI 2%–5%) on D10GR00T N1.7: 14% (95% CI 11%–17%) on D100GR00T N1.7: 19% (95% CI 16%–22%) on D300D10D100D300

D10, D100, and D300 indicate how many demonstration episodes per environment were used to train the model. Whiskers are 95% confidence intervals over the evaluation episodes.

Spatial Generalization

Spatial generalization by model: SuccessMolmoAct 2: Nominal 31% (95% CI 28%–34%), Interpolation 25% (95% CI 23%–28%), Extrapolation coming soon. Pi 0.5: Nominal 23% (95% CI 20%–26%), Interpolation 19% (95% CI 17%–22%), Extrapolation coming soon. DiT-Flow: Nominal 21% (95% CI 19%–24%), Interpolation 16% (95% CI 13%–18%), Extrapolation coming soon. GR00T N1.7: Nominal 13% (95% CI 11%–15%), Interpolation 12% (95% CI 10%–14%), Extrapolation coming soon.Success rate (%)03060Coming soonMolmoAct 2: 31% (95% CI 28%–34%) on NominalMolmoAct 2: 25% (95% CI 23%–28%) on InterpolationPi 0.5: 23% (95% CI 20%–26%) on NominalPi 0.5: 19% (95% CI 17%–22%) on InterpolationDiT-Flow: 21% (95% CI 19%–24%) on NominalDiT-Flow: 16% (95% CI 13%–18%) on InterpolationGR00T N1.7: 13% (95% CI 11%–15%) on NominalGR00T N1.7: 12% (95% CI 10%–14%) on InterpolationNominalInterpolationExtrapolation

Nominal uses training placements; interpolation uses unseen placements between them; extrapolation uses unseen placements outside their convex hull. Whiskers are 95% confidence intervals over the evaluation episodes.

Per Environment (Success)

Model performance by environmentMolmoAct 2 overall: Sort screws into bins 10%, Datum-corner alignment 44%, Pour screws into the funnel 52%, Connect DC jack 15%, Route cable through closed hoops 21%, Put tools in standing toolbox 4%, Place loaded boxes in container (bimanual) 42%, Place spray bottles upright, nozzle forward (bimanual handover) 52%, Open the toolbox with a screwdriver 18%, Clamp electrical component 24%. Pi 0.5 overall: Sort screws into bins 4%, Datum-corner alignment 39%, Pour screws into the funnel 66%, Connect DC jack 2%, Route cable through closed hoops 9%, Put tools in standing toolbox 3%, Place loaded boxes in container (bimanual) 13%, Place spray bottles upright, nozzle forward (bimanual handover) 59%, Open the toolbox with a screwdriver 13%, Clamp electrical component 2%. DiT-Flow overall: Sort screws into bins 2%, Datum-corner alignment 17%, Pour screws into the funnel 48%, Connect DC jack 2%, Route cable through closed hoops 0%, Put tools in standing toolbox 1%, Place loaded boxes in container (bimanual) 32%, Place spray bottles upright, nozzle forward (bimanual handover) 63%, Open the toolbox with a screwdriver 16%, Clamp electrical component 4%. GR00T N1.7 overall: Sort screws into bins 3%, Datum-corner alignment 14%, Pour screws into the funnel 25%, Connect DC jack 1%, Route cable through closed hoops 1%, Put tools in standing toolbox 1%, Place loaded boxes in container (bimanual) 14%, Place spray bottles upright, nozzle forward (bimanual handover) 50%, Open the toolbox with a screwdriver 13%, Clamp electrical component 0%.Sort screws into binse01Datum-corner alignmente02Pour screws into the funnele03Connect DC jacke04Route cable through closed hoopse05Put tools in standing toolboxe06Place loaded boxes in container (bimanual)e07Place spray bottles upright, nozzle forward (bimanual handover)e08Open the toolbox with a screwdrivere09Clamp electrical componente10Basic skillsHigh precisionBimanualTool use

e01–e10 represent the environments in PAW-GEN-10, each value in the selected metric averaged across every data tier and placement condition.

Share of episodes rated Ok or Excellent.

Rank and modelOverall score and number of evaluationsMid-training DataM0 has no mid-training data and post-trains the model directly on the environment demonstrations. M100 first mid-trains it on 100 hours of embodiment data (coming soon).Data ScalingScores after training on ~10, ~100 or ~300 demonstrations per environment (D10, D100, D300). The sets are nested.Spatial GeneralizationNominal uses the training placements for evaluation. Interpolation uses unseen placements between them. Extrapolation uses unseen placements outside their convex hull.Per EnvironmentThe model's score on every environment of the benchmark, e01 onwards. Hover an environment's column for its name.Open model page
#Model# EvalsOpen model
1MolmoAct 2*28%360028%Coming
soon
4%37%44%31%25%Coming soon10%44%52%15%21%4%42%52%18%24%
2Pi 0.5*21%360021%Coming
soon
5%26%33%23%19%Coming soon4%39%66%2%9%3%13%59%13%2%
3DiT-Flow19%360019%Coming
soon
2%24%29%21%16%Coming soon2%17%48%2%0%1%32%63%16%4%
4GR00T N1.7*12%360012%Coming
soon
4%14%19%13%12%Coming soon3%14%25%1%1%1%14%50%13%0%

* All VLAs were started from their official checkpoints, but fine-tuned by us on the target tasks and data regimes.

Environments

Environment group

Evaluation Protocol

Scene Reproducibility: Compare Apples with Apples

A score only means something next to another score measured the same way. Each environment fixes everything a test depends on: the embodiment, the station, the controls, the task, what counts as success, and the limits.

Each condition is evaluated on 30 scenes, and no two are alike: every object starts somewhere else in each one. Before every rollout, every object — every box, every screwdriver, every DC jack, every single screw — is set in its prescribed initial pose. Then the next scene is built from scratch. Every model faces the same 30. Across the benchmark, that is more than 40,000 objects placed exactly where they should be.

1 / 3
Sort screws into bins: placements of four policies compared

Eval Procedure

Each model gets 30 rollouts for every combination of the evaluation axes below.

Evaluation axisSettings# Settings
Environmentse01–e1010
Training-data tierD10, D100, D3003
Placement groupNominal, Interpolation, Extrapolation (coming soon)2
Mid-training data tierM0, M100 (unveiled soon)2
Rollouts per settingPer combination30
Rollouts per modelAll axes3,600
Rollouts across 4 modelsAll axes14,400

Every episode is rated Excellent or Ok when it succeeds, and Failed when it does not. Its subtasks are annotated and rated on the same scale, which shows step by step where a policy fails.

RatingEpisodeSubtask
ExcellentReaches the goal, and every subtask is Excellent.One fluid, efficient motion: no wasted travel, and the grasp holds on the first try.
OkReaches the goal, but at least one subtask is Ok, or Failed and then recovered.Meets its objective, but slowly or with a minor wobble.
FailedStops without reaching the goal.Misses its objective: a dropped object, a missed grasp or a broken motion. A later subtask can still recover it.

The success rate counts Excellent and Ok episodes alike; the Execution quality metric is the Excellent share of those successes.

Every evaluation is reviewed twice, and a sample of them a third time. The operator who ran it rates it first. A second, different operator then reviews it: they verify the placements, annotate the subtasks, and either correct the rating or request a re-evaluation. Finally, engineers review a sample of episodes for quality control and can send a series back for re-annotation or re-evaluation.

Evaluation rated by operatorExcellentOkFailed
Review + annotation by a different operatorCan correct the rating, e.g.ExcellenttoOk
  • Can require redoing an eval
Sampled quality-control episodes reviewed by engineers
  • Can require redoing a series of evals
  • Can require redoing the annotation of a series of evals

Evaluation Conditions

We evaluate all policies across three different axes.

Spatial Generalization

The real world is never set up exactly like the training data. A policy that only succeeds on demonstrated poses has memorised positions rather than learned the task. We therefore place each scene's objects in one of three ways relative to the demonstrations. Nominal poses sit on the training poses themselves. Interpolation poses are new but fall between them. Extrapolation poses fall outside everything the policy was shown. The gap between these three is how much of the task it really understood.

Environment shown: Pour screws into the funnel

MolmoAct 2: 23 of 30 nominal evaluation episodes succeeded: 20 of 30 interpolation evaluation episodes succeeded

Pi 0.5: 28 of 30 nominal evaluation episodes succeeded: 27 of 30 interpolation evaluation episodes succeeded

DiT-Flow: 18 of 30 nominal evaluation episodes succeeded: 16 of 30 interpolation evaluation episodes succeeded

GR00T N1.7: 10 of 30 nominal evaluation episodes succeeded: 2 of 30 interpolation evaluation episodes succeeded

Reproducible evaluation across modelsEvery model runs the same 30 scenes.

Task-Specific Demonstrations

Every demonstration costs operator time, so how fast a policy learns matters as much as how well. Each model is trained three times per environment, on ~10, ~100 or ~300 of its demonstrations (D10, D100, D300). The sets are nested: D10 is part of D100, and D100 is part of D300. Reading across the three shows how quickly each model turns data into skill, and whether more data still pays off.

Mid-Training (coming soon)

Public checkpoints are pretrained on other robots, other cameras and other tasks. Mid-training asks whether a policy gains from first learning the robot itself before it learns an environment. M0 fine-tunes the developer's public checkpoint directly on the task-specific demonstrations. M100 first mid-trains it on 100 hours of embodiment data from our stations, then fine-tunes on the same task-specific demonstrations. Comparing the two shows what experience on the embodiment is worth.

Training Protocol

All training datasets consist of only Excellent teleoperated demonstrations, with no recoveries or suboptimal episodes. For all VLA-based evaluations, each model trains one policy per environment and data tier (D10, D100 or D300; see the environments section to inspect them).

Each training starts from its developer's public checkpoint and recipe. DiT-Flow has no public default checkpoint, so it starts from an ImageNet vision encoder.

Training parameters

SettingMolmoAct 2Pi 0.5DiT-FlowGR00T N1.7
Starting checkpointallenai/MolmoAct2pi0.5 baseFrom scratchnvidia/GR00T-N1.7-3B
Frozen partNoneNoneNoneVision and language towers
Total batch size643225664
Learning rateAction expert 5e-5LLM 1e-5ViT 5e-6Connector 5e-65e-52e-41e-4
ScheduleWarmup, then cosine to 0.1× peakWarmup, then constantWarmup, then cosine to 0Warmup, then cosine to 0
Action chunk30323232
Action spaceDelta commanded TCP poseAbsolute TCP poseAbsolute TCP poseRelative TCP pose