Reality Check Leaderboard
PAW-GEN-10
Data Scaling
D10, D100, and D300 indicate how many demonstration episodes per environment were used to train the model. Whiskers are 95% confidence intervals over the evaluation episodes.
Spatial Generalization
Nominal uses training placements; interpolation uses unseen placements between them; extrapolation uses unseen placements outside their convex hull. Whiskers are 95% confidence intervals over the evaluation episodes.
Per Environment (SuccessProgressExecution qualityAvg execution timeExecution speedSmoothnessAvg jerkSafe failure rateMax contact force)
e01–e10 represent the environments in PAW-GEN-10, each value in the selected metric averaged across every data tier and placement condition.
Share of episodes rated Ok or Excellent.
Share of the task's steps completed successfully.
Excellent-vs-Ok share of successful episodes.
Mean active duration of successful episodes. Lower is better.
Task steps completed per active minute, pooled over the successful episodes.
Spectral arc length (SPARC) of the arm's motion, a unitless measure ≤ -1: the closer to -1, the smoother. It is measured over each whole episode and averaged over episodes.
How abruptly the arm's acceleration changed: the end effector's jerk, filtered at 10 Hz and averaged over each episode, then over episodes. It grows with speed, so a faster motion reads as rougher. Lower is better.
Share of failed episodes that ended with both arms ok, from the arms' own status when the recording stopped.
The largest force either arm exerted on its surroundings during an episode, averaged over episodes. It is based on the FR3's joint torque sensors. Lower is better.
| Rank and model | Overall score and number of evaluations | Mid-training DataM0 has no mid-training data and post-trains the model directly on the environment demonstrations. M100 first mid-trains it on 100 hours of embodiment data (coming soon). | Data ScalingScores after training on ~10, ~100 or ~300 demonstrations per environment (D10, D100, D300). The sets are nested. | Spatial GeneralizationNominal uses the training placements for evaluation. Interpolation uses unseen placements between them. Extrapolation uses unseen placements outside their convex hull. | Per EnvironmentThe model's score on every environment of the benchmark, e01 onwards. Hover an environment's column for its name. | Open model page | ||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| # | Model | # Evals | Open model | |||||||||||||||||||
| 1 | MolmoAct 2* | 28% | 3600 | 28% | Coming soon | 4% | 37% | 44% | 31% | 25% | Coming soon | 10% | 44% | 52% | 15% | 21% | 4% | 42% | 52% | 18% | 24% | |
| 2 | Pi 0.5* | 21% | 3600 | 21% | Coming soon | 5% | 26% | 33% | 23% | 19% | Coming soon | 4% | 39% | 66% | 2% | 9% | 3% | 13% | 59% | 13% | 2% | |
| 3 | DiT-Flow | 19% | 3600 | 19% | Coming soon | 2% | 24% | 29% | 21% | 16% | Coming soon | 2% | 17% | 48% | 2% | 0% | 1% | 32% | 63% | 16% | 4% | |
| 4 | GR00T N1.7* | 12% | 3600 | 12% | Coming soon | 4% | 14% | 19% | 13% | 12% | Coming soon | 3% | 14% | 25% | 1% | 1% | 1% | 14% | 50% | 13% | 0% | |
* All VLAs were started from their official checkpoints, but fine-tuned by us on the target tasks and data regimes.
Environments
- e01Sort screws into bins
Specification of Sort screws into bins
- Embodiment
- FR3 Duo [REAL]
- Scenario
- PAW benchmarking station v1
- Allowed controls
- Full arm and gripper control at 30 Hz
- Task
- Sort screws into bins
- Success criteria
- 3 black screws in the red bin
- 2 silver screws in the black bin
- Constraints
- Max episode duration: 120 s
- One screw per grasp
- Black → red bin
- Silver → black bin
- Steps
- Pick up a single black screw from the table. ×3
- Pick up a single silver screw from the table. ×2
- Place the screw held by the gripper into its respective bin: a black screw in the red bin, a silver screw in the black bin. ×3
Task-specific training datasets - e03Pour screws into the funnel
Specification of Pour screws into the funnel
- Embodiment
- FR3 Duo [REAL]
- Scenario
- PAW benchmarking station v1
- Allowed controls
- Full arm and gripper control at 30 Hz
- Task
- Pour screws into the funnel
- Success criteria
- Screws in the funnel
- Container back in its area
- Constraints
- Max episode duration: 90 s
- No screws poured outside the funnel
- Steps
- Pick up the container with screws.
- Pour the screws into the funnel.
- Place the container back on the designated area.
Task-specific training datasets - e04Connect DC jack
Specification of Connect DC jack
- Embodiment
- FR3 Duo [REAL]
- Scenario
- PAW benchmarking station v1
- Allowed controls
- Full arm and gripper control at 30 Hz
- Task
- Connect DC jack
- Success criteria
- Both pairs plugged in
- Both pairs on the table
- Constraints
- Max episode duration: 100 s
- One pair at a time
- Steps
- Pick up the first male DC jack connector.
- Pick up the first female DC jack connector.
- Connect the male and female DC jack connectors.
- Place the connected DC jack on the table.
- Pick up the second female DC jack connector.
- Pick up the second male DC jack connector.
- Connect the male and female DC jack connectors.
- Place the connected DC jack on the table.
Task-specific training datasets - e05Route cable through closed hoops
Specification of Route cable through closed hoops
- Embodiment
- FR3 Duo [REAL]
- Scenario
- PAW benchmarking station v1
- Allowed controls
- Full arm and gripper control at 30 Hz
- Task
- Route cable through closed hoops
- Success criteria
- Through the blue and white hoops
- End inside the orange hoop
- Constraints
- Max episode duration: 115 s
- Order: blue, white, then orange
- Steps
- Pick up the ethernet cable from the table.
- Insert the cable end into the first hoop.
- Push the cable through the first hoop until the end comes out the other side.
- Grab the cable end coming out from the other side of the first hoop.
- Insert the cable end into the second hoop.
- Push the cable through the second hoop until the end comes out the other side.
- Grab the cable end coming out from the other side of the second hoop.
- Insert the cable end into the third hoop.
Task-specific training datasets - e09Open the toolbox with a screwdriver
Specification of Open the toolbox with a screwdriver
- Embodiment
- FR3 Duo [REAL]
- Scenario
- PAW benchmarking station v1
- Allowed controls
- Full arm and gripper control at 30 Hz
- Task
- Open the toolbox with a screwdriver
- Success criteria
- Both latches open
- Screwdriver on the table
- Lid fully open
- Constraints
- Max episode duration: 90 s
- Toolbox pushed to the center without grasping it
- Latches opened using the screwdriver
- Steps
- Push the toolbox, without grasping it, to the center of the workspace.
- Pick up the screwdriver.
- Reorient the screwdriver in the gripper.
- Release the first toolbox latch with the screwdriver.
- Release the second toolbox latch with the screwdriver.
- Place the screwdriver on the table.
- Pick up the toolbox handle.
- Lift up the toolbox lid with the left arm.
- Put the toolbox lid down with the right arm.
Task-specific training datasets - e02Datum-corner alignmentHeld-out benchmark environmentSpecificationHeld outTask-specific training datasets
- Held out
- e06Put tools in standing toolboxHeld-out benchmark environmentSpecificationHeld outTask-specific training datasets
- Held out
- e07Place loaded boxes in container (bimanual)Held-out benchmark environmentSpecificationHeld outTask-specific training datasets
- Held out
- e08Place spray bottles upright, nozzle forward (bimanual handover)Held-out benchmark environmentSpecificationHeld outTask-specific training datasets
- Held out
- e10Clamp electrical componentHeld-out benchmark environmentSpecificationHeld outTask-specific training datasets
- Held out
Evaluation Protocol
Scene Reproducibility: Compare Apples with Apples
A score only means something next to another score measured the same way. Each environment fixes everything a test depends on: the embodiment, the station, the controls, the task, what counts as success, and the limits.
Each condition is evaluated on 30 scenes, and no two are alike: every object starts somewhere else in each one. Before every rollout, every object — every box, every screwdriver, every DC jack, every single screw — is set in its prescribed initial pose. Then the next scene is built from scratch. Every model faces the same 30. Across the benchmark, that is more than 40,000 objects placed exactly where they should be.
Eval Procedure
Each model gets 30 rollouts for every combination of the evaluation axes below.
| Evaluation axis | Settings | # Settings |
|---|---|---|
| Environments | e01–e10 | 10 |
| Training-data tier | D10, D100, D300 | 3 |
| Placement group | Nominal, Interpolation, Extrapolation (coming soon) | 2 |
| Mid-training data tier | M0, M100 (unveiled soon) | 2 |
| Rollouts per setting | Per combination | 30 |
| Rollouts per model | All axes | 3,600 |
| Rollouts across 4 models | All axes | 14,400 |
Every episode is rated Excellent or Ok when it succeeds, and Failed when it does not. Its subtasks are annotated and rated on the same scale, which shows step by step where a policy fails.
| Rating | Episode | Subtask |
|---|---|---|
| Reaches the goal, and every subtask is Excellent. | One fluid, efficient motion: no wasted travel, and the grasp holds on the first try. | |
| Reaches the goal, but at least one subtask is Ok, or Failed and then recovered. | Meets its objective, but slowly or with a minor wobble. | |
| Stops without reaching the goal. | Misses its objective: a dropped object, a missed grasp or a broken motion. A later subtask can still recover it. |
The success rate counts Excellent and Ok episodes alike; the Execution quality metric is the Excellent share of those successes.
Every evaluation is reviewed twice, and a sample of them a third time. The operator who ran it rates it first. A second, different operator then reviews it: they verify the placements, annotate the subtasks, and either correct the rating or request a re-evaluation. Finally, engineers review a sample of episodes for quality control and can send a series back for re-annotation or re-evaluation.
- Can require redoing an eval
- Can require redoing a series of evals
- Can require redoing the annotation of a series of evals
Evaluation Conditions
We evaluate all policies across three different axes.
Spatial Generalization
The real world is never set up exactly like the training data. A policy that only succeeds on demonstrated poses has memorised positions rather than learned the task. We therefore place each scene's objects in one of three ways relative to the demonstrations. Nominal poses sit on the training poses themselves. Interpolation poses are new but fall between them. Extrapolation poses fall outside everything the policy was shown. The gap between these three is how much of the task it really understood.
Environment shown: Pour screws into the funnel

MolmoAct 2: 23 of 30 nominal evaluation episodes succeeded: 20 of 30 interpolation evaluation episodes succeeded
Pi 0.5: 28 of 30 nominal evaluation episodes succeeded: 27 of 30 interpolation evaluation episodes succeeded
DiT-Flow: 18 of 30 nominal evaluation episodes succeeded: 16 of 30 interpolation evaluation episodes succeeded
GR00T N1.7: 10 of 30 nominal evaluation episodes succeeded: 2 of 30 interpolation evaluation episodes succeeded
Task-Specific Demonstrations
Every demonstration costs operator time, so how fast a policy learns matters as much as how well. Each model is trained three times per environment, on ~10, ~100 or ~300 of its demonstrations (D10, D100, D300). The sets are nested: D10 is part of D100, and D100 is part of D300. Reading across the three shows how quickly each model turns data into skill, and whether more data still pays off.
Mid-Training (coming soon)
Public checkpoints are pretrained on other robots, other cameras and other tasks. Mid-training asks whether a policy gains from first learning the robot itself before it learns an environment. M0 fine-tunes the developer's public checkpoint directly on the task-specific demonstrations. M100 first mid-trains it on 100 hours of embodiment data from our stations, then fine-tunes on the same task-specific demonstrations. Comparing the two shows what experience on the embodiment is worth.
Training Protocol
All training datasets consist of only Excellent teleoperated demonstrations, with no recoveries or suboptimal episodes. For all VLA-based evaluations, each model trains one policy per environment and data tier (D10, D100 or D300; see the environments section to inspect them).
Each training starts from its developer's public checkpoint and recipe. DiT-Flow has no public default checkpoint, so it starts from an ImageNet vision encoder.
Training parameters
| Setting | MolmoAct 2 | Pi 0.5 | DiT-Flow | GR00T N1.7 |
|---|---|---|---|---|
| Starting checkpoint | allenai/MolmoAct2 | pi0.5 base | From scratch | nvidia/GR00T-N1.7-3B |
| Frozen part | None | None | None | Vision and language towers |
| Total batch size | 64 | 32 | 256 | 64 |
| Learning rate | Action expert 5e-5LLM 1e-5ViT 5e-6Connector 5e-6 | 5e-5 | 2e-4 | 1e-4 |
| Schedule | Warmup, then cosine to 0.1× peak | Warmup, then constant | Warmup, then cosine to 0 | Warmup, then cosine to 0 |
| Action chunk | 30 | 32 | 32 | 32 |
| Action space | Delta commanded TCP pose | Absolute TCP pose | Absolute TCP pose | Relative TCP pose |