Request access

We Deleted 71 Robot Demos From NVIDIA's Dataset… and GR00T Got Better

Skyu found the 71 demos holding back NVIDIA's GR00T N1.5. Removing them lifted task success by 4 points. Removing 71 random demos did nothing.

10,055episodes scored
71episodes flagged (severity 6+)
+4 ptstask success after removing flagged episodes
0 ptsafter removing same number of random episodes

GR00T gr1_arms_waist.PlateToCardboardBox · GR-1 humanoid · RoboCasa tabletop

What we ran

THE DATA

gr1_arms_waist.PlateToCardboardBox
  • 10,055 teleop episodes, ~2M frames
  • GR-1 humanoid, RoboCasa tabletop sim
  • Pick target object from plate, place into box. 2 objects on plate

THE SCORING

Skyu scores every episode
  • Each episode compared against what a lightly trained policy expects
  • Single A10 GPU
  • Runs on-prem; no data leaves building

THE OUTPUT

Ranked manifest + HTML report
  • Severity 0–10 per episode with reasoning
  • Deterministic: rerun gives the same ranking, bit for bit
  • Full manifest, eval logs available

What flagged demos look like

SAMPLE FLAGGED DEMOObjects knocked over
SAMPLE FLAGGED DEMOObjects disappear mid-demo without reason

The test: train with and without them

ARMTRAINING SETEPISODES
A · Control929 random + 71 flagged1000
B · Flagged removedwithout 71 flagged929
C · Random removedwithout 71 random episodes929

Same recipe for every arm

  • GR00T N1.5 fine-tune, diffusion transformer tuned, per NVIDIA's recipe
  • 10000 steps, batch 32 (about 1.6 passes over the data), 2 seeds per arm: 6 models, ~78 min each on one A100
  • Evaluated in NVIDIA's RoboCasa GR-1 tabletop sim: 50 fixed scenes, 720 max steps, same scenes for every model

Why only 1000 episodes

The 71 flagged episodes are 0.7% of the full set. Removing 0.7% from a fine-tuned foundation model would show nothing either way. At 7% of the training set the effect is measurable if it exists. The claim is therefore scoped: in a 1000-episode training set, what do these 71 cost?

Removed flagged episodes: +4 points. Removed random ones: 0

Task success rate50 fixed scenes · 2 seeds per arm
NVIDIA's released checkpoint (control) A · 1,000 episodes B · removed 71 flagged C · removed 71 random
+4 ptsBetter task success rate without 71 flagged demos
0 ptsRemoving 71 random episodes did nothing

What this does and does not show

The evaluation is sound

NVIDIA's released checkpoint scores 42% (21/50) on our setup; NVIDIA reports 44% for this task. Every number in this post is on NVIDIA's own footing: NVIDIA's model, NVIDIA's recipe, NVIDIA's simulator.

Direction is right, size is unproven

+4 for flagged removal, 0 for random, over 100 rollouts per arm. But arm A's two seeds differ by 6 points, so this could still be luck. Five seeds per arm settles it: nine more fine-tunes, about 90 minutes on 8×A100, one day end to end.

Why the effect is small here

Sim teleop data is clean by construction: 0.7% anomalous. We had to concentrate the flagged episodes to 7% to measure anything. Real-robot teleop (resets, operator errors, dropouts) flags more, and each flag costs more. This benchmark is the hard case for the tool.

Run Skyu on your dataset. Request access