# We Deleted 71 Robot Demos From NVIDIA's Dataset… and GR00T Got Better

Skyu, Inc. · Berkeley SkyDeck · NVIDIA Inception
Published: 2026-10-01 · Category: Research · URL: /proof

## Summary

Skyu found the 71 demos holding back NVIDIA's GR00T N1.5. Removing them lifted task success by 4 points. Removing 71 random demos did nothing.

- 10,055 episodes scored
- 71 episodes flagged (severity 6+)
- +4 pts task success after removing flagged episodes
- 0 pts after removing the same number of random episodes

Task: GR00T gr1_arms_waist.PlateToCardboardBox · GR-1 humanoid · RoboCasa tabletop

## What we ran

### The data
- gr1_arms_waist.PlateToCardboardBox
- 10,055 teleop episodes, ~2M frames
- GR-1 humanoid, RoboCasa tabletop sim
- Pick target object from plate, place into box; 2 objects on plate

### The scoring
- Skyu scores every episode
- Each episode compared against what a lightly trained policy expects
- Single A10 GPU
- Runs on-prem; no data leaves the building

### The output
- Ranked manifest + HTML report
- Severity 0–10 per episode with reasoning
- Deterministic: rerun gives the same ranking, bit for bit
- Full manifest, eval logs available

## What flagged demos look like

Two sample flagged demos (severity 6+, removed in arm B):
- Objects knocked over
- Objects disappear mid-demo without reason

## The test: train with and without them

| Arm | Training set | Episodes |
|---|---|---|
| A · Control | 929 random + 71 flagged | 1,000 |
| B · Flagged removed | without 71 flagged | 929 |
| C · Random removed | without 71 random episodes | 929 |

Same recipe for every arm:
- GR00T N1.5 fine-tune, diffusion transformer tuned, per NVIDIA's recipe
- 10,000 steps, batch 32 (about 1.6 passes over the data), 2 seeds per arm: 6 models, ~78 min each on one A100
- Evaluated in NVIDIA's RoboCasa GR-1 tabletop sim: 50 fixed scenes, 720 max steps, same scenes for every model

Why only 1,000 episodes: the 71 flagged episodes are 0.7% of the full set. Removing 0.7% from a fine-tuned foundation model would show nothing either way. At 7% of the training set the effect is measurable if it exists. The claim is therefore scoped: in a 1,000-episode training set, what do these 71 cost?

## Results

Removed flagged episodes: +4 points. Removed random ones: 0.

| Model | Task success |
|---|---|
| NVIDIA's released checkpoint (control) | 42% |
| A · 1,000 episodes | 29% |
| B · removed 71 flagged | 33% |
| C · removed 71 random | 29% |

- +4 pts: better task success rate without 71 flagged demos
- 0 pts: removing 71 random episodes did nothing

## What this does and does not show

**The evaluation is sound.** NVIDIA's released checkpoint scores 42% (21/50) on our setup; NVIDIA reports 44% for this task. Every number here is on NVIDIA's own footing: NVIDIA's model, NVIDIA's recipe, NVIDIA's simulator.

**Direction is right, size is unproven.** +4 for flagged removal, 0 for random, over 100 rollouts per arm. But arm A's two seeds differ by 6 points, so this could still be luck. Five seeds per arm settles it: nine more fine-tunes, about 90 minutes on 8×A100, one day end to end.

**Why the effect is small here.** Sim teleop data is clean by construction: 0.7% anomalous. We had to concentrate the flagged episodes to 7% to measure anything. Real-robot teleop (resets, operator errors, dropouts) flags more, and each flag costs more. This benchmark is the hard case for the tool.

