Request access

[ Data scoring for physical AI ]

Train your model on your best data.

Skyu scores every training data sample by its real impact on your model, so you can train on the best subset that improves your model.

Find the demonstrations that are hurting your policy.

Skyu scores every episode in a robot-learning dataset against your trained policy, flags the ones worth a closer look, and writes a reproducible report. It runs offline, on your hardware.

$ pip install skyu-0.2.1-py3-none-any.whl
$ skyu score --dataset ./pusht --checkpoint ./ckpt/030000
  scored 1,204 episodes
  flagged 96 for review
  → report.html
report.html skyu score · 1,204 episodes
Score distribution one score per episode
flag threshold
0.00score →0.92
Episodes by rank known-bad episodes marked red
rank 1 · highest scorerank 1,204

Placeholder data — 12 injected corruptions, all recovered among the flagged episodes.

[ HOW IT WORKS ]

01

Point

Give Skyu a LeRobot-format dataset and a policy checkpoint. Nothing else — no labels, no cloud account, no instrumentation of your training run.

skyu score --dataset D --checkpoint C
02

Score

Skyu scores every episode in your dataset and ranks them, so the demonstrations most likely to be hurting your policy rise to the top.

every episode scored and ranked
03

Review

The report flags the episodes worth a closer look, and lists the exact command behind every number, so anyone on your team can rerun it and get the same result.

→ report.html · single file