Pareto ships per-episode reductions only. Its velocity-debiasing direction depends on a cross-episode number, and that number isn't shipped. So I computed it, on real public data, with no model in the loop.
Each bar is one demonstration. Red bars are the 5% tails. The dashed line is the task mean.
All four are computed inside one episode. None aggregates across demonstrations.
The same signal peak_velocity reads, aggregated across the demos of a task. This is the input velocity debiasing needs.
Twenty demos whose speed, not skill, drives the variance. These are the prune candidates a cross-episode metric would flag.
Computed from lerobot/pusht, 25,650 frames across 206 demos. No model in the loop. Reproduce with node local_test.mjs.