LeRobot documentation

Training Dataset Streaming

You are viewing main version, which requires installation from source. If you'd like regular pip install, checkout the latest stable version (v0.6.1).
Hugging Face's logo
Join the Hugging Face community

and get access to the augmented documentation experience

to get started

Training Dataset Streaming

Training-time dataset streaming lets lerobot-train consume a LeRobotDataset v3 without first downloading its complete Parquet and video payload. Enable it through the existing public switch:

lerobot-train \
  --dataset.repo_id=OWNER/DATASET \
  --dataset.streaming=true \
  --policy.type=act \
  --output_dir=outputs/train/act_streaming

This feature is independent of --dataset.streaming_encoding=true. streaming_encoding controls how videos are written while recording; dataset.streaming controls how an existing dataset is read during training. Recording and rollout encoders are not used by this training path.

How an epoch is read

Each distributed rank owns a deterministic, frame-balanced set of complete episodes. One logical exact-coverage pool per rank mixes those episodes while visiting every selected frame once per rank-local coverage epoch. Parquet columns and compressed MP4 byte ranges are prefetched from the same byte-aware admission frontier. Temporal history and future windows are resolved inside the complete episode, including the same boundary padding masks as map-style loading.

The default map-style LeRobotDataset behavior is unchanged when --dataset.streaming=false.

Workers, shards and concurrency

Streaming does not use “classical” DataLoader workers. With a map-style dataset, each DataLoader worker process samples and decodes its own items. With streaming, the DataLoader runs at most one worker process per rank; all concurrency happens inside that process, behind the DataLoader, in thread pools owned by the dataset:

StageConcurrencySetting
Rank shardingOne disjoint whole-episode shard per rankDistributed world size
Episode fetchParquet rows and MP4 byte ranges, in parallelmax_num_shards, set from --num_workers
Sample assembly and decodeTemporal windows and video frames, in parallel--dataset.streaming_decode_threads
Ordered sample bufferDecoded samples kept ahead, in planner order--dataset.streaming_decoded_queue_size
DataLoader worker (0 or 1)Collates batches off the training process--prefetch_factor

Shards and workers. Each rank owns a frame-balanced shard of complete episodes, so the number of episodes bounds the useful world size: a rank with no episode fails at startup. Inside a rank, --num_workers no longer means “N processes”. In lerobot-train it sets max_num_shards, the number of episodes fetched concurrently (also capped by streaming_episode_pool_size + streaming_prefetch_episodes). Raising it speeds up episode admission. It does not split episodes further, add samplers, or multiply the byte budget. --num_workers=0 keeps the dataset in the training process, and any positive value moves it into a single DataLoader worker.

One worker per rank keeps a single episode pool, byte budget and decoder cache per rank, so memory stays bounded, episode mixing uses the whole pool, and the exactly-once order is deterministic and resumable from a single sample offset. Video decoding mostly releases the GIL, so decode threads scale within that process.

MP4 sidecars and the first run

Video streaming uses an MP4 index sidecar. Dataset initialization first checks the revision-keyed local cache, then looks for a valid published sidecar. If neither is available, LeRobot builds the sidecar locally while holding a process lock and installs it atomically. A failed or interrupted build does not replace the previous valid file.

Hub payload reads are pinned to the metadata snapshot’s commit. Bucket indexes are checked against a fresh object listing, including content hashes, so replacing a video at the same path invalidates its sidecar even if the file size is unchanged. This listing adds startup work, but does not download video payloads. Keep metadata and payloads unchanged during a run. Generic remote filesystems without stable object identities rebuild their index on each open. Existing Bucket sidecars without object fingerprints are rebuilt locally once; publishing the updated sidecar remains an explicit maintainer action.

Training is read-only: it never uploads a sidecar or modifies the dataset repository. On a cluster with node-local caches, the first job may build once per node. A shared LeRobot cache avoids that duplication.

The compressed v3 sidecar is converted once into a read-only, memory-mapped index under $HF_LEROBOT_HOME/streaming-indexes. Conversion is locked and published atomically; ranks sharing that cache reuse the same file-backed array pages instead of decompressing a private copy of every array. Each rank constructs records only for source files used by its episodes. The derived index needs disk space for the uncompressed sample tables, but stores no video payloads. Its first conversion has a cost; subsequent opens reuse it. Replacing a sidecar creates a new immutable mapped-index generation without invalidating live readers.

This does not make all metadata constant-size: episode metadata and span tables still grow with the dataset, and shuffled rank shards can touch overlapping source files. OS-resident mapped pages also count toward RSS; use proportional set size (PSS) when measuring memory shared by ranks. Parquet reads prune row groups using episode statistics and then filter to the complete episode. Groups without statistics remain candidates, so files with no useful bounds can still require a whole-file projected read.

Dataset maintainers can build a sidecar ahead of time:

lerobot-build-mp4-sidecar \
  --repo-id=OWNER/DATASET \
  --revision=COMMIT_SHA \
  --data-root=hf://datasets/OWNER/DATASET@COMMIT_SHA \
  --output=/tmp/dataset-mp4-sidecar.npz

Publication is always explicit. Add --push only after validating the complete-dataset sidecar. Subset sidecars cannot be published.

For dataset repositories, --push creates or updates the dedicated lerobot-sidecars branch. The source dataset branch/tag is unchanged. Training looks there for an index keyed to the pinned video-source commit and validates it before use. A tag and its resolved commit share the same published index; existing revision-keyed local caches remain reusable. Buckets still publish directly under meta/mp4-sidecars/. Publishing requires repository or bucket write access; training only needs read access and never creates a branch or uploads files.

Sidecar schema v3 preserves MP4 composition timing and encoder-delay edits, including B-frame videos. Older sidecars are rebuilt automatically into a separate revision-keyed cache entry; they must also be rebuilt before explicit publication. Reordered videos include the end of the last GOP in each fetched span so decoded frame indices remain contiguous. Repeated or non-unit-rate MP4 edit timelines are rejected rather than silently returning incorrectly aligned frames.

Configuration

The production defaults are:

OptionDefaultMeaning
streaming_episode_pool_size32Maximum complete episodes mixed by each rank
streaming_sampling_strategyremainingRemaining-frame weighting or shuffled round_robin
streaming_prefetch_episodes8Episodes fetched ahead of the active pool
streaming_byte_budget_gb8Maximum synthesized MP4 bytes per rank
streaming_decode_threads2Parallel sample assembly and video decode workers
streaming_decoded_queue_size8Decoded samples buffered ahead, in planner order
video_decoder_cache_sizepool × camerasOpen video decoder LRU cap per rank
streaming_native_http_connectionsunsetNative HTTP connection cap per rank
streaming_native_http_subranges1Concurrent subranges per native HTTP range read
max_num_shards (--num_workers)16Episodes fetched concurrently per rank

All options except max_num_shards are --dataset.* flags. max_num_shards is a StreamingLeRobotDataset argument; lerobot-train sets it from --num_workers (at least 1).

For more even episode mixing, opt in with --dataset.streaming_sampling_strategy=round_robin. Each round shuffles the currently resident episodes and samples one shuffled frame from each; episodes admitted during a round join the next one. Both strategies retain exactly-once epoch coverage and the same resource bounds. Neither is a global uniform permutation: round-robin favors shorter episodes in finite prefixes and can increase decoder churn and reduce throughput. The order comes from --seed (default 1000), together with the policy initialization and the other seeded random generators. Different seeds give different episode and anchor orders, and the same seed repeats the order. --seed=null uses the fixed streaming seed 42. Keep the same strategy, --seed and pool settings when resuming a checkpoint. The checkpoint stores the seed, so a plain --resume keeps the order. A checkpoint from a run that started before --seed reached the streaming dataset used seed 42 for its order: pass --seed=42 when you resume it to keep that order.

The active episode set is capped by both episode count and the exact synthesized mini-MP4 sizes computed from the sidecar. An episode larger than the complete rank budget fails before training fetches its payload. Reservations cover pending fetches, completed prefetches, and episodes still being decoded. At a full budget, speculative prefetch pauses and replacement admission waits for the last decode to release its bytes. EpisodeByteCache.reserved_bytes reports these reservations; resident_bytes reports all completed payloads, even before their first sample is requested.

This is a compressed-payload budget, not a process RSS limit: range-fetch/synthesis scratch buffers, decoder state, Parquet tables, and decoded tensors need additional memory. The decoder LRU and decoded-sample queue have separate limits. Start with a smaller pool or budget on memory-constrained hosts:

lerobot-train \
  --dataset.repo_id=OWNER/DATASET \
  --dataset.streaming=true \
  --dataset.streaming_episode_pool_size=16 \
  --dataset.streaming_prefetch_episodes=4 \
  --dataset.streaming_byte_budget_gb=4 \
  --dataset.streaming_decode_threads=2 \
  --dataset.streaming_decoded_queue_size=8 \
  --num_workers=4 \
  --policy.type=act \
  --output_dir=outputs/train/act_streaming

For a complete dataset stored in an HF Storage Bucket, use the bucket’s identifier and repository type, or pass the bucket URI as --dataset.root=hf://buckets/OWNER/BUCKET. Metadata (meta/), Parquet and video files must all be present at the root of that bucket:

lerobot-train \
  --dataset.repo_id=OWNER/BUCKET \
  --dataset.repo_type=bucket \
  --dataset.streaming=true \
  --policy.type=act \
  --output_dir=outputs/train/act_bucket_streaming

Hub datasets use repo_type=dataset by default. For a complete local dataset, set --dataset.root to its directory. Training resolves metadata and payloads from the same source; no separate streaming payload-location flag is needed.

Resume and shuffle migration

The earlier streaming reader used a bounded row shuffle buffer. The episode reader instead has deterministic exact-coverage ordering derived from the seed and epoch. Checkpoint resume restores the per-rank sample offset using the checkpoint batch size. Changing distributed world size or batch size changes ownership or batch boundaries. For sample-exact comparisons, resume with the same world size and batch size. Keep the same internal fetch concurrency when comparing performance.

The checkpoint metadata does not identify the earlier multi-worker sampler. Its checkpoints cannot be distinguished reliably by the world-size and batch-size checks and must not be used for sample-exact streaming resume. Start a new run from their saved policy weights instead. The resume guarantee concerns anchor order, not bitwise replay of stochastic image transforms. Random image transforms run on the decode threads, in parallel with decoding, so a fixed seed does not reproduce the same augmentations from one run to the next.

Troubleshooting

SymptomWhat to do
The first batch takes a long timeCheck the logs for a sidecar build. Reuse a shared cache or explicitly publish a validated complete sidecar.
A sidecar lock times outAnother process may still be indexing the same revision. Confirm it is healthy before removing a stale lock.
A rank owns no dataReduce the number of ranks so every rank owns at least one selected episode.
The byte budget is exceededLower the episode pool, raise the per-rank byte budget, or use a payload layout with smaller episode ranges.
Remote reads cannot authenticateMake the Hugging Face token or fsspec credentials available to every worker. Sidecars never embed credentials.
Refill stalls are highCompare p95/p99 batch wait, reduce network contention, raise prefetch gradually, check the sidecar matches revision.
Update on GitHub