LeRobot documentation
Training Dataset Streaming
Training Dataset Streaming
Training-time dataset streaming lets lerobot-train consume a LeRobotDataset v3 without first
downloading its complete Parquet and video payload. Enable it through the existing public switch:
lerobot-train \
--dataset.repo_id=OWNER/DATASET \
--dataset.streaming=true \
--policy.type=act \
--output_dir=outputs/train/act_streamingThis feature is independent of --dataset.streaming_encoding=true. streaming_encoding controls
how videos are written while recording; dataset.streaming controls how an existing dataset is read
during training. Recording and rollout encoders are not used by this training path.
How an epoch is read
Each distributed rank owns a deterministic, frame-balanced set of complete episodes. One logical exact-coverage pool per rank mixes those episodes while visiting every selected frame once per rank-local coverage epoch. Parquet columns and compressed MP4 byte ranges are prefetched from the same byte-aware admission frontier. Temporal history and future windows are resolved inside the complete episode, including the same boundary padding masks as map-style loading.
The default map-style LeRobotDataset behavior is unchanged when --dataset.streaming=false.
Workers, shards and concurrency
Streaming does not use “classical” DataLoader workers. With a map-style dataset, each DataLoader worker process samples and decodes its own items. With streaming, the DataLoader runs at most one worker process per rank; all concurrency happens inside that process, behind the DataLoader, in thread pools owned by the dataset:
| Stage | Concurrency | Setting |
|---|---|---|
| Rank sharding | One disjoint whole-episode shard per rank | Distributed world size |
| Episode fetch | Parquet rows and MP4 byte ranges, in parallel | max_num_shards, set from --num_workers |
| Sample assembly and decode | Temporal windows and video frames, in parallel | --dataset.streaming_decode_threads |
| Ordered sample buffer | Decoded samples kept ahead, in planner order | --dataset.streaming_decoded_queue_size |
| DataLoader worker (0 or 1) | Collates batches off the training process | --prefetch_factor |
Shards and workers. Each rank owns a frame-balanced shard of complete episodes, so the number
of episodes bounds the useful world size: a rank with no episode fails at startup. Inside a
rank, --num_workers no longer means “N processes”. In lerobot-train it sets max_num_shards, the number of episodes fetched concurrently (also capped by streaming_episode_pool_size + streaming_prefetch_episodes). Raising it speeds up episode
admission. It does not split episodes further, add samplers, or multiply the byte budget. --num_workers=0 keeps the dataset in the training process, and any positive value moves it
into a single DataLoader worker.
One worker per rank keeps a single episode pool, byte budget and decoder cache per rank, so memory stays bounded, episode mixing uses the whole pool, and the exactly-once order is deterministic and resumable from a single sample offset. Video decoding mostly releases the GIL, so decode threads scale within that process.
MP4 sidecars and the first run
Video streaming uses an MP4 index sidecar. Dataset initialization first checks the revision-keyed local cache, then looks for a valid published sidecar. If neither is available, LeRobot builds the sidecar locally while holding a process lock and installs it atomically. A failed or interrupted build does not replace the previous valid file.
Hub payload reads are pinned to the metadata snapshot’s commit. Bucket indexes are checked against a fresh object listing, including content hashes, so replacing a video at the same path invalidates its sidecar even if the file size is unchanged. This listing adds startup work, but does not download video payloads. Keep metadata and payloads unchanged during a run. Generic remote filesystems without stable object identities rebuild their index on each open. Existing Bucket sidecars without object fingerprints are rebuilt locally once; publishing the updated sidecar remains an explicit maintainer action.
Training is read-only: it never uploads a sidecar or modifies the dataset repository. On a cluster with node-local caches, the first job may build once per node. A shared LeRobot cache avoids that duplication.
The compressed v3 sidecar is converted once into a read-only, memory-mapped index under $HF_LEROBOT_HOME/streaming-indexes. Conversion is locked and published atomically; ranks
sharing that cache reuse the same file-backed array pages instead of decompressing a private
copy of every array. Each rank constructs records only for source files used by its episodes.
The derived index needs disk space for the uncompressed sample tables, but stores no video
payloads. Its first conversion has a cost; subsequent opens reuse it. Replacing a sidecar
creates a new immutable mapped-index generation without invalidating live readers.
This does not make all metadata constant-size: episode metadata and span tables still grow with the dataset, and shuffled rank shards can touch overlapping source files. OS-resident mapped pages also count toward RSS; use proportional set size (PSS) when measuring memory shared by ranks. Parquet reads prune row groups using episode statistics and then filter to the complete episode. Groups without statistics remain candidates, so files with no useful bounds can still require a whole-file projected read.
Dataset maintainers can build a sidecar ahead of time:
lerobot-build-mp4-sidecar \ --repo-id=OWNER/DATASET \ --revision=COMMIT_SHA \ --data-root=hf://datasets/OWNER/DATASET@COMMIT_SHA \ --output=/tmp/dataset-mp4-sidecar.npz
Publication is always explicit. Add --push only after validating the complete-dataset sidecar.
Subset sidecars cannot be published.
For dataset repositories, --push creates or updates the dedicated lerobot-sidecars branch.
The source dataset branch/tag is unchanged. Training looks there for an index keyed to the
pinned video-source commit and validates it before use. A tag and its resolved commit share
the same published index; existing revision-keyed local caches remain reusable. Buckets still
publish directly under meta/mp4-sidecars/. Publishing requires repository or bucket write access;
training only needs read access and never creates a branch or uploads files.
Sidecar schema v3 preserves MP4 composition timing and encoder-delay edits, including B-frame videos. Older sidecars are rebuilt automatically into a separate revision-keyed cache entry; they must also be rebuilt before explicit publication. Reordered videos include the end of the last GOP in each fetched span so decoded frame indices remain contiguous. Repeated or non-unit-rate MP4 edit timelines are rejected rather than silently returning incorrectly aligned frames.
Configuration
The production defaults are:
| Option | Default | Meaning |
|---|---|---|
streaming_episode_pool_size | 32 | Maximum complete episodes mixed by each rank |
streaming_sampling_strategy | remaining | Remaining-frame weighting or shuffled round_robin |
streaming_prefetch_episodes | 8 | Episodes fetched ahead of the active pool |
streaming_byte_budget_gb | 8 | Maximum synthesized MP4 bytes per rank |
streaming_decode_threads | 2 | Parallel sample assembly and video decode workers |
streaming_decoded_queue_size | 8 | Decoded samples buffered ahead, in planner order |
video_decoder_cache_size | pool × cameras | Open video decoder LRU cap per rank |
streaming_native_http_connections | unset | Native HTTP connection cap per rank |
streaming_native_http_subranges | 1 | Concurrent subranges per native HTTP range read |
max_num_shards (--num_workers) | 16 | Episodes fetched concurrently per rank |
All options except max_num_shards are --dataset.* flags. max_num_shards is a StreamingLeRobotDataset argument; lerobot-train sets it from --num_workers (at least 1).
For more even episode mixing, opt in with --dataset.streaming_sampling_strategy=round_robin.
Each round shuffles the currently resident episodes and samples one shuffled frame from each;
episodes admitted during a round join the next one. Both strategies retain exactly-once epoch
coverage and the same resource bounds. Neither is a global uniform permutation: round-robin
favors shorter episodes in finite prefixes and can increase decoder churn and reduce throughput.
The order comes from --seed (default 1000), together with the policy initialization and the
other seeded random generators. Different seeds give different episode and anchor orders, and the
same seed repeats the order. --seed=null uses the fixed streaming seed 42. Keep the same strategy, --seed and pool settings when resuming a checkpoint. The checkpoint stores the seed, so a plain --resume keeps the order. A checkpoint from a run that started before --seed reached the
streaming dataset used seed 42 for its order: pass --seed=42 when you resume it to keep that order.
The active episode set is capped by both episode count and the exact synthesized mini-MP4 sizes
computed from the sidecar. An episode larger than the complete rank budget fails before training
fetches its payload. Reservations cover pending fetches, completed prefetches, and episodes still
being decoded. At a full budget, speculative prefetch pauses and replacement admission waits for
the last decode to release its bytes. EpisodeByteCache.reserved_bytes reports these reservations; resident_bytes reports all completed payloads, even before their first sample is requested.
This is a compressed-payload budget, not a process RSS limit: range-fetch/synthesis scratch buffers, decoder state, Parquet tables, and decoded tensors need additional memory. The decoder LRU and decoded-sample queue have separate limits. Start with a smaller pool or budget on memory-constrained hosts:
lerobot-train \
--dataset.repo_id=OWNER/DATASET \
--dataset.streaming=true \
--dataset.streaming_episode_pool_size=16 \
--dataset.streaming_prefetch_episodes=4 \
--dataset.streaming_byte_budget_gb=4 \
--dataset.streaming_decode_threads=2 \
--dataset.streaming_decoded_queue_size=8 \
--num_workers=4 \
--policy.type=act \
--output_dir=outputs/train/act_streamingFor a complete dataset stored in an HF Storage Bucket, use the bucket’s identifier and repository
type, or pass the bucket URI as --dataset.root=hf://buckets/OWNER/BUCKET. Metadata (meta/),
Parquet and video files must all be present at the root of that bucket:
lerobot-train \
--dataset.repo_id=OWNER/BUCKET \
--dataset.repo_type=bucket \
--dataset.streaming=true \
--policy.type=act \
--output_dir=outputs/train/act_bucket_streamingHub datasets use repo_type=dataset by default. For a complete local dataset, set --dataset.root to its directory. Training resolves metadata and payloads from the same source; no separate
streaming payload-location flag is needed.
Resume and shuffle migration
The earlier streaming reader used a bounded row shuffle buffer. The episode reader instead has deterministic exact-coverage ordering derived from the seed and epoch. Checkpoint resume restores the per-rank sample offset using the checkpoint batch size. Changing distributed world size or batch size changes ownership or batch boundaries. For sample-exact comparisons, resume with the same world size and batch size. Keep the same internal fetch concurrency when comparing performance.
The checkpoint metadata does not identify the earlier multi-worker sampler. Its checkpoints cannot be distinguished reliably by the world-size and batch-size checks and must not be used for sample-exact streaming resume. Start a new run from their saved policy weights instead. The resume guarantee concerns anchor order, not bitwise replay of stochastic image transforms. Random image transforms run on the decode threads, in parallel with decoding, so a fixed seed does not reproduce the same augmentations from one run to the next.
Troubleshooting
| Symptom | What to do |
|---|---|
| The first batch takes a long time | Check the logs for a sidecar build. Reuse a shared cache or explicitly publish a validated complete sidecar. |
| A sidecar lock times out | Another process may still be indexing the same revision. Confirm it is healthy before removing a stale lock. |
| A rank owns no data | Reduce the number of ranks so every rank owns at least one selected episode. |
| The byte budget is exceeded | Lower the episode pool, raise the per-rank byte budget, or use a payload layout with smaller episode ranges. |
| Remote reads cannot authenticate | Make the Hugging Face token or fsspec credentials available to every worker. Sidecars never embed credentials. |
| Refill stalls are high | Compare p95/p99 batch wait, reduce network contention, raise prefetch gradually, check the sidecar matches revision. |