Streaming dataset abstraction for sources with no random access or no known length (sharded files, generators) – pulsatrix's analogue of PyTorch's IterableDataset / tf.data's source-op model. DataLoader treats this and Dataset via a common internal adapter (data_loader.hpp) so both share one fetch/collate path.
More...
#include <iterable_dataset.hpp>
Streaming dataset abstraction for sources with no random access or no known length (sharded files, generators) – pulsatrix's analogue of PyTorch's IterableDataset / tf.data's source-op model. DataLoader treats this and Dataset via a common internal adapter (data_loader.hpp) so both share one fetch/collate path.
- Note
- Sharding across multiple DataLoader workers (each worker owning a distinct stream shard) is out of scope for campaign_exai_dl_library_data_pipeline's Phase 1 – DataLoader guards num_workers > 1 against an IterableDataset rather than silently producing duplicate/missing records.
◆ ~IterableDataset()
| virtual pulsatrix::IterableDataset::~IterableDataset |
( |
| ) |
|
|
virtualdefault |
◆ next()
| virtual std::optional< Sample > pulsatrix::IterableDataset::next |
( |
| ) |
|
|
pure virtual |
Fetches the next sample.
- Returns
- The next Sample, or std::nullopt once the stream is exhausted.
◆ reset()
| virtual void pulsatrix::IterableDataset::reset |
( |
| ) |
|
|
pure virtual |
Resets iteration to the beginning. Called once per epoch by DataLoader.
The documentation for this class was generated from the following file: