pulsatrix
Loading...
Searching...
No Matches
pulsatrix::IterableDataset Class Referenceabstract

Streaming dataset abstraction for sources with no random access or no known length (sharded files, generators) – pulsatrix's analogue of PyTorch's IterableDataset / tf.data's source-op model. DataLoader treats this and Dataset via a common internal adapter (data_loader.hpp) so both share one fetch/collate path. More...

#include <iterable_dataset.hpp>

Public Member Functions

virtual ~IterableDataset ()=default
 
virtual void reset ()=0
 Resets iteration to the beginning. Called once per epoch by DataLoader.
 
virtual std::optional< Sample > next ()=0
 Fetches the next sample.
 

Detailed Description

Streaming dataset abstraction for sources with no random access or no known length (sharded files, generators) – pulsatrix's analogue of PyTorch's IterableDataset / tf.data's source-op model. DataLoader treats this and Dataset via a common internal adapter (data_loader.hpp) so both share one fetch/collate path.

Note
Sharding across multiple DataLoader workers (each worker owning a distinct stream shard) is out of scope for campaign_exai_dl_library_data_pipeline's Phase 1 – DataLoader guards num_workers > 1 against an IterableDataset rather than silently producing duplicate/missing records.

Constructor & Destructor Documentation

◆ ~IterableDataset()

virtual pulsatrix::IterableDataset::~IterableDataset ( )
virtualdefault

Member Function Documentation

◆ next()

virtual std::optional< Sample > pulsatrix::IterableDataset::next ( )
pure virtual

Fetches the next sample.

Returns
The next Sample, or std::nullopt once the stream is exhausted.

◆ reset()

virtual void pulsatrix::IterableDataset::reset ( )
pure virtual

Resets iteration to the beginning. Called once per epoch by DataLoader.


The documentation for this class was generated from the following file: