Data Loading, Transformation & Validation¶
This section covers how to get data from files on disk into batched Tensors for training or
explanation. Use it to read CSV tables, image folders, text corpora, WAV audio or extracted
video frames, transform them, and check them for data-quality problems.
The generic core is Dataset/IterableDataset, DataLoader, Transform/Compose and
CollateFn. Everything else is a concrete implementation built on that core, one per modality:
tabular, image, text, audio and video (frames only). Generic dataset validation sits alongside.
What's inside¶
- Core:
Sample,Dataset,IterableDataset(streaming),Sampler(SequentialSampler/ShuffleSampler),Transform/Compose/TransformDataset,Batch/CollateFn/DefaultCollate,DataLoader/DataLoaderOptions.BoundedQueue/DataThreadPoolare built but not yet wired intoDataLoader(see below). - Tabular:
CsvReader/CsvTable,CsvDataset - Image:
ImageDecoder(backed by stb_image),ResizeTransform/CenterCropTransform/NormalizeTransform/HorizontalFlipTransform,ImageFolderDataset - Text:
Tokenizer,Vocabulary/BuildVocabulary,TextDataset,PadCollate - Audio:
WavReader/WavData,AudioFolderDataset,ResampleTransform,AudioPadCollate - Video (pre-extracted frames only; see Notes):
VideoFrameDirectoryDataset,UniformFrameSampleTransform - Validation:
FieldStatistics/DatasetStatistics/ValidationIssue,DatasetValidator::ComputeStatistics/DetectIssues. These compute descriptive statistics and flag missing values and outliers (see Notes). - MNIST:
MnistDatasetAdapterwraps the MNIST IDX loader (MnistIdxLoader,mnist_loader.hpp) as aDataset.
DataLoaderOptions fields: batch_size (default 1), shuffle and shuffle_seed,
drop_last, collate_fn (default DefaultCollate), num_workers and prefetch_batches.
DataLoader also accepts an IterableDataset; with one, num_workers must be 0 or 1.
Full API reference: Doxygen: Data Loading, Transformation & Validation
How to implement¶
Loading a CSV file through DataLoader¶
#include "pulsatrix/cpu_backend.hpp"
#include "pulsatrix/csv_dataset.hpp"
#include "pulsatrix/data_loader.hpp"
using namespace pulsatrix;
CPUBackend backend;
auto dataset = std::make_shared<CsvDataset>("data.csv", std::vector<std::string>{"x1", "x2"},
/*label_column=*/"y", &backend);
DataLoaderOptions options;
options.batch_size = 4;
options.shuffle = true;
DataLoader loader(dataset, &backend, options);
while (auto batch = loader.next_batch()) {
// batch->fields[0]: (N, 2) features, batch->fields[1]: (N,) labels
}
What's happening: CsvDataset::get(i) parses row i into a (1, num_features) feature
Tensor and a (1,) label Tensor. MnistDatasetAdapter uses the same one-row-per-sample
convention. DataLoader walks the dataset's indices with a Sampler, shuffled or sequential.
DefaultCollate then calls Tensor::Stack to stack each field along dimension 0 into a batch.
Loading is synchronous: num_workers defaults to 0, so no threads are spawned.
BoundedQueue and DataThreadPool exist and are unit-tested, but parallel fetching through
them is not wired in yet.
Recipe: CSV + DataLoader training.
Transforming images¶
#include "pulsatrix/cpu_backend.hpp"
#include "pulsatrix/data_loader.hpp"
#include "pulsatrix/image_folder_dataset.hpp"
#include "pulsatrix/image_transforms.hpp"
#include "pulsatrix/transform.hpp"
using namespace pulsatrix;
CPUBackend backend;
// root/<class_name>/<image_file>; labels are the sorted class-folder indices
auto images = std::make_shared<ImageFolderDataset>("data/images", &backend);
auto pipeline = std::make_shared<Compose>(std::vector<std::shared_ptr<Transform>>{
std::make_shared<ResizeTransform>(32, 32, &backend),
std::make_shared<NormalizeTransform>(std::vector<float>{0.5f, 0.5f, 0.5f}, // mean, one per channel
std::vector<float>{0.5f, 0.5f, 0.5f}), // std, one per channel
});
auto dataset = std::make_shared<TransformDataset>(images, pipeline);
DataLoaderOptions options;
options.batch_size = 16;
DataLoader loader(dataset, &backend, options);
// batch->fields[0]: (N, C, 32, 32) images, batch->fields[1]: (N,) labels
What's happening: TransformDataset applies the Composed transforms, in order, to each
sample as it is fetched. Images are decoded lazily, one per get() call. NormalizeTransform
needs one mean and one std per channel; it throws if the counts don't match the image.
Handling variable-length data with a custom CollateFn¶
DefaultCollate can only stack samples of the same shape. For ragged data such as text,
swap in a padding collate function:
#include "pulsatrix/cpu_backend.hpp"
#include "pulsatrix/data_loader.hpp"
#include "pulsatrix/text_collate.hpp"
#include "pulsatrix/text_dataset.hpp"
#include "pulsatrix/tokenizer.hpp"
#include "pulsatrix/vocabulary.hpp"
#include <fstream>
using namespace pulsatrix;
CPUBackend backend;
// Build the vocabulary from the same corpus, one token list per line.
std::vector<std::vector<std::string>> tokenized;
std::ifstream file("corpus.txt");
for (std::string line; std::getline(file, line);) {
tokenized.push_back(Tokenizer::Tokenize(line));
}
Vocabulary vocab = BuildVocabulary(tokenized);
auto dataset = std::make_shared<TextDataset>("corpus.txt", &vocab, &backend);
DataLoaderOptions options;
options.batch_size = 8;
options.collate_fn = PadCollate(); // pads each batch to its longest sequence
DataLoader loader(dataset, &backend, options);
while (auto batch = loader.next_batch()) {
// batch->fields[0]: (N, max_len) right-padded token indices
// batch->fields[1]: (N,) lengths before padding
}
What's happening: TextDataset::get(i) turns line i into a (1, seq_len) tensor of
token indices, and seq_len varies per line. PadCollate() pads every sample to the batch's
longest length and adds a lengths field. Dataset, DataLoader and Batch need no changes.
AudioPadCollate() applies the same pattern to variable-length audio.
The default pad value, 0, is also the index of the unknown token <unk>. Use the lengths
field, not the token value, to find padding.
Validating a dataset¶
#include "pulsatrix/cpu_backend.hpp"
#include "pulsatrix/csv_dataset.hpp"
#include "pulsatrix/dataset_validator.hpp"
using namespace pulsatrix;
CPUBackend backend;
CsvDataset dataset("data.csv", {"x1", "x2"}, /*label_column=*/"y", &backend); // any Dataset works
DatasetStatistics stats = DatasetValidator::ComputeStatistics(dataset);
std::vector<ValidationIssue> issues =
DatasetValidator::DetectIssues(dataset, stats, /*z_score_threshold=*/3.0f);
for (const ValidationIssue& issue : issues) {
// issue.sample_index, issue.field_index, issue.description
}
What's happening: ComputeStatistics scans every sample and computes per-field
statistics. It throws if samples have different numbers of fields. DetectIssues then flags
missing (NaN) values and elements more than z_score_threshold standard deviations from their
field's mean. Both work on any Dataset.
Notes¶
- Token IDs are stored as float32
Tensorvalues, not a separate integer type.Tensoris float32-only throughout this library, and float32 represents every integer up to 2^24 exactly, far above any vocabulary size.CsvDataset's label column andMnistDatasetAdapter's class label use the same convention. - Video support reads pre-extracted frames only.
VideoFrameDirectoryDatasetreads directories of frame images throughImageDecoder. It does not decode video containers or codecs, so there is no FFmpeg dependency. Clips sampled byUniformFrameSampleTransformare(N, C, H, W), the same shape convention as an image batch. - Validation is descriptive only.
DatasetValidatorcomputes per-field statistics and flags missing values and z-score outliers. Distribution drift detection and bias/fairness metrics are out of scope for now.
Recipes¶
See also examples/mnist_dataloader_demo.cpp
(CMake target mnist_dataloader_demo). It trains MnistConvNet on MNIST, loading the data
through DataLoader and MnistDatasetAdapter.