Recipe: CSV + DataLoader Training¶
What you'll build: a LinearModule(2,1) regressor trained to recover y = 2*x1 - 3*x2 + 1
from a synthetic CSV file. It pulls shuffled batches through CsvDataset + DataLoader instead
of building Tensors by hand (compare the XOR recipe's
hardcoded four-example dataset).
CMake target: csv_dataloader_training_recipe (examples/recipes/csv_dataloader_training.cpp).
Run it: ./build/csv_dataloader_training_recipe (Windows:
build\Release\csv_dataloader_training_recipe.exe).
Code¶
auto dataset = std::make_shared<CsvDataset>(csv_path, std::vector<std::string>{"x1", "x2"}, "y", &backend);
DataLoaderOptions options;
options.batch_size = 4;
options.shuffle = true;
options.shuffle_seed = 42;
LinearModule model(/*in_features=*/2, /*out_features=*/1, &backend);
AdamOptimizer optimizer(0.05f, &backend);
MSELoss loss_fn(&backend);
for (int epoch = 1; epoch <= 50; ++epoch) {
DataLoader loader(dataset, &backend, options); // fresh Sampler -> reshuffled order each epoch
while (auto batch = loader.next_batch()) {
optimizer.zero_grad(model);
Tensor prediction = model.forward(batch->fields[0]);
float loss_value = loss_fn.forward(prediction, batch->fields[1]);
Tensor grad_prediction = loss_fn.backward();
(void)model.backward(grad_prediction);
optimizer.step(model);
}
}
Full source: examples/recipes/csv_dataloader_training.cpp.
Expected output¶
CSV + DataLoader training recipe -- Linear(2,1) regressing y = 2*x1 - 3*x2 + 1
20 rows, batch_size=4, shuffled each epoch
epoch 1 | mean batch MSE 53.0817
epoch 10 | mean batch MSE 13.8806
epoch 20 | mean batch MSE 3.6161
epoch 30 | mean batch MSE 0.6891
epoch 40 | mean batch MSE 0.1259
epoch 50 | mean batch MSE 0.0374
What's happening¶
At startup the recipe writes a 20-row CSV file, csv_dataloader_training_recipe_data.csv, into
the current working directory. It deletes the file on exit, so there's nothing to download.
CsvDataset loads the file and turns each row into a (1, 2) feature Tensor and a (1,)
label Tensor. A new DataLoader is built each epoch. Its ShuffleSampler reshuffles on every
construction, so each epoch sees the batches in a different order.
LinearModule starts with all weights at zero. That's fine for a single linear layer, whose
gradient doesn't vanish at zero weights (unlike a multi-layer network with ReLU in between).
Training converges toward the true y = 2*x1 - 3*x2 + 1 using only the batches DataLoader
produces.
See also: Data Loading, Transformation & Validation.