Dataset over a line-delimited text corpus: each line becomes one sample, tokenized via Tokenizer::Tokenize and indexed via a Vocabulary into a (1, seq_len) float32 Tensor of token indices – Decision Point 6's resolved representation (token IDs as float32 values, the same integer-as-float32 pattern CsvDataset's label column and MnistDatasetAdapter's class label already use). seq_len varies per sample; no padding here (see PadCollate, Mission 10) – an empty line produces a zero-element (1, 0) Tensor, a valid, non-error Tensor state.
More...
#include <text_dataset.hpp>
Dataset over a line-delimited text corpus: each line becomes one sample, tokenized via Tokenizer::Tokenize and indexed via a Vocabulary into a (1, seq_len) float32 Tensor of token indices – Decision Point 6's resolved representation (token IDs as float32 values, the same integer-as-float32 pattern CsvDataset's label column and MnistDatasetAdapter's class label already use). seq_len varies per sample; no padding here (see PadCollate, Mission 10) – an empty line produces a zero-element (1, 0) Tensor, a valid, non-error Tensor state.
◆ TextDataset()
| pulsatrix::TextDataset::TextDataset |
( |
const std::string & |
corpus_path, |
|
|
const Vocabulary * |
vocabulary, |
|
|
DeviceBackend * |
backend |
|
) |
| |
- Parameters
-
| corpus_path | Path to a line-delimited text file. |
| vocabulary | Vocabulary to index tokens through. Not owned; must outlive this Dataset. |
| backend | Backend to allocate token-index Tensors through. Not owned. |
- Exceptions
-
| std::runtime_error | if corpus_path can't be opened – external boundary (file content, not an internal invariant). |
◆ get()
| Sample pulsatrix::TextDataset::get |
( |
int64_t |
index | ) |
const |
|
overridevirtual |
◆ size()
| int64_t pulsatrix::TextDataset::size |
( |
| ) |
const |
|
overridevirtual |
The documentation for this class was generated from the following file: