Generic data validation over any Dataset implementation – descriptive statistics and missingness/outlier detection only (campaign_exai_dl_library_data_pipeline Decision Point 5: distributional drift detection and bias/fairness metrics are explicitly out of scope, named follow-ups for a future campaign).
More...
#include <dataset_validator.hpp>
Generic data validation over any Dataset implementation – descriptive statistics and missingness/outlier detection only (campaign_exai_dl_library_data_pipeline Decision Point 5: distributional drift detection and bias/fairness metrics are explicitly out of scope, named follow-ups for a future campaign).
- Note
- Operates purely against the Dataset interface – no modality-specific code. Field statistics aggregate every element of every sample's Tensor for that field position (flattened, not per-position), since field shapes may legitimately vary in size across samples (e.g. text sequence length).
◆ ComputeStatistics()
Computes per-field descriptive statistics across every sample in dataset.
- Exceptions
-
| std::invalid_argument | if dataset is empty. |
| std::runtime_error | if samples have inconsistent field counts (a schema violation, detected while scanning) – external boundary: dataset content is caller-supplied, not an internal invariant. |
◆ DetectIssues()
Scans dataset for missing (NaN) values and statistical outliers.
- Parameters
-
| stats | Per-field baseline statistics (from ComputeStatistics) used for z-score outlier detection. |
| z_score_threshold | Elements with |value - mean| / std_dev exceeding this threshold are flagged as outliers. A field with std_dev == 0 (constant data) is skipped for outlier detection (z-score is undefined), not flagged. |
- Returns
- Every detected issue, in sample-then-field-then-element scan order.
The documentation for this class was generated from the following file: