pulsatrix
Loading...
Searching...
No Matches
pulsatrix::Tokenizer Class Reference

Splits text into lowercase word/punctuation tokens – pulsatrix's first text primitive (campaign_exai_dl_library_data_pipeline, Phase 3, Decision Point 6: a minimal whitespace/punctuation tokenizer, not BPE – training a real subword-merge algorithm is a project-sized undertaking on its own). More...

#include <tokenizer.hpp>

Static Public Member Functions

static std::vector< std::string > Tokenize (const std::string &text)
 Tokenizes text: lowercase-normalizes, splits on whitespace (a pure separator, never emitted), emits runs of alphanumeric/apostrophe characters as single word tokens, and emits every other non-whitespace character as its own single-character punctuation token.
 

Detailed Description

Splits text into lowercase word/punctuation tokens – pulsatrix's first text primitive (campaign_exai_dl_library_data_pipeline, Phase 3, Decision Point 6: a minimal whitespace/punctuation tokenizer, not BPE – training a real subword-merge algorithm is a project-sized undertaking on its own).

Note
ASCII-oriented (uses std::tolower/std::isalnum/std::isspace per byte) – Unicode normalization is out of scope for this phase.

Member Function Documentation

◆ Tokenize()

static std::vector< std::string > pulsatrix::Tokenizer::Tokenize ( const std::string &  text)
static

Tokenizes text: lowercase-normalizes, splits on whitespace (a pure separator, never emitted), emits runs of alphanumeric/apostrophe characters as single word tokens, and emits every other non-whitespace character as its own single-character punctuation token.

Parameters
textInput text.
Returns
The token sequence, in order. Empty if text has no non-whitespace content.

The documentation for this class was generated from the following file: