|
pulsatrix
|
Splits text into lowercase word/punctuation tokens – pulsatrix's first text primitive (campaign_exai_dl_library_data_pipeline, Phase 3, Decision Point 6: a minimal whitespace/punctuation tokenizer, not BPE – training a real subword-merge algorithm is a project-sized undertaking on its own). More...
#include <tokenizer.hpp>
Static Public Member Functions | |
| static std::vector< std::string > | Tokenize (const std::string &text) |
| Tokenizes text: lowercase-normalizes, splits on whitespace (a pure separator, never emitted), emits runs of alphanumeric/apostrophe characters as single word tokens, and emits every other non-whitespace character as its own single-character punctuation token. | |
Splits text into lowercase word/punctuation tokens – pulsatrix's first text primitive (campaign_exai_dl_library_data_pipeline, Phase 3, Decision Point 6: a minimal whitespace/punctuation tokenizer, not BPE – training a real subword-merge algorithm is a project-sized undertaking on its own).
|
static |
Tokenizes text: lowercase-normalizes, splits on whitespace (a pure separator, never emitted), emits runs of alphanumeric/apostrophe characters as single word tokens, and emits every other non-whitespace character as its own single-character punctuation token.
| text | Input text. |