Token<->index lookup table. Index 0 is always the reserved "<unk>" token – guaranteed by construction, not caller convention: the constructor takes the ranked list of real tokens and prepends "<unk>" itself.
More...
#include <vocabulary.hpp>
|
| | Vocabulary (std::vector< std::string > ranked_tokens) |
| |
| int64_t | IndexOf (const std::string &token) const |
| | The token's index, or kUnkIndex if the token isn't in this vocabulary.
|
| |
| const std::string & | TokenAt (int64_t index) const |
| | The token at a given index.
|
| |
| int64_t | size () const |
| | Total token count, including the reserved <unk> at index 0.
|
| |
Token<->index lookup table. Index 0 is always the reserved "<unk>" token – guaranteed by construction, not caller convention: the constructor takes the ranked list of real tokens and prepends "<unk>" itself.
◆ Vocabulary()
| pulsatrix::Vocabulary::Vocabulary |
( |
std::vector< std::string > |
ranked_tokens | ) |
|
|
explicit |
- Parameters
-
| ranked_tokens | Real (non-<unk>) tokens, in the order they should be indexed from 1. |
◆ IndexOf()
| int64_t pulsatrix::Vocabulary::IndexOf |
( |
const std::string & |
token | ) |
const |
The token's index, or kUnkIndex if the token isn't in this vocabulary.
◆ size()
| int64_t pulsatrix::Vocabulary::size |
( |
| ) |
const |
|
inline |
Total token count, including the reserved <unk> at index 0.
◆ TokenAt()
| const std::string & pulsatrix::Vocabulary::TokenAt |
( |
int64_t |
index | ) |
const |
The token at a given index.
- Note
- PULSATRIX_ASSERT-gated, not throw – internal invariant: every call site in this codebase passes an index already known to be in range (from IndexOf's own return value or a corpus's known vocabulary size), matching Shape::dim()'s classification.
◆ kUnkIndex
| constexpr int64_t pulsatrix::Vocabulary::kUnkIndex = 0 |
|
staticconstexpr |
◆ kUnkToken
| constexpr const char* pulsatrix::Vocabulary::kUnkToken = "<unk>" |
|
staticconstexpr |
The documentation for this class was generated from the following file: