|
pulsatrix
|
Linear probe – trains one linear classifier to decode a binary concept from a layer's activations, answering "is this concept linearly represented here?". More...
#include <cmath>#include <cstdint>#include <random>#include <stdexcept>#include <string>#include <vector>#include "pulsatrix/determinism.hpp"#include "pulsatrix/bce_with_logits_loss.hpp"#include "pulsatrix/linear_module.hpp"
Go to the source code of this file.
Classes | |
| class | pulsatrix::LinearProbe |
A linear probe: LinearModule(activation_dim, 1) + BCEWithLogitsLoss, trained on (activation, binary concept label) pairs. High post-training accuracy means the concept is linearly decodable from those activations; chance-level accuracy means it is not (at least not linearly). More... | |
Namespaces | |
| namespace | pulsatrix |
Linear probe – trains one linear classifier to decode a binary concept from a layer's activations, answering "is this concept linearly represented here?".