y = (mask_i ? x_i / (1 - p) : 0) at training time (inverted dropout – scaling happens at training time so eval-time forward needs no rescaling); y = x at eval time or when p == 0. No parameters.
More...
|
| | DropoutModule (float p, DeviceBackend *backend, uint64_t seed) |
| | Constructs a dropout layer.
|
| |
| | DropoutModule (float p, DeviceBackend *backend) |
| | Seeded from the global seed stream (next_seed(), FND-7), so every layer built this way draws its own masks – reproducibly for a given set_seed().
|
| |
| Tensor | backward (const Tensor &grad_output) override |
| | Computes the gradient w.r.t. this module's input: grad_output * mask * scale, the true gradient of the actual (masked/scaled) forward computation. If the most recent forward() ran in eval mode, mask is all-ones and scale == 1, so this correctly reduces to identity with no special-casing needed here.
|
| |
| OpType | op_type () const override |
| | Elementwise per charter's closed OpType set – a per-element scale-or-zero operation, an honest fit for the existing category (no new OpType needed).
|
| |
| Tensor | propagate_relevance (const Tensor &relevance_out, const LRPRuleConfig &config) override |
| | Unconditional identity LRP relevance propagation.
|
| |
| bool | supports_lrp_rule (LRPRule) const override |
| | Pass-through relevance is the same under every rule: supports all of them.
|
| |
| std::optional< DeviceType > | compute_device () const override |
| | Where this layer computes, so forward() rejects an input on another device (FND-8).
|
| |
| virtual | ~Module ()=default |
| |
| Tensor | forward (const Tensor &input) |
| | Runs this module's forward computation.
|
| |
| std::pair< Tensor, NodeId > | forward_traced (const Tensor &input, NodeId input_node, ComputationGraph &graph, Autograd &autograd) |
| | Runs forward() while also registering a ComputationGraph node (tagged with this module's op_type(), parented to input_node) and wiring an Autograd backward function that reuses this module's own backward() – the opt-in traced/explainable path, per Phase 2 Mission 0.
|
| |
| virtual std::vector< NamedParamRef > | named_parameters () |
| | This module's trainable parameters, each with its hierarchical name – the one place a module declares its parameters (roadmap FND-1).
|
| |
| virtual std::vector< ParamRef > | parameters () |
| | This module's trainable parameters and their gradients, for an optimizer to update uniformly across module types.
|
| |
| void | set_requires_grad (bool requires_grad, const std::string &prefix="") |
| | Freezes (false) or unfreezes (true) parameters by name (roadmap FND-2).
|
| |
| virtual void | set_training (bool training) |
| | Sets this module's training/eval mode. Defaults to training (matches every mainstream framework's Module default).
|
| |
| bool | is_training () const |
| | Whether this module is currently in training mode.
|
| |
y = (mask_i ? x_i / (1 - p) : 0) at training time (inverted dropout – scaling happens at training time so eval-time forward needs no rescaling); y = x at eval time or when p == 0. No parameters.
- Note
- propagate_relevance is unconditional identity pass-through, reusing ReluModule's established precedent (Montavon et al. 2019: activation-/ regularization-like pointwise operations pass relevance through unchanged) – the same masked-backward/unconditional-identity-relevance split ReluModule already uses, not a new rule invented for this module. Dropout is conventionally disabled during inference/explanation in every mainstream framework, so at is_training() == false (the expected state when running LRP) forward is already identity, making this the mathematically exact treatment for that case, not just an approximation carried over from ReLU.