h_t = tanh(x_t @ W_xh + h_{t-1} @ W_hh + b_h), h_0 = 0 (zero-initialized, not learnable – a deliberate scope cut, see the class-level conservation note below). Input (N, L, input_size) -> output (N, L, hidden_size), the full hidden-state sequence. Single layer, tanh only, no bidirectional/multi-layer/ variable-length support.
More...
#include <rnn_module.hpp>
|
| | RNNModule (int64_t input_size, int64_t hidden_size, DeviceBackend *backend) |
| | Constructs an RNN layer with zero-initialized weights/bias.
|
| |
| Tensor | backward (const Tensor &grad_output) override |
| | Real backpropagation-through-time (BPTT): accumulates W_xh/W_hh/b_h gradients across every timestep into the same buffers via Tensor::accumulate().
|
| |
| OpType | op_type () const override |
| | Recurrent per charter's closed OpType set – a compound accumulate-over-time operation, not any existing category.
|
| |
| void | set_weight_xh (std::initializer_list< float > values) |
| | Overwrites the input-to-hidden weight buffer – test/initialization use only.
|
| |
| void | set_weight_hh (std::initializer_list< float > values) |
| | Overwrites the hidden-to-hidden weight buffer – test/initialization use only.
|
| |
| void | set_bias (std::initializer_list< float > values) |
| | Overwrites the hidden bias buffer – test/initialization use only.
|
| |
| const Tensor & | weight_xh () const |
| |
| const Tensor & | weight_hh () const |
| |
| const Tensor & | bias () const |
| |
| const Tensor & | weight_xh_grad () const |
| |
| const Tensor & | weight_hh_grad () const |
| |
| const Tensor & | bias_grad () const |
| |
| Tensor | propagate_relevance (const Tensor &relevance_out, const LRPRuleConfig &config) override |
| | Epsilon/z-rule LRP relevance propagation, generalized to two weighted sources, tanh treated as identity (Arras et al. 2017).
|
| |
| std::vector< NamedParamRef > | named_parameters () override |
| | This module's trainable parameters, each with its hierarchical name – the one place a module declares its parameters (roadmap FND-1).
|
| |
| std::optional< DeviceType > | compute_device () const override |
| | Where this layer computes, so forward() rejects an input on another device (FND-8).
|
| |
| virtual | ~Module ()=default |
| |
| Tensor | forward (const Tensor &input) |
| | Runs this module's forward computation.
|
| |
| std::pair< Tensor, NodeId > | forward_traced (const Tensor &input, NodeId input_node, ComputationGraph &graph, Autograd &autograd) |
| | Runs forward() while also registering a ComputationGraph node (tagged with this module's op_type(), parented to input_node) and wiring an Autograd backward function that reuses this module's own backward() – the opt-in traced/explainable path, per Phase 2 Mission 0.
|
| |
| virtual bool | supports_lrp_rule (LRPRule rule) const |
| | Whether propagate_relevance() implements rule (no silent fallback: callers such as ExplainerContext::relevance_pass() throw rather than run a module on a rule it does not implement).
|
| |
| virtual std::vector< ParamRef > | parameters () |
| | This module's trainable parameters and their gradients, for an optimizer to update uniformly across module types.
|
| |
| void | set_requires_grad (bool requires_grad, const std::string &prefix="") |
| | Freezes (false) or unfreezes (true) parameters by name (roadmap FND-2).
|
| |
| virtual void | set_training (bool training) |
| | Sets this module's training/eval mode. Defaults to training (matches every mainstream framework's Module default).
|
| |
| bool | is_training () const |
| | Whether this module is currently in training mode.
|
| |
|
| Tensor | forward_impl (const Tensor &input) override |
| | The actual forward computation – per-timestep tied-weight recurrence.
|
| |
h_t = tanh(x_t @ W_xh + h_{t-1} @ W_hh + b_h), h_0 = 0 (zero-initialized, not learnable – a deliberate scope cut, see the class-level conservation note below). Input (N, L, input_size) -> output (N, L, hidden_size), the full hidden-state sequence. Single layer, tanh only, no bidirectional/multi-layer/ variable-length support.
- Note
- Device-generic (GPU-native-kernels Mission 5): forward(), backward() and propagate_relevance() run entirely through DeviceBackend primitives (gemm/gemm_ex, copy_2d timestep slicing, elementwise Tanh, recurrent_cell(RnnBackward), accumulate_rows, lrp_linear), reproducing the former host loops' evaluation order so CPU results are bit-identical.
-
LRP rule (Arras et al. 2017, already cited in this charter for RNN/LSTM): epsilon/z-rule generalized to two weighted sources sharing one pre-activation (x_t's branch and h_{t-1}'s branch), tanh treated as identity pass-through (same precedent as ReluModule/DropoutModule, Montavon et al. 2019). Processed in reverse time order, threading a carried relevance accumulator exactly like backward()'s BPTT threads a gradient accumulator. Because h_0 is zero (this module's own scope cut), the relevance that would otherwise "leak" into the non-existent input before h_0 is provably exactly zero (the epsilon-rule numerator is h_prev*weight, and h_prev == 0 at t=0) – end-to-end conservation is exact, not approximate, verified numerically by a dedicated conservation test.
◆ RNNModule()
| pulsatrix::RNNModule::RNNModule |
( |
int64_t |
input_size, |
|
|
int64_t |
hidden_size, |
|
|
DeviceBackend * |
backend |
|
) |
| |
Constructs an RNN layer with zero-initialized weights/bias.
- Parameters
-
| input_size | Input feature dimension. |
| hidden_size | Hidden state dimension. |
| backend | Backend to allocate/compute through. Not owned; must outlive this module. |
- Exceptions
-
| std::invalid_argument | if input_size <= 0 or hidden_size <= 0 – external boundary (construction arguments can originate from Phase 5's Python bindings with no upstream validation). |
◆ backward()
| Tensor pulsatrix::RNNModule::backward |
( |
const Tensor & |
grad_output | ) |
|
|
overridevirtual |
Real backpropagation-through-time (BPTT): accumulates W_xh/W_hh/b_h gradients across every timestep into the same buffers via Tensor::accumulate().
- Parameters
-
| grad_output | Gradient w.r.t. this module's output. Must be (N, L, hidden_size) matching the most recent forward() call's output shape. |
- Returns
- Gradient w.r.t. this module's input, shape (N, L, input_size).
- Exceptions
-
| std::logic_error | if forward() has never been called. |
| std::invalid_argument | if grad_output's shape doesn't match the cached forward output shape. |
- Note
- Device-generic (GPU-native-kernels Mission 5): every step runs through DeviceBackend primitives, bit-identical to the former host loops on CPU.
Implements pulsatrix::Module.
◆ bias()
| const Tensor & pulsatrix::RNNModule::bias |
( |
| ) |
const |
|
inline |
◆ bias_grad()
| const Tensor & pulsatrix::RNNModule::bias_grad |
( |
| ) |
const |
|
inline |
◆ compute_device()
| std::optional< DeviceType > pulsatrix::RNNModule::compute_device |
( |
| ) |
const |
|
inlineoverridevirtual |
◆ forward_impl()
| Tensor pulsatrix::RNNModule::forward_impl |
( |
const Tensor & |
input | ) |
|
|
overrideprotectedvirtual |
The actual forward computation – per-timestep tied-weight recurrence.
- Exceptions
-
| std::invalid_argument | if input isn't rank-3 (N, L, input_size), or its last dimension doesn't match input_size. |
- Note
- Device-generic (GPU-native-kernels Mission 5): every step runs through DeviceBackend primitives, bit-identical to the former host loops on CPU.
Implements pulsatrix::Module.
◆ named_parameters()
| std::vector< NamedParamRef > pulsatrix::RNNModule::named_parameters |
( |
| ) |
|
|
inlineoverridevirtual |
This module's trainable parameters, each with its hierarchical name – the one place a module declares its parameters (roadmap FND-1).
- Returns
- {name, {value, grad}} entries pointing directly at this module's own members, in a fixed order. Names are unique within the module tree. Default: empty (a parameterless module like ReluModule needs no override).
- Note
- Override this, not parameters(): saving, loading, freezing by name and optimizer parameter groups all key on these names.
Reimplemented from pulsatrix::Module.
◆ op_type()
| OpType pulsatrix::RNNModule::op_type |
( |
| ) |
const |
|
inlineoverridevirtual |
Recurrent per charter's closed OpType set – a compound accumulate-over-time operation, not any existing category.
Implements pulsatrix::Module.
◆ propagate_relevance()
Epsilon/z-rule LRP relevance propagation, generalized to two weighted sources, tanh treated as identity (Arras et al. 2017).
- Parameters
-
| relevance_out | Relevance at this module's output. Must be (N, L, hidden_size) matching the most recent forward() call's output shape. |
| config | Selects epsilon. |
- Returns
- Relevance at this module's input, shape (N, L, input_size). Conserves exactly – see the class-level note.
- Exceptions
-
| std::logic_error | if forward() has never been called. |
| std::invalid_argument | if relevance_out's shape doesn't match the cached forward output shape. |
Implements pulsatrix::Module.
◆ set_bias()
| void pulsatrix::RNNModule::set_bias |
( |
std::initializer_list< float > |
values | ) |
|
Overwrites the hidden bias buffer – test/initialization use only.
◆ set_weight_hh()
| void pulsatrix::RNNModule::set_weight_hh |
( |
std::initializer_list< float > |
values | ) |
|
Overwrites the hidden-to-hidden weight buffer – test/initialization use only.
◆ set_weight_xh()
| void pulsatrix::RNNModule::set_weight_xh |
( |
std::initializer_list< float > |
values | ) |
|
Overwrites the input-to-hidden weight buffer – test/initialization use only.
◆ weight_hh()
| const Tensor & pulsatrix::RNNModule::weight_hh |
( |
| ) |
const |
|
inline |
◆ weight_hh_grad()
| const Tensor & pulsatrix::RNNModule::weight_hh_grad |
( |
| ) |
const |
|
inline |
◆ weight_xh()
| const Tensor & pulsatrix::RNNModule::weight_xh |
( |
| ) |
const |
|
inline |
◆ weight_xh_grad()
| const Tensor & pulsatrix::RNNModule::weight_xh_grad |
( |
| ) |
const |
|
inline |
The documentation for this class was generated from the following file: