|
pulsatrix
|
Single-layer GRU – gated recurrence with a reset-gated candidate. More...

Go to the source code of this file.
Classes | |
| class | pulsatrix::GRUModule |
Standard GRU recurrence (Cho et al. 2014), h_0 = 0 (zero-initialized, not learnable – the same deliberate scope cut RNNModule/LSTMModule made, and the same thing that makes this module's conservation exact; see the LRP note below): z_t = sigmoid(x_t @ W_xz + h_{t-1} @ W_hz + b_z) (update gate), r_t = sigmoid(x_t @ W_xr + h_{t-1} @ W_hr + b_r) (reset gate), hn_prev_t = h_{t-1} @ W_hn (internal projection, no bias), n_t = tanh(x_t @ W_xn + r_t * hn_prev_t + b_n) (candidate), h_t = (1 - z_t) * h_{t-1} + z_t * n_t. Note the reset gate multiplies the projected previous hidden state hn_prev_t, not h_{t-1} itself – that projection is a distinct cached intermediate, and it is what gives GRU's LRP rule a different shape from LSTM's. Input (N, L, input_size) -> output (N, L, hidden_size), the full hidden-state sequence (matches RNNModule's/LSTMModule's convention). Single layer, no bidirectional/multi-layer/variable-length support. More... | |
Namespaces | |
| namespace | pulsatrix |
Single-layer GRU – gated recurrence with a reset-gated candidate.