95 [[nodiscard]] int64_t
d_model()
const {
return d_model_; }
96 [[nodiscard]] int64_t
d_ff()
const {
return d_ff_; }
132 int64_t last_n_flat_ = 0;
136 bool has_forwarded_ =
false;
Vendor-agnostic compute/memory backend. CPUBackend, CUDABackend (Phase 1.5), and HIPBackend (Phase 1....
Definition device_backend.hpp:219
virtual DeviceType device() const noexcept=0
Which device this backend's buffers reside on.
y = x @ W + b, batched (x is (N, in_features), y is (N, out_features)) – migrated from the original u...
Definition linear_module.hpp:37
Base class for every layer type (LinearModule, Conv2DModule, activations, ...).
Definition module.hpp:58
An N-dimensional shape. A plain aggregate of dimensions with no invariant beyond "non-negative dimens...
Definition shape.hpp:24
down_proj(silu(gate_proj(x)) * up_proj(x)), the gated feedforward block used in place of a plain two-...
Definition swiglu_module.hpp:49
OpType op_type() const override
Elementwise per this module's own op_type() note above.
Definition swiglu_module.hpp:76
LinearModule & up_proj()
Definition swiglu_module.hpp:101
std::optional< DeviceType > compute_device() const override
Where this layer computes, so forward() rejects an input on another device (FND-8).
Definition swiglu_module.hpp:107
SwiGLUModule(int64_t d_model, int64_t d_ff, DeviceBackend *backend)
Constructs a SwiGLU block with zero-initialized projections.
int64_t d_ff() const
Definition swiglu_module.hpp:96
Tensor forward_impl(const Tensor &input) override
Runs: project (gate, up) -> silu(gate) -> gate*up -> project (down).
Tensor backward(const Tensor &grad_output) override
Gradient w.r.t. this module's input; sub-module parameter gradients accumulate inside gate_proj_/up_p...
Tensor propagate_relevance(const Tensor &relevance_out, const LRPRuleConfig &config) override
LRP relevance propagation: down_proj_'s epsilon rule, then the diagonal Eq. 15 split into gate/up sha...
int64_t d_model() const
Definition swiglu_module.hpp:95
std::vector< NamedParamRef > named_parameters() override
gate_proj_'s, up_proj_'s, and down_proj_'s parameters, flattened.
LinearModule & gate_proj()
Definition swiglu_module.hpp:100
LinearModule & down_proj()
Definition swiglu_module.hpp:102
N-dimensional tensor. Owns its data buffer exclusively; a DeviceBackend* is injected (not owned) – th...
Definition tensor.hpp:29
Dense/fully-connected layer – the reference Module implementation.
Abstract base every layer subclasses – NVI forward(), pure-virtual LRP contract.
Definition acquisition_functions.hpp:16
OpType
The op-type tag a Node carries. Charter Part 2 §3: nodes are tagged by a small closed set of op types...
Definition op_type.hpp:19
Configuration for LRP relevance propagation: which rule a module applies and its hyperparameters....
Definition lrp_rule_config.hpp:57