y = x @ W + b, batched (x is (N, in_features), y is (N, out_features)) – migrated from the original unbatched (rank-1) scope by campaign_exai_dl_library_batch_dimension_support (breaking migration to always-batched; a single example is N=1, not a structurally different case).
More...
|
| | LinearModule (int64_t in_features, int64_t out_features, DeviceBackend *backend, DeviceType device) |
| | Constructs a linear layer with zero-initialized weight/bias.
|
| |
| | LinearModule (int64_t in_features, int64_t out_features, DeviceBackend *backend) |
| | As above, on backend's own device (backend->device()).
|
| |
| Tensor | backward (const Tensor &grad_output) override |
| | Computes the gradient w.r.t. this module's input, and accumulates the weight/bias gradients internally (summed across the batch).
|
| |
| OpType | op_type () const override |
| | Linear per charter's closed OpType set.
|
| |
| void | set_weight (std::initializer_list< float > values) |
| | Overwrites the weight buffer – test/initialization use only.
|
| |
| void | set_bias (std::initializer_list< float > values) |
| | Overwrites the bias buffer – test/initialization use only.
|
| |
| void | set_weight (const std::vector< float > &values) |
| | Vector overload for runtime-sized sources – see Tensor's own vector ctor.
|
| |
| void | set_bias (const std::vector< float > &values) |
| | Vector overload for runtime-sized sources – see Tensor's own vector ctor.
|
| |
| const Tensor & | weight () const |
| |
| const Tensor & | bias () const |
| |
| const Tensor & | weight_grad () const |
| |
| const Tensor & | bias_grad () const |
| |
| Tensor | propagate_relevance (const Tensor &relevance_out, const LRPRuleConfig &config) override |
| | Epsilon-rule LRP relevance propagation (Bach et al. 2015), applied independently per example in the batch.
|
| |
| bool | supports_lrp_rule (LRPRule) const override |
| | Implements every LRPRule.
|
| |
| std::vector< NamedParamRef > | named_parameters () override |
| | This module's trainable parameters, each with its hierarchical name – the one place a module declares its parameters (roadmap FND-1).
|
| |
| std::optional< DeviceType > | compute_device () const override |
| | Where this layer computes, so forward() rejects an input on another device (FND-8).
|
| |
| virtual | ~Module ()=default |
| |
| Tensor | forward (const Tensor &input) |
| | Runs this module's forward computation.
|
| |
| std::pair< Tensor, NodeId > | forward_traced (const Tensor &input, NodeId input_node, ComputationGraph &graph, Autograd &autograd) |
| | Runs forward() while also registering a ComputationGraph node (tagged with this module's op_type(), parented to input_node) and wiring an Autograd backward function that reuses this module's own backward() – the opt-in traced/explainable path, per Phase 2 Mission 0.
|
| |
| virtual std::vector< ParamRef > | parameters () |
| | This module's trainable parameters and their gradients, for an optimizer to update uniformly across module types.
|
| |
| void | set_requires_grad (bool requires_grad, const std::string &prefix="") |
| | Freezes (false) or unfreezes (true) parameters by name (roadmap FND-2).
|
| |
| virtual void | set_training (bool training) |
| | Sets this module's training/eval mode. Defaults to training (matches every mainstream framework's Module default).
|
| |
| bool | is_training () const |
| | Whether this module is currently in training mode.
|
| |
y = x @ W + b, batched (x is (N, in_features), y is (N, out_features)) – migrated from the original unbatched (rank-1) scope by campaign_exai_dl_library_batch_dimension_support (breaking migration to always-batched; a single example is N=1, not a structurally different case).
- Note
- Weight layout is (in_features, out_features), not the more common (out_features, in_features) PyTorch convention – chosen specifically so forward (x @ W) and the weight-gradient step (X^T @ grad_Y) both use DeviceBackend::gemm directly with no transpose of the weight operand. The input-gradient step (grad_Y @ W^T) and the weight-gradient step (X^T @ grad_Y, the batched sum of outer products) read their transposed operand in place through DeviceBackend::gemm_ex – no transposed copy is built.
-
Weights/biases are owned here as member Tensors, not ComputationGraph nodes. Parameter gradients accumulate via Tensor::accumulate() across backward() calls until something (the optimizer) resets them. Batch-dimension gradient reduction (summing weight/bias gradient contributions across the N examples in a batch) needs no new Tensor primitive – see campaign_exai_dl_library_batch_dimension_support's mission_tensor_shape_foundation.md: weight_grad's reduction falls out of gemm's own k-dimension summation (accumulated in place, beta = 1); bias_grad's is DeviceBackend::column_sums, likewise accumulated in place. Both are device-resident (GPU-native-kernels Mission 1).