Adam (Kingma & Ba, 2015): per-parameter moving averages of gradient (m) and squared gradient (v), with bias correction.
More...
#include <adam_optimizer.hpp>
|
| | AdamOptimizer (float learning_rate, DeviceBackend *backend, float beta1=0.9f, float beta2=0.999f, float eps=1e-8f) |
| | Constructs an Adam optimizer.
|
| |
| void | step (Module &module) |
| | Applies one Adam update to every parameter the module exposes.
|
| |
| void | zero_grad (Module &module) |
| | Resets every parameter's gradient to zero. Does not reset Adam's moment state.
|
| |
| float | learning_rate () const |
| | Current step size.
|
| |
| void | set_learning_rate (float learning_rate) |
| | Overwrites the step size used by every subsequent step() call – necessary infrastructure for any mid-training hyperparameter schedule (e.g. Population Based Training's own explore step), found necessary by campaign_exai_dl_library_evolutionary_deep_learning's Phase 4 Mission 0, logged as a small addition beyond this class's original fixed-at-construction scope. Does not reset Adam's own moment state (m/v), matching zero_grad()'s own precedent that state and gradient are independent concerns.
|
| |
Adam (Kingma & Ba, 2015): per-parameter moving averages of gradient (m) and squared gradient (v), with bias correction.
- Note
- Per-parameter state is keyed by the parameter Tensor's pointer identity (
ParamRef::value), which is stable for as long as the owning Module exists – LinearModule/Conv2DModule's weight_/bias_ members never move or get reallocated.
◆ AdamOptimizer()
| pulsatrix::AdamOptimizer::AdamOptimizer |
( |
float |
learning_rate, |
|
|
DeviceBackend * |
backend, |
|
|
float |
beta1 = 0.9f, |
|
|
float |
beta2 = 0.999f, |
|
|
float |
eps = 1e-8f |
|
) |
| |
|
explicit |
Constructs an Adam optimizer.
- Parameters
-
| learning_rate | Step size. |
| backend | Backend used to allocate per-parameter moment-tracking tensors. |
| beta1 | First moment decay rate. |
| beta2 | Second moment decay rate. |
| eps | Denominator stabilizer. |
◆ learning_rate()
| float pulsatrix::AdamOptimizer::learning_rate |
( |
| ) |
const |
|
inline |
◆ set_learning_rate()
| void pulsatrix::AdamOptimizer::set_learning_rate |
( |
float |
learning_rate | ) |
|
|
inline |
Overwrites the step size used by every subsequent step() call – necessary infrastructure for any mid-training hyperparameter schedule (e.g. Population Based Training's own explore step), found necessary by campaign_exai_dl_library_evolutionary_deep_learning's Phase 4 Mission 0, logged as a small addition beyond this class's original fixed-at-construction scope. Does not reset Adam's own moment state (m/v), matching zero_grad()'s own precedent that state and gradient are independent concerns.
◆ step()
| void pulsatrix::AdamOptimizer::step |
( |
Module & |
module | ) |
|
Applies one Adam update to every parameter the module exposes.
- Parameters
-
| module | Module to update. Safe no-op if it has no parameters. |
- Note
- Device-generic: one fused DeviceBackend::adam_step per parameter. Moments are allocated through this optimizer's backend, so every parameter must live on that backend's device.
- Exceptions
-
| std::invalid_argument | if a parameter's device differs from the backend's. |
◆ zero_grad()
| void pulsatrix::AdamOptimizer::zero_grad |
( |
Module & |
module | ) |
|
Resets every parameter's gradient to zero. Does not reset Adam's moment state.
- Parameters
-
| module | Module whose gradients to reset. Safe no-op if it has no parameters. |
The documentation for this class was generated from the following file: