Fixed-length, fill-once on-policy trajectory buffer of (observation, action, reward, log_prob, done) steps, with discounted return-to-go computation – the storage REINFORCE/A2C/PPO collect a rollout into.
More...
#include <rollout_buffer.hpp>
|
| | RolloutBuffer (int64_t max_length, int64_t observation_dim, int64_t action_dim, DeviceBackend *backend) |
| | Constructs an empty buffer with all storage pre-allocated and zero-filled.
|
| |
| void | add (const Tensor &observation, const Tensor &action, float reward, float log_prob, bool done) |
| | Appends one step to the rollout.
|
| |
| RolloutBatch | compute_returns (float gamma) const |
| | Reduces the stored rollout to per-step discounted return-to-go, G_t = sum_{k=t}^{T-1} gamma^(k-t) * r_k, alongside the stored observations/actions/log_probs.
|
| |
| Tensor | rewards () const |
| | The raw per-step rewards exactly as add()-ed, shape (size(), 1), in stored order.
|
| |
| Tensor | dones () const |
| | The raw per-step termination flags as 0.0f/1.0f floats, shape (size(), 1), in stored order – the same encoding ReplayBatch, ComputeDQNTarget and ComputeGAE use.
|
| |
| void | clear () |
| | Empties the rollout, making the buffer reusable for the next one.
|
| |
| int64_t | size () const |
| | Steps currently stored – rises to max_length(), never past it.
|
| |
| int64_t | max_length () const |
| | Steps one rollout holds, as passed to the constructor.
|
| |
| int64_t | observation_dim () const |
| | Width of an observation row, as passed to the constructor.
|
| |
| int64_t | action_dim () const |
| | Width of an action row, as passed to the constructor.
|
| |
Fixed-length, fill-once on-policy trajectory buffer of (observation, action, reward, log_prob, done) steps, with discounted return-to-go computation – the storage REINFORCE/A2C/PPO collect a rollout into.
The usage protocol is a cycle: add() exactly max_length() steps (or fewer, then stop collecting), compute_returns(gamma), do one gradient update, clear(), repeat.
- Note
- Deliberately NOT circular, unlike ReplayBuffer. An on-policy update consumes the whole rollout and then discards it – reusing stale data for a second update is the on-policy/off-policy distinction itself – so there is no rolling window to overwrite into. add() past max_length() therefore throws std::logic_error rather than wrapping around: the training loop forgot its compute-and-clear step, and silently dropping the oldest step of a trajectory would corrupt the return computation rather than merely losing data.
-
std::logic_error, not the std::invalid_argument every other throw here uses, and the distinction is deliberate: the arguments to an over-capacity add() are perfectly well-formed. What is wrong is the sequence of calls, a usage-protocol violation, not a malformed argument. The two are separately catchable on purpose.
-
Stores log_prob as a scalar per step, not the action distribution's parameters. That scalar is exactly what REINFORCE's policy-gradient estimator and PPO's probability ratio consume directly, and it is distribution-agnostic: a categorical, Gaussian or any other family's log-probability is still one float by the time it reaches this buffer, so the buffer never has to know which family produced it.
-
No
value slot. Value-function estimates are needed by A2C/PPO's advantage computation but not by plain REINFORCE; baking one into a Phase 1 generic buffer would be speculative, and extending or wrapping this buffer when a later phase actually needs one is additive. Same "core mechanism, not the full block" discipline RWKVModule applied to a layer, applied here to a data structure.
-
Not a Module subclass, and not an Environment/Agent either. No parameters, no gradient, no forward/backward and nothing for an LRP rule to explain – forcing it into Module would put a pure-virtual propagate_relevance() on a type for which the concept is undefined, exactly the failure mode module.hpp's charter note rules out. Same disposition as ReplayBuffer, MSELoss and Reparameterize: a plain utility class.
-
Storage is five pre-allocated (max_length, *)-shaped row-major host blocks written into at the current length, not a std::vector of per-step tensors – one allocation per field for the buffer's whole lifetime. clear() resets the length only; it never deallocates, so a training loop's repeated rollouts allocate nothing.
-
Host boundary (GPU-native-kernels Mission 7): the store is only ever read and written by the host, so it lives in host memory. add() accepts rows on any device (one device->host copy of each row per call); compute_returns(), rewards() and dones() hand back tensors uploaded through the buffer's own backend, so a buffer built with a GPU backend hands out GPU tensors.
◆ RolloutBuffer()
| pulsatrix::RolloutBuffer::RolloutBuffer |
( |
int64_t |
max_length, |
|
|
int64_t |
observation_dim, |
|
|
int64_t |
action_dim, |
|
|
DeviceBackend * |
backend |
|
) |
| |
Constructs an empty buffer with all storage pre-allocated and zero-filled.
- Parameters
-
| max_length | Number of steps one rollout holds. Must be >= 1. |
| observation_dim | Width of an observation row. Must be >= 1. |
| action_dim | Width of an action row (Environment::action_dim()). Must be >= 1. |
| backend | Backend to allocate through. Not owned; must outlive this buffer. |
- Exceptions
-
| std::invalid_argument | if max_length, observation_dim or action_dim is <= 0 – external boundary, the same classification as ReplayBuffer's capacity check. |
◆ action_dim()
| int64_t pulsatrix::RolloutBuffer::action_dim |
( |
| ) |
const |
|
inline |
Width of an action row, as passed to the constructor.
◆ add()
| void pulsatrix::RolloutBuffer::add |
( |
const Tensor & |
observation, |
|
|
const Tensor & |
action, |
|
|
float |
reward, |
|
|
float |
log_prob, |
|
|
bool |
done |
|
) |
| |
Appends one step to the rollout.
- Parameters
-
| observation | State the action was chosen from, shape (1, observation_dim()). |
| action | Action taken, shape (1, action_dim()). |
| reward | Scalar reward received. |
| log_prob | Log-probability the acting policy assigned to action. |
| done | Whether the episode ended on this step (segments the return computation). |
- Exceptions
-
| std::invalid_argument | if either tensor has the wrong shape – external boundary: a mismatched-shape Tensor can arrive from any caller, and silently writing it would corrupt neighbouring rows of the storage block. |
| std::logic_error | if size() already equals max_length(). See the class note: a full rollout is a protocol violation, not a malformed argument, and the two exception types are distinct so a caller can tell them apart. |
- Note
- Host boundary:
observation and action may live on any device; each is copied to the host once (Tensor::to_host_vector()).
◆ clear()
| void pulsatrix::RolloutBuffer::clear |
( |
| ) |
|
Empties the rollout, making the buffer reusable for the next one.
- Note
- Resets the length to 0 only. The storage blocks stay allocated at max_length() capacity and their stale contents are simply unreachable, since every read path is bounded by size(); zero-filling them would be work no observer can detect.
◆ compute_returns()
| RolloutBatch pulsatrix::RolloutBuffer::compute_returns |
( |
float |
gamma | ) |
const |
Reduces the stored rollout to per-step discounted return-to-go, G_t = sum_{k=t}^{T-1} gamma^(k-t) * r_k, alongside the stored observations/actions/log_probs.
- Parameters
-
| gamma | Discount factor, must be in (0, 1]. gamma == 1 is legal (undiscounted). |
- Returns
- The whole rollout as a RolloutBatch, leading dimension size().
- Exceptions
-
| std::invalid_argument | if gamma <= 0 or gamma > 1 – external boundary; those are not discount factors, and a negative or >1 gamma would produce sign-alternating or divergent returns rather than an error the caller notices. |
- Note
- Computed in a single reverse pass,
G_t = r_t + gamma * G_{t+1}, with the running accumulator reset to 0 at every step whose done flag is set, before that step's own reward is added. A rollout may span several episodes when episodes are shorter than max_length(), and without that reset episode k+1's return would silently leak backwards into episode k's final steps – a bootstrap across a terminal state, which is exactly the value the done flag exists to forbid. The returns are therefore per-episode, not rollout-cumulative.
-
const: this is a pure reduction over the stored steps. It does not consume or clear the buffer, so calling it twice yields the same batch; clear() is the caller's explicit, separate step once the update has been applied.
◆ dones()
| Tensor pulsatrix::RolloutBuffer::dones |
( |
| ) |
const |
The raw per-step termination flags as 0.0f/1.0f floats, shape (size(), 1), in stored order – the same encoding ReplayBatch, ComputeDQNTarget and ComputeGAE use.
- Returns
- A fresh (size(), 1) Tensor; the buffer's own storage is not exposed.
- Note
- Same rationale as rewards(): ComputeGAE() cuts both its bootstrap and its trace recursion on these flags, so a PPO training loop needs them per step rather than only as the episode segmentation compute_returns() already applied internally.
◆ max_length()
| int64_t pulsatrix::RolloutBuffer::max_length |
( |
| ) |
const |
|
inline |
Steps one rollout holds, as passed to the constructor.
◆ observation_dim()
| int64_t pulsatrix::RolloutBuffer::observation_dim |
( |
| ) |
const |
|
inline |
Width of an observation row, as passed to the constructor.
◆ rewards()
| Tensor pulsatrix::RolloutBuffer::rewards |
( |
| ) |
const |
The raw per-step rewards exactly as add()-ed, shape (size(), 1), in stored order.
- Returns
- A fresh (size(), 1) Tensor; the buffer's own storage is not exposed.
- Note
- Added for PPO (Phase 3 Mission 4), and deliberately not speculative API surface. The class note on RolloutBatch reasoned that once
returns has been computed the raw rewards/dones "carry no further information a REINFORCE/A2C/PPO update uses" – true of REINFORCE and A2C, and false of PPO: ComputeGAE() performs its own, differently-weighted reduction of the raw rewards/dones and never calls compute_returns(). So this is not "re-deriving a reduction this buffer already
performed"; it is the input to a different reduction the buffer does not perform.
-
Purely a read of already-stored state – no new logic, and add()/compute_returns()/ clear()/RolloutBatch are all untouched by its addition.
◆ size()
| int64_t pulsatrix::RolloutBuffer::size |
( |
| ) |
const |
|
inline |
Steps currently stored – rises to max_length(), never past it.
The documentation for this class was generated from the following file: