pulsatrix
Loading...
Searching...
No Matches
pulsatrix::RolloutBuffer Class Reference

Fixed-length, fill-once on-policy trajectory buffer of (observation, action, reward, log_prob, done) steps, with discounted return-to-go computation – the storage REINFORCE/A2C/PPO collect a rollout into. More...

#include <rollout_buffer.hpp>

Public Member Functions

 RolloutBuffer (int64_t max_length, int64_t observation_dim, int64_t action_dim, DeviceBackend *backend)
 Constructs an empty buffer with all storage pre-allocated and zero-filled.
 
void add (const Tensor &observation, const Tensor &action, float reward, float log_prob, bool done)
 Appends one step to the rollout.
 
RolloutBatch compute_returns (float gamma) const
 Reduces the stored rollout to per-step discounted return-to-go, G_t = sum_{k=t}^{T-1} gamma^(k-t) * r_k, alongside the stored observations/actions/log_probs.
 
Tensor rewards () const
 The raw per-step rewards exactly as add()-ed, shape (size(), 1), in stored order.
 
Tensor dones () const
 The raw per-step termination flags as 0.0f/1.0f floats, shape (size(), 1), in stored order – the same encoding ReplayBatch, ComputeDQNTarget and ComputeGAE use.
 
void clear ()
 Empties the rollout, making the buffer reusable for the next one.
 
int64_t size () const
 Steps currently stored – rises to max_length(), never past it.
 
int64_t max_length () const
 Steps one rollout holds, as passed to the constructor.
 
int64_t observation_dim () const
 Width of an observation row, as passed to the constructor.
 
int64_t action_dim () const
 Width of an action row, as passed to the constructor.
 

Detailed Description

Fixed-length, fill-once on-policy trajectory buffer of (observation, action, reward, log_prob, done) steps, with discounted return-to-go computation – the storage REINFORCE/A2C/PPO collect a rollout into.

The usage protocol is a cycle: add() exactly max_length() steps (or fewer, then stop collecting), compute_returns(gamma), do one gradient update, clear(), repeat.

Note
Deliberately NOT circular, unlike ReplayBuffer. An on-policy update consumes the whole rollout and then discards it – reusing stale data for a second update is the on-policy/off-policy distinction itself – so there is no rolling window to overwrite into. add() past max_length() therefore throws std::logic_error rather than wrapping around: the training loop forgot its compute-and-clear step, and silently dropping the oldest step of a trajectory would corrupt the return computation rather than merely losing data.
std::logic_error, not the std::invalid_argument every other throw here uses, and the distinction is deliberate: the arguments to an over-capacity add() are perfectly well-formed. What is wrong is the sequence of calls, a usage-protocol violation, not a malformed argument. The two are separately catchable on purpose.
Stores log_prob as a scalar per step, not the action distribution's parameters. That scalar is exactly what REINFORCE's policy-gradient estimator and PPO's probability ratio consume directly, and it is distribution-agnostic: a categorical, Gaussian or any other family's log-probability is still one float by the time it reaches this buffer, so the buffer never has to know which family produced it.
No value slot. Value-function estimates are needed by A2C/PPO's advantage computation but not by plain REINFORCE; baking one into a Phase 1 generic buffer would be speculative, and extending or wrapping this buffer when a later phase actually needs one is additive. Same "core mechanism, not the full block" discipline RWKVModule applied to a layer, applied here to a data structure.
Not a Module subclass, and not an Environment/Agent either. No parameters, no gradient, no forward/backward and nothing for an LRP rule to explain – forcing it into Module would put a pure-virtual propagate_relevance() on a type for which the concept is undefined, exactly the failure mode module.hpp's charter note rules out. Same disposition as ReplayBuffer, MSELoss and Reparameterize: a plain utility class.
Storage is five pre-allocated (max_length, *)-shaped row-major host blocks written into at the current length, not a std::vector of per-step tensors – one allocation per field for the buffer's whole lifetime. clear() resets the length only; it never deallocates, so a training loop's repeated rollouts allocate nothing.
Host boundary (GPU-native-kernels Mission 7): the store is only ever read and written by the host, so it lives in host memory. add() accepts rows on any device (one device->host copy of each row per call); compute_returns(), rewards() and dones() hand back tensors uploaded through the buffer's own backend, so a buffer built with a GPU backend hands out GPU tensors.

Constructor & Destructor Documentation

◆ RolloutBuffer()

pulsatrix::RolloutBuffer::RolloutBuffer ( int64_t  max_length,
int64_t  observation_dim,
int64_t  action_dim,
DeviceBackend *  backend 
)

Constructs an empty buffer with all storage pre-allocated and zero-filled.

Parameters
max_lengthNumber of steps one rollout holds. Must be >= 1.
observation_dimWidth of an observation row. Must be >= 1.
action_dimWidth of an action row (Environment::action_dim()). Must be >= 1.
backendBackend to allocate through. Not owned; must outlive this buffer.
Exceptions
std::invalid_argumentif max_length, observation_dim or action_dim is <= 0 – external boundary, the same classification as ReplayBuffer's capacity check.

Member Function Documentation

◆ action_dim()

int64_t pulsatrix::RolloutBuffer::action_dim ( ) const
inline

Width of an action row, as passed to the constructor.

◆ add()

void pulsatrix::RolloutBuffer::add ( const Tensor &  observation,
const Tensor &  action,
float  reward,
float  log_prob,
bool  done 
)

Appends one step to the rollout.

Parameters
observationState the action was chosen from, shape (1, observation_dim()).
actionAction taken, shape (1, action_dim()).
rewardScalar reward received.
log_probLog-probability the acting policy assigned to action.
doneWhether the episode ended on this step (segments the return computation).
Exceptions
std::invalid_argumentif either tensor has the wrong shape – external boundary: a mismatched-shape Tensor can arrive from any caller, and silently writing it would corrupt neighbouring rows of the storage block.
std::logic_errorif size() already equals max_length(). See the class note: a full rollout is a protocol violation, not a malformed argument, and the two exception types are distinct so a caller can tell them apart.
Note
Host boundary: observation and action may live on any device; each is copied to the host once (Tensor::to_host_vector()).

◆ clear()

void pulsatrix::RolloutBuffer::clear ( )

Empties the rollout, making the buffer reusable for the next one.

Note
Resets the length to 0 only. The storage blocks stay allocated at max_length() capacity and their stale contents are simply unreachable, since every read path is bounded by size(); zero-filling them would be work no observer can detect.

◆ compute_returns()

RolloutBatch pulsatrix::RolloutBuffer::compute_returns ( float  gamma) const

Reduces the stored rollout to per-step discounted return-to-go, G_t = sum_{k=t}^{T-1} gamma^(k-t) * r_k, alongside the stored observations/actions/log_probs.

Parameters
gammaDiscount factor, must be in (0, 1]. gamma == 1 is legal (undiscounted).
Returns
The whole rollout as a RolloutBatch, leading dimension size().
Exceptions
std::invalid_argumentif gamma <= 0 or gamma > 1 – external boundary; those are not discount factors, and a negative or >1 gamma would produce sign-alternating or divergent returns rather than an error the caller notices.
Note
Computed in a single reverse pass, G_t = r_t + gamma * G_{t+1}, with the running accumulator reset to 0 at every step whose done flag is set, before that step's own reward is added. A rollout may span several episodes when episodes are shorter than max_length(), and without that reset episode k+1's return would silently leak backwards into episode k's final steps – a bootstrap across a terminal state, which is exactly the value the done flag exists to forbid. The returns are therefore per-episode, not rollout-cumulative.
const: this is a pure reduction over the stored steps. It does not consume or clear the buffer, so calling it twice yields the same batch; clear() is the caller's explicit, separate step once the update has been applied.

◆ dones()

Tensor pulsatrix::RolloutBuffer::dones ( ) const

The raw per-step termination flags as 0.0f/1.0f floats, shape (size(), 1), in stored order – the same encoding ReplayBatch, ComputeDQNTarget and ComputeGAE use.

Returns
A fresh (size(), 1) Tensor; the buffer's own storage is not exposed.
Note
Same rationale as rewards(): ComputeGAE() cuts both its bootstrap and its trace recursion on these flags, so a PPO training loop needs them per step rather than only as the episode segmentation compute_returns() already applied internally.

◆ max_length()

int64_t pulsatrix::RolloutBuffer::max_length ( ) const
inline

Steps one rollout holds, as passed to the constructor.

◆ observation_dim()

int64_t pulsatrix::RolloutBuffer::observation_dim ( ) const
inline

Width of an observation row, as passed to the constructor.

◆ rewards()

Tensor pulsatrix::RolloutBuffer::rewards ( ) const

The raw per-step rewards exactly as add()-ed, shape (size(), 1), in stored order.

Returns
A fresh (size(), 1) Tensor; the buffer's own storage is not exposed.
Note
Added for PPO (Phase 3 Mission 4), and deliberately not speculative API surface. The class note on RolloutBatch reasoned that once returns has been computed the raw rewards/dones "carry no further information a REINFORCE/A2C/PPO update uses" – true of REINFORCE and A2C, and false of PPO: ComputeGAE() performs its own, differently-weighted reduction of the raw rewards/dones and never calls compute_returns(). So this is not "re-deriving a reduction this buffer already performed"; it is the input to a different reduction the buffer does not perform.
Purely a read of already-stored state – no new logic, and add()/compute_returns()/ clear()/RolloutBatch are all untouched by its addition.

◆ size()

int64_t pulsatrix::RolloutBuffer::size ( ) const
inline

Steps currently stored – rises to max_length(), never past it.


The documentation for this class was generated from the following file: