pulsatrix
Loading...
Searching...
No Matches
pulsatrix::RolloutBatch Struct Reference

One whole stored rollout, reduced to what a policy-gradient update consumes: the visited observations, the actions taken, the discounted return-to-go of each step, and the log-probability the acting policy assigned to each action. More...

#include <rollout_buffer.hpp>

Collaboration diagram for pulsatrix::RolloutBatch:

Public Attributes

Tensor observations
 
Tensor actions
 
Tensor returns
 
Tensor log_probs
 

Detailed Description

One whole stored rollout, reduced to what a policy-gradient update consumes: the visited observations, the actions taken, the discounted return-to-go of each step, and the log-probability the acting policy assigned to each action.

Note
Plain data, no behavior – the same small named return-struct precedent as ReparamGrad (reparameterize.hpp), StepResult (environment.hpp) and ReplayBatch (replay_buffer.hpp): four positional members in a std::tuple would be unreadable at every call site.
All four tensors share the same leading dimension, RolloutBuffer::size() at the time compute_returns() was called, and row t of every tensor belongs to the same stored step. observations is (size, observation_dim), actions is (size, action_dim), returns and log_probs are (size, 1).
No dones member, unlike ReplayBatch. The done flags exist only to segment the return computation at episode boundaries; once returns has been computed they carry no further information a REINFORCE/A2C/PPO update uses, and handing them back would invite a caller to re-derive a reduction this buffer has already performed.

Member Data Documentation

◆ actions

Tensor pulsatrix::RolloutBatch::actions

◆ log_probs

Tensor pulsatrix::RolloutBatch::log_probs

◆ observations

Tensor pulsatrix::RolloutBatch::observations

◆ returns

Tensor pulsatrix::RolloutBatch::returns

The documentation for this struct was generated from the following file: