One whole stored rollout, reduced to what a policy-gradient update consumes: the visited observations, the actions taken, the discounted return-to-go of each step, and the log-probability the acting policy assigned to each action.
More...
#include <rollout_buffer.hpp>
One whole stored rollout, reduced to what a policy-gradient update consumes: the visited observations, the actions taken, the discounted return-to-go of each step, and the log-probability the acting policy assigned to each action.
- Note
- Plain data, no behavior – the same small named return-struct precedent as ReparamGrad (reparameterize.hpp), StepResult (environment.hpp) and ReplayBatch (replay_buffer.hpp): four positional members in a std::tuple would be unreadable at every call site.
-
All four tensors share the same leading dimension, RolloutBuffer::size() at the time compute_returns() was called, and row t of every tensor belongs to the same stored step. observations is (size, observation_dim), actions is (size, action_dim), returns and log_probs are (size, 1).
-
No
dones member, unlike ReplayBatch. The done flags exist only to segment the return computation at episode boundaries; once returns has been computed they carry no further information a REINFORCE/A2C/PPO update uses, and handing them back would invite a caller to re-derive a reduction this buffer has already performed.
◆ actions
| Tensor pulsatrix::RolloutBatch::actions |
◆ log_probs
| Tensor pulsatrix::RolloutBatch::log_probs |
◆ observations
| Tensor pulsatrix::RolloutBatch::observations |
◆ returns
| Tensor pulsatrix::RolloutBatch::returns |
The documentation for this struct was generated from the following file: