|
pulsatrix
|
On-policy trajectory storage: fill-once fixed-length rollout + discounted returns. More...
#include <cstdint>#include <vector>#include "pulsatrix/device_backend.hpp"#include "pulsatrix/tensor.hpp"
Go to the source code of this file.
Classes | |
| struct | pulsatrix::RolloutBatch |
| One whole stored rollout, reduced to what a policy-gradient update consumes: the visited observations, the actions taken, the discounted return-to-go of each step, and the log-probability the acting policy assigned to each action. More... | |
| class | pulsatrix::RolloutBuffer |
| Fixed-length, fill-once on-policy trajectory buffer of (observation, action, reward, log_prob, done) steps, with discounted return-to-go computation – the storage REINFORCE/A2C/PPO collect a rollout into. More... | |
Namespaces | |
| namespace | pulsatrix |
On-policy trajectory storage: fill-once fixed-length rollout + discounted returns.