One uniformly-sampled minibatch of transitions, one Tensor per transition field.
More...
#include <replay_buffer.hpp>
One uniformly-sampled minibatch of transitions, one Tensor per transition field.
- Note
- Plain data, no behavior – the same small named return-struct precedent as ReparamGrad (reparameterize.hpp) and StepResult (environment.hpp): five positional members in a std::tuple would be unreadable at every call site, and five out- parameters would be worse.
-
All five tensors share the same leading dimension, the batch_size passed to ReplayBuffer::sample(), and row b of every tensor belongs to the same transition. observations/next_observations are (batch_size, observation_dim), actions is (batch_size, action_dim), rewards and dones are (batch_size, 1).
-
dones holds 0.0f/1.0f floats rather than bools – this codebase's established "everything is a Tensor" convention, the same one that makes Environment encode a discrete action index as a float. A DQN target computation wants (1 - done) as a float multiplier anyway, so no conversion is imposed on the caller.
◆ actions
| Tensor pulsatrix::ReplayBatch::actions |
◆ dones
| Tensor pulsatrix::ReplayBatch::dones |
◆ next_observations
| Tensor pulsatrix::ReplayBatch::next_observations |
◆ observations
| Tensor pulsatrix::ReplayBatch::observations |
◆ rewards
| Tensor pulsatrix::ReplayBatch::rewards |
The documentation for this struct was generated from the following file: