|
pulsatrix
|
The (action, log_prob) pair TanhGaussianPolicy::forward() produces. More...
#include <tanh_gaussian_policy.hpp>

Public Attributes | |
| Tensor | action |
| Tensor | log_prob |
The (action, log_prob) pair TanhGaussianPolicy::forward() produces.
action is (N, action_dim) – the bounded action handed to the environment. log_prob is (N, 1): the per-sample log-density of the joint action, i.e. already summed over action_dim (a diagonal Gaussian's dimensions are independent, so the joint log-density is the sum of the per-dimension ones). For action_dim == 1 the sum is a one-term sum, not a special case. | Tensor pulsatrix::TanhGaussianSample::action |
| Tensor pulsatrix::TanhGaussianSample::log_prob |