Samples from softmax(mask(policy_network(observation))) – a categorical policy restricted to a caller...
Definition gflownet_forward_policy.hpp:47
An n-dimensional grid: state is an integer coordinate in [0, H-1]^ndim, actions increment one coordin...
Definition hypergrid_env.hpp:31
Masked stochastic categorical policy over a logit-producing Module – P_F for GFlowNet training object...
HyperGrid – the standard minimal GFlowNet correctness-check environment.
Definition acquisition_functions.hpp:16
GFlowNetTrajectory sample_gflownet_trajectory(HyperGridEnv &env, GFlowNetForwardPolicy &forward_policy)
Samples one full trajectory: resets env, then repeatedly samples a masked action from forward_policy ...
One full sampled trajectory: every state the forward policy acted from, the action taken at each,...
Definition gflownet_trajectory.hpp:26
std::vector< int64_t > actions
The action index sampled at each corresponding entry of states.
Definition gflownet_trajectory.hpp:30
float sum_log_pf
Σ_t log P_F(s_{t+1}|s_t) over the whole trajectory, including the final stop decision.
Definition gflownet_trajectory.hpp:33
float terminal_reward
R(x) at the trajectory's terminal state.
Definition gflownet_trajectory.hpp:38
float sum_log_pb
Σ_t log P_B(s_t|s_{t+1}) over every real move (the stop action does not move the state,...
Definition gflownet_trajectory.hpp:36
std::vector< Tensor > states
The state the policy acted from at each decision point, in order.
Definition gflownet_trajectory.hpp:28
N-dimensional tensor – owns a buffer via DeviceBackend*, RAII (Rule of Five).