Base class for every RL environment (CartPoleEnv, and whatever later phases add).
Definition environment.hpp:52
virtual int64_t observation_dim() const =0
Number of components in an observation vector.
virtual bool is_discrete() const =0
Whether action_dim() counts discrete choices (true) or vector components.
virtual int64_t action_dim() const =0
The action space's size: the number of distinct actions when is_discrete(), otherwise the dimension o...
virtual ~Environment()=default
virtual Tensor reset()=0
Starts a new episode from an implementation-chosen initial state.
virtual Tensor reset(const Tensor &initial_state)=0
Starts a new episode from an exact caller-supplied initial state.
virtual StepResult step(const Tensor &action)=0
Advances the environment one timestep under the given action.
N-dimensional tensor. Owns its data buffer exclusively; a DeviceBackend* is injected (not owned) – th...
Definition tensor.hpp:29
Definition acquisition_functions.hpp:16
What one Environment::step() produces: the next observation, this step's reward, and whether the epis...
Definition environment.hpp:25
Tensor observation
Definition environment.hpp:26
float reward
Definition environment.hpp:27
bool done
Definition environment.hpp:28
N-dimensional tensor – owns a buffer via DeviceBackend*, RAII (Rule of Five).