pulsatrix
Loading...
Searching...
No Matches
environment.hpp
Go to the documentation of this file.
1
5#pragma once
6
7#include <cstdint>
8
10
11namespace pulsatrix {
12
25struct StepResult {
27 float reward;
28 bool done;
29};
30
53public:
54 virtual ~Environment() = default;
55
64 [[nodiscard]] virtual Tensor reset() = 0;
65
77 [[nodiscard]] virtual Tensor reset(const Tensor& initial_state) = 0;
78
86 [[nodiscard]] virtual StepResult step(const Tensor& action) = 0;
87
89 [[nodiscard]] virtual int64_t observation_dim() const = 0;
90
95 [[nodiscard]] virtual int64_t action_dim() const = 0;
96
98 [[nodiscard]] virtual bool is_discrete() const = 0;
99};
100
101} // namespace pulsatrix
Base class for every RL environment (CartPoleEnv, and whatever later phases add).
Definition environment.hpp:52
virtual int64_t observation_dim() const =0
Number of components in an observation vector.
virtual bool is_discrete() const =0
Whether action_dim() counts discrete choices (true) or vector components.
virtual int64_t action_dim() const =0
The action space's size: the number of distinct actions when is_discrete(), otherwise the dimension o...
virtual ~Environment()=default
virtual Tensor reset()=0
Starts a new episode from an implementation-chosen initial state.
virtual Tensor reset(const Tensor &initial_state)=0
Starts a new episode from an exact caller-supplied initial state.
virtual StepResult step(const Tensor &action)=0
Advances the environment one timestep under the given action.
N-dimensional tensor. Owns its data buffer exclusively; a DeviceBackend* is injected (not owned) – th...
Definition tensor.hpp:29
Definition acquisition_functions.hpp:16
What one Environment::step() produces: the next observation, this step's reward, and whether the epis...
Definition environment.hpp:25
Tensor observation
Definition environment.hpp:26
float reward
Definition environment.hpp:27
bool done
Definition environment.hpp:28
N-dimensional tensor – owns a buffer via DeviceBackend*, RAII (Rule of Five).