Base class for every RL environment (CartPoleEnv, and whatever later phases add).
More...
#include <environment.hpp>
|
| virtual | ~Environment ()=default |
| |
| virtual Tensor | reset ()=0 |
| | Starts a new episode from an implementation-chosen initial state.
|
| |
| virtual Tensor | reset (const Tensor &initial_state)=0 |
| | Starts a new episode from an exact caller-supplied initial state.
|
| |
| virtual StepResult | step (const Tensor &action)=0 |
| | Advances the environment one timestep under the given action.
|
| |
| virtual int64_t | observation_dim () const =0 |
| | Number of components in an observation vector.
|
| |
| virtual int64_t | action_dim () const =0 |
| | The action space's size: the number of distinct actions when is_discrete(), otherwise the dimension of a continuous action vector.
|
| |
| virtual bool | is_discrete () const =0 |
| | Whether action_dim() counts discrete choices (true) or vector components.
|
| |
Base class for every RL environment (CartPoleEnv, and whatever later phases add).
- Note
- Deliberately NOT a Module subclass. A Module is a differentiable, parameterized layer with an LRP rule and a gradient; an environment has no parameters, no gradient, and no relevance to propagate. Forcing the two together would put a pure-virtual propagate_relevance() on a type for which the concept is undefined – exactly the "no default rule" failure mode module.hpp's charter note rules out. Environment/Agent are therefore a new, independent interface family.
-
Gymnasium-API-shaped on purpose (campaign Risk Register mitigation): reset/step, reward, done. There is no Python dependency here, but a future pybind11 binding should be able to wrap this in a
gymnasium.Env with minimal adaptation rather than a translation layer.
-
Observations are always shape (1, observation_dim()) – this codebase's established always-batched convention (SinusoidalTimestepEmbedding's own (1, D) decision), so an observation feeds straight into any Module::forward() without a reshape.
-
Actions are always a Tensor, including for discrete action spaces, where the action is a shape-(1,1) Tensor holding the action index as a float. is_discrete() tells a caller which interpretation applies. This keeps the "everything is a Tensor" uniformity the rest of the codebase already has.
◆ ~Environment()
| virtual pulsatrix::Environment::~Environment |
( |
| ) |
|
|
virtualdefault |
◆ action_dim()
| virtual int64_t pulsatrix::Environment::action_dim |
( |
| ) |
const |
|
pure virtual |
◆ is_discrete()
| virtual bool pulsatrix::Environment::is_discrete |
( |
| ) |
const |
|
pure virtual |
◆ observation_dim()
| virtual int64_t pulsatrix::Environment::observation_dim |
( |
| ) |
const |
|
pure virtual |
◆ reset() [1/2]
| virtual Tensor pulsatrix::Environment::reset |
( |
| ) |
|
|
pure virtual |
◆ reset() [2/2]
| virtual Tensor pulsatrix::Environment::reset |
( |
const Tensor & |
initial_state | ) |
|
|
pure virtual |
Starts a new episode from an exact caller-supplied initial state.
- Parameters
-
- Returns
- The initial observation (the state just set), shape (1, observation_dim()).
- Exceptions
-
| std::invalid_argument | if initial_state's shape is wrong – external boundary. |
- Note
- Pure-virtual rather than a base-class default that validates and delegates to a protected set_state hook: the mission leaves the choice to the implementer, and a pure-virtual pair keeps Environment a true pure interface with no state and no protected extension point, which is the simpler contract to subclass against. Each environment implements both overloads directly.
Implemented in pulsatrix::CartPoleEnv, pulsatrix::ContinuousCartPoleEnv, and pulsatrix::HyperGridEnv.
◆ step()
Advances the environment one timestep under the given action.
- Parameters
-
| action | The action to take, shape (1, 1) for a discrete environment. |
- Returns
- The resulting observation, reward and done flag.
- Exceptions
-
| std::invalid_argument | if the action is malformed or the environment has not been reset yet (implementation-defined which preconditions apply). |
Implemented in pulsatrix::CartPoleEnv, pulsatrix::ContinuousCartPoleEnv, and pulsatrix::HyperGridEnv.
The documentation for this class was generated from the following file: