pulsatrix
Loading...
Searching...
No Matches
pulsatrix::Environment Class Referenceabstract

Base class for every RL environment (CartPoleEnv, and whatever later phases add). More...

#include <environment.hpp>

Inheritance diagram for pulsatrix::Environment:

Public Member Functions

virtual ~Environment ()=default
 
virtual Tensor reset ()=0
 Starts a new episode from an implementation-chosen initial state.
 
virtual Tensor reset (const Tensor &initial_state)=0
 Starts a new episode from an exact caller-supplied initial state.
 
virtual StepResult step (const Tensor &action)=0
 Advances the environment one timestep under the given action.
 
virtual int64_t observation_dim () const =0
 Number of components in an observation vector.
 
virtual int64_t action_dim () const =0
 The action space's size: the number of distinct actions when is_discrete(), otherwise the dimension of a continuous action vector.
 
virtual bool is_discrete () const =0
 Whether action_dim() counts discrete choices (true) or vector components.
 

Detailed Description

Base class for every RL environment (CartPoleEnv, and whatever later phases add).

Note
Deliberately NOT a Module subclass. A Module is a differentiable, parameterized layer with an LRP rule and a gradient; an environment has no parameters, no gradient, and no relevance to propagate. Forcing the two together would put a pure-virtual propagate_relevance() on a type for which the concept is undefined – exactly the "no default rule" failure mode module.hpp's charter note rules out. Environment/Agent are therefore a new, independent interface family.
Gymnasium-API-shaped on purpose (campaign Risk Register mitigation): reset/step, reward, done. There is no Python dependency here, but a future pybind11 binding should be able to wrap this in a gymnasium.Env with minimal adaptation rather than a translation layer.
Observations are always shape (1, observation_dim()) – this codebase's established always-batched convention (SinusoidalTimestepEmbedding's own (1, D) decision), so an observation feeds straight into any Module::forward() without a reshape.
Actions are always a Tensor, including for discrete action spaces, where the action is a shape-(1,1) Tensor holding the action index as a float. is_discrete() tells a caller which interpretation applies. This keeps the "everything is a Tensor" uniformity the rest of the codebase already has.

Constructor & Destructor Documentation

◆ ~Environment()

virtual pulsatrix::Environment::~Environment ( )
virtualdefault

Member Function Documentation

◆ action_dim()

virtual int64_t pulsatrix::Environment::action_dim ( ) const
pure virtual

The action space's size: the number of distinct actions when is_discrete(), otherwise the dimension of a continuous action vector.

Implemented in pulsatrix::CartPoleEnv, pulsatrix::ContinuousCartPoleEnv, and pulsatrix::HyperGridEnv.

◆ is_discrete()

virtual bool pulsatrix::Environment::is_discrete ( ) const
pure virtual

Whether action_dim() counts discrete choices (true) or vector components.

Implemented in pulsatrix::CartPoleEnv, pulsatrix::ContinuousCartPoleEnv, and pulsatrix::HyperGridEnv.

◆ observation_dim()

virtual int64_t pulsatrix::Environment::observation_dim ( ) const
pure virtual

Number of components in an observation vector.

Implemented in pulsatrix::CartPoleEnv, pulsatrix::ContinuousCartPoleEnv, and pulsatrix::HyperGridEnv.

◆ reset() [1/2]

virtual Tensor pulsatrix::Environment::reset ( )
pure virtual

Starts a new episode from an implementation-chosen initial state.

Returns
The initial observation, shape (1, observation_dim()).
Note
Any randomness is expected to come from an internal deterministic generator seeded at construction, not from a global/nondeterministic source – the same "randomness is a reproducible, caller-controlled input" convention as Reparameterize's caller-supplied epsilon and NoiseSchedule's epsilon/z.

Implemented in pulsatrix::CartPoleEnv, pulsatrix::ContinuousCartPoleEnv, and pulsatrix::HyperGridEnv.

◆ reset() [2/2]

virtual Tensor pulsatrix::Environment::reset ( const Tensor &  initial_state)
pure virtual

Starts a new episode from an exact caller-supplied initial state.

Parameters
initial_stateState to start from, shape (1, observation_dim()).
Returns
The initial observation (the state just set), shape (1, observation_dim()).
Exceptions
std::invalid_argumentif initial_state's shape is wrong – external boundary.
Note
Pure-virtual rather than a base-class default that validates and delegates to a protected set_state hook: the mission leaves the choice to the implementer, and a pure-virtual pair keeps Environment a true pure interface with no state and no protected extension point, which is the simpler contract to subclass against. Each environment implements both overloads directly.

Implemented in pulsatrix::CartPoleEnv, pulsatrix::ContinuousCartPoleEnv, and pulsatrix::HyperGridEnv.

◆ step()

virtual StepResult pulsatrix::Environment::step ( const Tensor &  action)
pure virtual

Advances the environment one timestep under the given action.

Parameters
actionThe action to take, shape (1, 1) for a discrete environment.
Returns
The resulting observation, reward and done flag.
Exceptions
std::invalid_argumentif the action is malformed or the environment has not been reset yet (implementation-defined which preconditions apply).

Implemented in pulsatrix::CartPoleEnv, pulsatrix::ContinuousCartPoleEnv, and pulsatrix::HyperGridEnv.


The documentation for this class was generated from the following file: