pulsatrix
Loading...
Searching...
No Matches
pulsatrix::ContinuousCartPoleEnv Class Reference

The cart-pole balancing task with a continuous action: the action is a single force fraction in [-1, 1] rather than a discrete left/right choice. More...

#include <continuous_cartpole_env.hpp>

Inheritance diagram for pulsatrix::ContinuousCartPoleEnv:
Collaboration diagram for pulsatrix::ContinuousCartPoleEnv:

Public Member Functions

 ContinuousCartPoleEnv (DeviceBackend *backend, int64_t max_steps=200, uint32_t seed=42)
 Constructs a fresh, not-yet-reset continuous cart-pole environment.
 
Tensor reset () override
 Starts a new episode from a small random state: each of the 4 state variables drawn i.i.d. uniform in [-0.05, 0.05] from the internal LCG.
 
Tensor reset (const Tensor &initial_state) override
 Starts a new episode from an exact caller-supplied state, bypassing the LCG.
 
StepResult step (const Tensor &action) override
 Advances the physics one timestep (tau = 0.02 s) under the given continuous action.
 
int64_t observation_dim () const override
 4 – (x, x_dot, theta, theta_dot).
 
int64_t action_dim () const override
 1 – a single continuous force fraction.
 
bool is_discrete () const override
 False – this environment's action space is continuous.
 
int64_t step_count () const
 Steps taken since the last reset().
 
int64_t max_steps () const
 Episode length limit this environment was constructed with.
 
- Public Member Functions inherited from pulsatrix::Environment
virtual ~Environment ()=default
 

Static Public Attributes

static constexpr double kXThreshold = 2.4
 Cart position past which the episode terminates.
 
static constexpr double kThetaThreshold = 0.20943951
 Pole angle (radians) past which the episode terminates (~12 degrees).
 
static constexpr double kActionRangeTolerance = 1e-4
 How far outside [-1, 1] an action may sit before step() rejects it.
 

Detailed Description

The cart-pole balancing task with a continuous action: the action is a single force fraction in [-1, 1] rather than a discrete left/right choice.

The physics, constants, integration order, reward and termination thresholds are exactly CartPoleEnv's (Barto, Sutton & Anderson 1983 / OpenAI Gym's CartPoleEnv equations, cited as an independently-verifiable reference, not reused as a dependency). The only difference is how action becomes force:

So action = +1.0 is exactly CartPoleEnv's discrete action = 1, action = -1.0 is exactly its discrete action = 0, and every value between is a proportionally weaker push. That equivalence at the extremes is a test in this class's suite, tying its physics to the already-verified discrete environment rather than only to a re-derivation.

State is (x, x_dot, theta, theta_dot); the observation is that 4-vector, shape (1, 4). Every step yields reward 1.0, the terminating step included (Gym's convention). The episode ends when |x| > 2.4, |theta| > 0.20943951 rad (~12 degrees), or max_steps steps have been taken.

Note
Exists because Phase 4 (SAC) needs a continuous-control test case and CartPoleEnv is discrete-only. Deliberately a sibling class rather than a mode flag or a template parameter on CartPoleEnv: the discrete environment is already verified and depended on by four training-integration tests, and an action-interpretation branch inside its step() would make every one of those tests pay for a case it never takes. The duplicated physics block is ~10 lines and is pinned to the original by a direct trajectory-equality test, which is a stronger guarantee than shared code without one.
The physics constants are fixed, not constructor-configurable – the same deliberate scope cut CartPoleEnv documents. A knob with exactly one used value is speculative.
Integration is explicit (forward) Euler in Gym's exact order: x and theta are updated from the pre-update velocities. Semi-implicit Euler is a one-line difference that produces measurably different trajectories; do not "fix" this.
Physics is computed and stored in double; only the observation Tensor is float. Repeated float-precision Euler steps drift enough over a 200-step episode to make a hand-derived reference trajectory unreproducible.

Constructor & Destructor Documentation

◆ ContinuousCartPoleEnv()

pulsatrix::ContinuousCartPoleEnv::ContinuousCartPoleEnv ( DeviceBackend *  backend,
int64_t  max_steps = 200,
uint32_t  seed = 42 
)
explicit

Constructs a fresh, not-yet-reset continuous cart-pole environment.

Parameters
backendBackend to allocate observation tensors through. Not owned; must outlive this object.
max_stepsEpisode length limit. Must be >= 1. Defaults to 200, matching CartPoleEnv.
seedSeed for the internal deterministic LCG used by the no-argument reset().
Exceptions
std::invalid_argumentif max_steps < 1 – external boundary.

Member Function Documentation

◆ action_dim()

int64_t pulsatrix::ContinuousCartPoleEnv::action_dim ( ) const
inlineoverridevirtual

1 – a single continuous force fraction.

Implements pulsatrix::Environment.

◆ is_discrete()

bool pulsatrix::ContinuousCartPoleEnv::is_discrete ( ) const
inlineoverridevirtual

False – this environment's action space is continuous.

Implements pulsatrix::Environment.

◆ max_steps()

int64_t pulsatrix::ContinuousCartPoleEnv::max_steps ( ) const
inline

Episode length limit this environment was constructed with.

◆ observation_dim()

int64_t pulsatrix::ContinuousCartPoleEnv::observation_dim ( ) const
inlineoverridevirtual

4 – (x, x_dot, theta, theta_dot).

Implements pulsatrix::Environment.

◆ reset() [1/2]

Tensor pulsatrix::ContinuousCartPoleEnv::reset ( )
overridevirtual

Starts a new episode from a small random state: each of the 4 state variables drawn i.i.d. uniform in [-0.05, 0.05] from the internal LCG.

Returns
The initial observation, shape (1, 4).
Note
Same Numerical-Recipes LCG, same seeding and advance-across-resets behavior as CartPoleEnv, so a given seed produces an identical reset sequence in both classes – deterministic and therefore testable, unlike \<random\>'s implementation-defined engines.

Implements pulsatrix::Environment.

◆ reset() [2/2]

Tensor pulsatrix::ContinuousCartPoleEnv::reset ( const Tensor &  initial_state)
overridevirtual

Starts a new episode from an exact caller-supplied state, bypassing the LCG.

Parameters
initial_state(x, x_dot, theta, theta_dot), shape (1, 4).
Returns
The initial observation (a copy of initial_state's values), shape (1, 4).
Exceptions
std::invalid_argumentif initial_state's shape is not (1, 4) – external boundary.
Note
Does not check the state against the termination thresholds: pinning an already-terminal state and confirming the very next step() reports done is a legitimate (and tested) use.

Implements pulsatrix::Environment.

◆ step()

StepResult pulsatrix::ContinuousCartPoleEnv::step ( const Tensor &  action)
overridevirtual

Advances the physics one timestep (tau = 0.02 s) under the given continuous action.

Parameters
actionShape (1, 1), holding a force fraction in [-1 - kActionRangeTolerance, 1 + kActionRangeTolerance]. The value is clamped to exactly [-1, 1] before force = clamped * force_mag, so a tiny float-rounding excess cannot produce a force beyond +/-force_mag.
Returns
The resulting observation (1, 4), reward 1.0, and whether the episode ended.
Exceptions
std::invalid_argumentif neither reset() overload has been called yet, if the action's shape is not (1, 1), or if the action is outside the tolerated range. All external boundary: an action can originate from an untrusted policy output or, eventually, Python bindings.
Note
Host boundary (GPU-native-kernels Mission 7): the physics is a scalar double-precision update on the host. action may live on any device – one device->host copy of it per call – and the observation is returned through this environment's own backend.
NaN actions are rejected: every comparison against a NaN is false, so the range check is written as "reject unless inside the band" rather than "reject if outside it," which would let NaN through and poison the state permanently.
Stepping past a done=true result is allowed and keeps integrating; the episode loop is the caller's responsibility, exactly as in CartPoleEnv.

Implements pulsatrix::Environment.

◆ step_count()

int64_t pulsatrix::ContinuousCartPoleEnv::step_count ( ) const
inline

Steps taken since the last reset().

Member Data Documentation

◆ kActionRangeTolerance

constexpr double pulsatrix::ContinuousCartPoleEnv::kActionRangeTolerance = 1e-4
staticconstexpr

How far outside [-1, 1] an action may sit before step() rejects it.

A tanh-squashed policy (what Phase 4's SAC actor is) mathematically stays inside [-1, 1] up to float rounding, so a tolerance this small accommodates that round-trip without silently accepting a genuinely out-of-range action. It is the same 1e-4 action-tolerance convention CartPoleEnv established, applied to a continuous bound instead of an integer check.

Note
Compared in double, not float: 1.0f + 1e-4f rounds to the same float as 1.0001f (floats near 1.0 are spaced ~1.2e-7 apart, and the two values differ by ~1.6e-8), so a float-domain comparison would accept 1.0001f – a value that is unambiguously out of range. Widening to double makes the boundary mean what it says.

◆ kThetaThreshold

constexpr double pulsatrix::ContinuousCartPoleEnv::kThetaThreshold = 0.20943951
staticconstexpr

Pole angle (radians) past which the episode terminates (~12 degrees).

◆ kXThreshold

constexpr double pulsatrix::ContinuousCartPoleEnv::kXThreshold = 2.4
staticconstexpr

Cart position past which the episode terminates.


The documentation for this class was generated from the following file: