|
| | ContinuousCartPoleEnv (DeviceBackend *backend, int64_t max_steps=200, uint32_t seed=42) |
| | Constructs a fresh, not-yet-reset continuous cart-pole environment.
|
| |
| Tensor | reset () override |
| | Starts a new episode from a small random state: each of the 4 state variables drawn i.i.d. uniform in [-0.05, 0.05] from the internal LCG.
|
| |
| Tensor | reset (const Tensor &initial_state) override |
| | Starts a new episode from an exact caller-supplied state, bypassing the LCG.
|
| |
| StepResult | step (const Tensor &action) override |
| | Advances the physics one timestep (tau = 0.02 s) under the given continuous action.
|
| |
| int64_t | observation_dim () const override |
| | 4 – (x, x_dot, theta, theta_dot).
|
| |
| int64_t | action_dim () const override |
| | 1 – a single continuous force fraction.
|
| |
| bool | is_discrete () const override |
| | False – this environment's action space is continuous.
|
| |
| int64_t | step_count () const |
| | Steps taken since the last reset().
|
| |
| int64_t | max_steps () const |
| | Episode length limit this environment was constructed with.
|
| |
| virtual | ~Environment ()=default |
| |
The cart-pole balancing task with a continuous action: the action is a single force fraction in [-1, 1] rather than a discrete left/right choice.
The physics, constants, integration order, reward and termination thresholds are exactly CartPoleEnv's (Barto, Sutton & Anderson 1983 / OpenAI Gym's CartPoleEnv equations, cited as an independently-verifiable reference, not reused as a dependency). The only difference is how action becomes force:
So action = +1.0 is exactly CartPoleEnv's discrete action = 1, action = -1.0 is exactly its discrete action = 0, and every value between is a proportionally weaker push. That equivalence at the extremes is a test in this class's suite, tying its physics to the already-verified discrete environment rather than only to a re-derivation.
State is (x, x_dot, theta, theta_dot); the observation is that 4-vector, shape (1, 4). Every step yields reward 1.0, the terminating step included (Gym's convention). The episode ends when |x| > 2.4, |theta| > 0.20943951 rad (~12 degrees), or max_steps steps have been taken.
- Note
- Exists because Phase 4 (SAC) needs a continuous-control test case and CartPoleEnv is discrete-only. Deliberately a sibling class rather than a mode flag or a template parameter on CartPoleEnv: the discrete environment is already verified and depended on by four training-integration tests, and an action-interpretation branch inside its step() would make every one of those tests pay for a case it never takes. The duplicated physics block is ~10 lines and is pinned to the original by a direct trajectory-equality test, which is a stronger guarantee than shared code without one.
-
The physics constants are fixed, not constructor-configurable – the same deliberate scope cut CartPoleEnv documents. A knob with exactly one used value is speculative.
-
Integration is explicit (forward) Euler in Gym's exact order: x and theta are updated from the pre-update velocities. Semi-implicit Euler is a one-line difference that produces measurably different trajectories; do not "fix" this.
-
Physics is computed and stored in double; only the observation Tensor is float. Repeated float-precision Euler steps drift enough over a 200-step episode to make a hand-derived reference trajectory unreproducible.