The classic cart-pole balancing task (Barto, Sutton & Anderson 1983), the same equations OpenAI Gym's own CartPoleEnv implements – cited as the canonical, independently-verifiable reference for this class's correctness tests, not reused as a code or runtime dependency.
More...
|
| | CartPoleEnv (DeviceBackend *backend, int64_t max_steps=200, uint32_t seed=42) |
| | Constructs a fresh, not-yet-reset CartPole environment.
|
| |
| Tensor | reset () override |
| | Starts a new episode from a small random state: each of the 4 state variables drawn i.i.d. uniform in [-0.05, 0.05] from the internal LCG.
|
| |
| Tensor | reset (const Tensor &initial_state) override |
| | Starts a new episode from an exact caller-supplied state, bypassing the LCG.
|
| |
| StepResult | step (const Tensor &action) override |
| | Advances the physics one timestep (tau = 0.02 s) under the given action.
|
| |
| int64_t | observation_dim () const override |
| | 4 – (x, x_dot, theta, theta_dot).
|
| |
| int64_t | action_dim () const override |
| | 2 – push left (0) or push right (1).
|
| |
| bool | is_discrete () const override |
| | True – CartPole's action space is discrete.
|
| |
| int64_t | step_count () const |
| | Steps taken since the last reset().
|
| |
| int64_t | max_steps () const |
| | Episode length limit this environment was constructed with.
|
| |
| virtual | ~Environment ()=default |
| |
The classic cart-pole balancing task (Barto, Sutton & Anderson 1983), the same equations OpenAI Gym's own CartPoleEnv implements – cited as the canonical, independently-verifiable reference for this class's correctness tests, not reused as a code or runtime dependency.
State is (x, x_dot, theta, theta_dot): cart position, cart velocity, pole angle from vertical (radians), pole angular velocity. The observation is exactly that 4-vector, shape (1, 4). Two discrete actions: 0 pushes the cart left, 1 pushes it right.
Every step yields reward 1.0, including the terminating step (Gym's own convention – the agent is rewarded for having survived through this step). The episode ends when |x| > 2.4, |theta| > 0.20943951 rad (~12 degrees), or max_steps steps have been taken.
- Note
- The physics constants (gravity, masses, pole length, force magnitude, integration timestep, termination thresholds) are fixed, not constructor-configurable – a deliberate scope cut. Parameterizing them is a future mission's job if a real need appears; a knob with exactly one used value is speculative generality.
-
Integration is explicit (forward) Euler in Gym's exact order: x and theta are updated from the pre-update velocities. Semi-implicit Euler is a one-line difference that produces measurably different trajectories; do not "fix" this.
-
Physics is computed in double and stored in double; only the observation Tensor is float. Repeated float-precision Euler steps drift enough over a 200-step episode to make a hand-derived reference trajectory unreproducible, which would undermine this class's whole reason for existing.