111 [[nodiscard]] int64_t
action_dim()
const override {
return 2; }
117 [[nodiscard]] int64_t
step_count()
const {
return step_count_; }
120 [[nodiscard]] int64_t
max_steps()
const {
return max_steps_; }
124 [[nodiscard]]
Tensor observation()
const;
127 [[nodiscard]]
double next_uniform();
136 double theta_dot_ = 0.0;
138 int64_t step_count_ = 0;
139 bool has_reset_ =
false;
The classic cart-pole balancing task (Barto, Sutton & Anderson 1983), the same equations OpenAI Gym's...
Definition cartpole_env.hpp:41
int64_t action_dim() const override
2 – push left (0) or push right (1).
Definition cartpole_env.hpp:111
int64_t step_count() const
Steps taken since the last reset().
Definition cartpole_env.hpp:117
bool is_discrete() const override
True – CartPole's action space is discrete.
Definition cartpole_env.hpp:114
Tensor reset() override
Starts a new episode from a small random state: each of the 4 state variables drawn i....
StepResult step(const Tensor &action) override
Advances the physics one timestep (tau = 0.02 s) under the given action.
Tensor reset(const Tensor &initial_state) override
Starts a new episode from an exact caller-supplied state, bypassing the LCG.
int64_t max_steps() const
Episode length limit this environment was constructed with.
Definition cartpole_env.hpp:120
static constexpr double kThetaThreshold
Pole angle (radians) past which the episode terminates (~12 degrees).
Definition cartpole_env.hpp:46
static constexpr double kXThreshold
Cart position past which the episode terminates.
Definition cartpole_env.hpp:44
int64_t observation_dim() const override
4 – (x, x_dot, theta, theta_dot).
Definition cartpole_env.hpp:108
CartPoleEnv(DeviceBackend *backend, int64_t max_steps=200, uint32_t seed=42)
Constructs a fresh, not-yet-reset CartPole environment.
Vendor-agnostic compute/memory backend. CPUBackend, CUDABackend (Phase 1.5), and HIPBackend (Phase 1....
Definition device_backend.hpp:219
Base class for every RL environment (CartPoleEnv, and whatever later phases add).
Definition environment.hpp:52
N-dimensional tensor. Owns its data buffer exclusively; a DeviceBackend* is injected (not owned) – th...
Definition tensor.hpp:29
Abstract interface isolating vendor-specific memory/compute operations from Tensor/ComputationGraph.
Abstract RL environment interface (gymnasium-shaped reset/step) + StepResult.
Definition acquisition_functions.hpp:16
What one Environment::step() produces: the next observation, this step's reward, and whether the epis...
Definition environment.hpp:25
N-dimensional tensor – owns a buffer via DeviceBackend*, RAII (Rule of Five).