pulsatrix
Loading...
Searching...
No Matches
continuous_cartpole_env.hpp
Go to the documentation of this file.
1
5#pragma once
6
7#include <cstdint>
8
11#include "pulsatrix/tensor.hpp"
12
13namespace pulsatrix {
14
54public:
56 static constexpr double kXThreshold = 2.4;
58 static constexpr double kThetaThreshold = 0.20943951;
73 static constexpr double kActionRangeTolerance = 1e-4;
74
84 explicit ContinuousCartPoleEnv(DeviceBackend* backend, int64_t max_steps = 200, uint32_t seed = 42);
85
95 [[nodiscard]] Tensor reset() override;
96
107 [[nodiscard]] Tensor reset(const Tensor& initial_state) override;
108
130 [[nodiscard]] StepResult step(const Tensor& action) override;
131
133 [[nodiscard]] int64_t observation_dim() const override { return 4; }
134
136 [[nodiscard]] int64_t action_dim() const override { return 1; }
137
139 [[nodiscard]] bool is_discrete() const override { return false; }
140
142 [[nodiscard]] int64_t step_count() const { return step_count_; }
143
145 [[nodiscard]] int64_t max_steps() const { return max_steps_; }
146
147private:
149 [[nodiscard]] Tensor observation() const;
150
152 [[nodiscard]] double next_uniform();
153
154 DeviceBackend* backend_;
155 int64_t max_steps_;
156 uint32_t lcg_state_;
157
158 double x_ = 0.0;
159 double x_dot_ = 0.0;
160 double theta_ = 0.0;
161 double theta_dot_ = 0.0;
162
163 int64_t step_count_ = 0;
164 bool has_reset_ = false;
165};
166
167} // namespace pulsatrix
The cart-pole balancing task with a continuous action: the action is a single force fraction in [-1,...
Definition continuous_cartpole_env.hpp:53
int64_t max_steps() const
Episode length limit this environment was constructed with.
Definition continuous_cartpole_env.hpp:145
static constexpr double kActionRangeTolerance
How far outside [-1, 1] an action may sit before step() rejects it.
Definition continuous_cartpole_env.hpp:73
ContinuousCartPoleEnv(DeviceBackend *backend, int64_t max_steps=200, uint32_t seed=42)
Constructs a fresh, not-yet-reset continuous cart-pole environment.
int64_t action_dim() const override
1 – a single continuous force fraction.
Definition continuous_cartpole_env.hpp:136
bool is_discrete() const override
False – this environment's action space is continuous.
Definition continuous_cartpole_env.hpp:139
static constexpr double kThetaThreshold
Pole angle (radians) past which the episode terminates (~12 degrees).
Definition continuous_cartpole_env.hpp:58
Tensor reset(const Tensor &initial_state) override
Starts a new episode from an exact caller-supplied state, bypassing the LCG.
StepResult step(const Tensor &action) override
Advances the physics one timestep (tau = 0.02 s) under the given continuous action.
int64_t step_count() const
Steps taken since the last reset().
Definition continuous_cartpole_env.hpp:142
static constexpr double kXThreshold
Cart position past which the episode terminates.
Definition continuous_cartpole_env.hpp:56
Tensor reset() override
Starts a new episode from a small random state: each of the 4 state variables drawn i....
int64_t observation_dim() const override
4 – (x, x_dot, theta, theta_dot).
Definition continuous_cartpole_env.hpp:133
Vendor-agnostic compute/memory backend. CPUBackend, CUDABackend (Phase 1.5), and HIPBackend (Phase 1....
Definition device_backend.hpp:219
Base class for every RL environment (CartPoleEnv, and whatever later phases add).
Definition environment.hpp:52
N-dimensional tensor. Owns its data buffer exclusively; a DeviceBackend* is injected (not owned) – th...
Definition tensor.hpp:29
Abstract interface isolating vendor-specific memory/compute operations from Tensor/ComputationGraph.
Abstract RL environment interface (gymnasium-shaped reset/step) + StepResult.
Definition acquisition_functions.hpp:16
What one Environment::step() produces: the next observation, this step's reward, and whether the epis...
Definition environment.hpp:25
N-dimensional tensor – owns a buffer via DeviceBackend*, RAII (Rule of Five).