|
pulsatrix
|
Environments, agents, and training algorithms (DQN, REINFORCE, A2C, PPO, SAC) built on the Deep Learning Modules and Layers tensor/autograd core. More...
Files | |
| file | agent.hpp |
| Abstract RL agent interface – the inference-time policy contract, act() only. | |
| file | cartpole_env.hpp |
| CartPole-v1 environment – classic cart-pole balancing physics, discrete actions. | |
| file | categorical_policy_agent.hpp |
| Stochastic categorical (discrete-action) policy over a logit-producing Module. | |
| file | continuous_cartpole_env.hpp |
| Continuous-action variant of CartPoleEnv – identical physics, force is a fraction. | |
| file | dqn_agent.hpp |
| Epsilon-greedy DQN policy over an arbitrary Q-network Module. | |
| file | dqn_loss.hpp |
| DQN's masked-MSE Bellman loss – squared error on the taken action only. | |
| file | dqn_target.hpp |
| DQN Bellman target computation (vanilla + Double DQN) and target-network hard sync. | |
| file | environment.hpp |
| Abstract RL environment interface (gymnasium-shaped reset/step) + StepResult. | |
| file | gae.hpp |
| Generalized Advantage Estimation (Schulman et al. 2016) over a stored rollout. | |
| file | policy_gradient_loss.hpp |
| REINFORCE's return-weighted negative log-likelihood loss over a batched rollout. | |
| file | polyak_update.hpp |
| Soft (Polyak / exponential-moving-average) target-network update, as used by SAC. | |
| file | ppo_clipped_loss.hpp |
| PPO's clipped surrogate objective (Schulman et al. 2017) over a batched rollout. | |
| file | replay_buffer.hpp |
| Off-policy experience replay: fixed-capacity circular transition store + sampling. | |
| file | rollout_buffer.hpp |
| On-policy trajectory storage: fill-once fixed-length rollout + discounted returns. | |
| file | tanh_gaussian_policy.hpp |
| SAC's reparameterized, tanh-squashed Gaussian policy sampling (action + log-prob). | |
Environments, agents, and training algorithms (DQN, REINFORCE, A2C, PPO, SAC) built on the Deep Learning Modules and Layers tensor/autograd core.