pulsatrix
Loading...
Searching...
No Matches
Reinforcement Learning

Environments, agents, and training algorithms (DQN, REINFORCE, A2C, PPO, SAC) built on the Deep Learning Modules and Layers tensor/autograd core. More...

Files

file  agent.hpp
 Abstract RL agent interface – the inference-time policy contract, act() only.
 
file  cartpole_env.hpp
 CartPole-v1 environment – classic cart-pole balancing physics, discrete actions.
 
file  categorical_policy_agent.hpp
 Stochastic categorical (discrete-action) policy over a logit-producing Module.
 
file  continuous_cartpole_env.hpp
 Continuous-action variant of CartPoleEnv – identical physics, force is a fraction.
 
file  dqn_agent.hpp
 Epsilon-greedy DQN policy over an arbitrary Q-network Module.
 
file  dqn_loss.hpp
 DQN's masked-MSE Bellman loss – squared error on the taken action only.
 
file  dqn_target.hpp
 DQN Bellman target computation (vanilla + Double DQN) and target-network hard sync.
 
file  environment.hpp
 Abstract RL environment interface (gymnasium-shaped reset/step) + StepResult.
 
file  gae.hpp
 Generalized Advantage Estimation (Schulman et al. 2016) over a stored rollout.
 
file  policy_gradient_loss.hpp
 REINFORCE's return-weighted negative log-likelihood loss over a batched rollout.
 
file  polyak_update.hpp
 Soft (Polyak / exponential-moving-average) target-network update, as used by SAC.
 
file  ppo_clipped_loss.hpp
 PPO's clipped surrogate objective (Schulman et al. 2017) over a batched rollout.
 
file  replay_buffer.hpp
 Off-policy experience replay: fixed-capacity circular transition store + sampling.
 
file  rollout_buffer.hpp
 On-policy trajectory storage: fill-once fixed-length rollout + discounted returns.
 
file  tanh_gaussian_policy.hpp
 SAC's reparameterized, tanh-squashed Gaussian policy sampling (action + log-prob).
 

Detailed Description

Environments, agents, and training algorithms (DQN, REINFORCE, A2C, PPO, SAC) built on the Deep Learning Modules and Layers tensor/autograd core.