Reinforcement Learning¶
Use this section when you want an agent to learn by trial and error: it acts in an environment, receives rewards, and improves its policy (the rule that picks actions). Pulsatrix ships five algorithms: DQN (with Double DQN), REINFORCE, A2C, PPO, and SAC.
Every network here is a plain Module from
Deep Learning Modules and Layers. You train it with the same
forward()/backward()/optimizer loop as any other network. The environment and agent
interfaces follow the shape of Gymnasium (the standard Python RL toolkit): reset(), step(),
and act().
Which algorithm should I use?¶
| Algorithm | Action space | Learns from | Use it when |
|---|---|---|---|
| DQN / Double DQN | Discrete | Replayed past experience (off-policy) | You have a small, discrete action set and want sample efficiency. |
| REINFORCE | Discrete | Its own latest episodes (on-policy) | You want the simplest policy-gradient baseline. |
| A2C | Discrete | Its own latest rollouts (on-policy) | You want lower variance than REINFORCE by adding a critic. |
| PPO | Discrete | Its own latest rollouts (on-policy) | You want a stable, general-purpose default. |
| SAC | Continuous | Replayed past experience (off-policy) | Your actions are real numbers (e.g. a force or a torque). |
Each algorithm has an integration test that must reach a fixed score on CartPole
(tests/*_cartpole_integration_test.cpp).
What's inside¶
- Environments:
Environment(interface),CartPoleEnv,ContinuousCartPoleEnv - Agents:
Agent(interface),DQNAgent,CategoricalPolicyAgent(REINFORCE/A2C/PPO) - Policies:
TanhGaussianPolicy, SAC's continuous-action sampler. It is not aModuleor anAgent: it takes three tensors and returns two. - Buffers:
ReplayBuffer(off-policy: DQN, SAC),RolloutBuffer(on-policy: REINFORCE/A2C/PPO) - Value estimation:
ComputeGAE(Generalized Advantage Estimation),ComputeDQNTarget/ComputeDoubleDQNTarget - Target-network updates:
SyncTargetNetwork(hard periodic copy, used by DQN),PolyakUpdate(small blend every step, used by SAC) - Losses:
DQNLoss,PolicyGradientLoss(REINFORCE/A2C),PPOClippedLoss
Full API reference: Doxygen: Reinforcement Learning
Runnable demos¶
Each algorithm has a demo target, built by default (PULSATRIX_BUILD_EXAMPLES=ON):
| Target | Algorithm |
|---|---|
dqn_cartpole_demo |
Double DQN on CartPoleEnv |
reinforce_cartpole_demo |
REINFORCE on CartPoleEnv |
a2c_cartpole_demo |
A2C on CartPoleEnv |
ppo_cartpole_demo |
PPO on CartPoleEnv |
sac_continuous_cartpole_demo |
SAC on ContinuousCartPoleEnv |
cmake --build build --target ppo_cartpole_demo --config Release
How to implement¶
The environment/agent loop¶
#include "pulsatrix/cartpole_env.hpp"
#include "pulsatrix/categorical_policy_agent.hpp"
#include "pulsatrix/cpu_backend.hpp"
#include "pulsatrix/linear_module.hpp"
using namespace pulsatrix;
CPUBackend backend;
CartPoleEnv env(&backend); // 4-dim observation, 2 discrete actions
LinearModule policy_net(4, 2, &backend); // observation -> action logits
CategoricalPolicyAgent agent(&policy_net, /*action_dim=*/2, &backend);
Tensor obs = env.reset(); // shape (1, 4)
float episode_return = 0.0f;
for (bool done = false; !done;) {
Tensor action = agent.act(obs); // samples an action from the policy
StepResult step = env.step(action);
episode_return += step.reward;
obs = step.observation;
done = step.done;
}
What's happening: reset() starts an episode and returns the first observation. Each
step() applies one action and returns a StepResult: the next observation, the reward, and
whether the episode ended. Training adds a buffer and a loss around this loop. The demos above
show the full version for each algorithm.
Generalized Advantage Estimation¶
An advantage measures how much better an action turned out than the critic (a network that predicts future reward) expected. Policy-gradient methods like A2C and PPO weight each update by it.
#include "pulsatrix/cpu_backend.hpp"
#include "pulsatrix/gae.hpp"
using namespace pulsatrix;
CPUBackend backend;
// One row per rollout step, each shape (N, 1). Here N = 3 and the last step ends the episode.
Tensor rewards(Shape({3, 1}), &backend, {1.0f, 1.0f, 1.0f});
Tensor dones(Shape({3, 1}), &backend, {0.0f, 0.0f, 1.0f});
Tensor values(Shape({3, 1}), &backend, {2.5f, 1.8f, 0.9f}); // critic's V(s_t)
GAEResult gae = ComputeGAE(rewards, dones, values, /*bootstrap_value=*/0.0f,
/*gamma=*/0.99f, /*lambda=*/0.95f, &backend);
// gae.advantages: how much better each action was than the critic expected (actor weight)
// gae.returns: advantages + values, the critic's regression target
What's happening: ComputeGAE walks the rollout backward once. At each step it computes
the one-step prediction error delta[t] = r[t] + gamma * V[t+1] - V[t] and adds it to a
decaying running sum. lambda sets the trade-off:
lambda == 0uses only the one-step error: low variance, but biased by critic mistakes.lambda == 1uses the full episode return: unbiased, but high variance.
Most training loops pick something in between; 0.95 is PPO's usual default. A done flag
stops the sum from crossing into the next episode.
Recipe: GAE and PPO's clipped objective.
Soft target-network updates¶
Off-policy methods train against a target network: a slowly changing copy of the critic that keeps the learning target stable.
#include "pulsatrix/cpu_backend.hpp"
#include "pulsatrix/linear_module.hpp"
#include "pulsatrix/polyak_update.hpp"
using namespace pulsatrix;
CPUBackend backend;
LinearModule online_critic(4, 1, &backend);
LinearModule target_critic(4, 1, &backend); // same architecture
PolyakUpdate(online_critic, target_critic, /*tau=*/0.005f);
// target_critic moves 0.5% of the way toward online_critic; call this after every update
What's happening: PolyakUpdate blends each parameter in place:
target[i] = tau * online[i] + (1 - tau) * target[i]. The target drifts a little on every
step. SyncTargetNetwork instead copies the online network outright every k steps, so the
target stays frozen and then jumps. SAC (like DDPG and TD3) uses the soft blend; DQN uses the
hard copy. tau must be in (0, 1], and both networks must have matching parameter shapes.
The SAC demo (sac_continuous_cartpole_demo, training loop in
examples/sac_continuous_cartpole_training.hpp) calls SyncTargetNetwork once at start-up and
PolyakUpdate after every update. The DQN on CartPole recipe
uses the periodic SyncTargetNetwork copy.