pulsatrix
Loading...
Searching...
No Matches
evolution_strategies.hpp File Reference

Evolution Strategies (Salimans et al. 2017, "Evolution Strategies as a Scalable Alternative to Reinforcement Learning"): a black-box, gradient-free optimizer over an arbitrary flat parameter vector, driven entirely by a scalar fitness function – zero inherent RL dependency, applicable to any fixed-topology network's flattened weight vector or any other real-valued parameterization. More...

#include <cstddef>
#include <random>
#include <stdexcept>
#include <vector>
Include dependency graph for evolution_strategies.hpp:

Go to the source code of this file.

Classes

struct  pulsatrix::ESResult
 Result of a full Evolution Strategies run. More...
 

Namespaces

namespace  pulsatrix
 

Functions

std::vector< double > pulsatrix::ESUpdateGivenPerturbations (const std::vector< double > &theta, const std::vector< std::vector< double > > &epsilons, const std::vector< double > &fitnesses, double alpha, double sigma)
 Pure core: the exact ES parameter update given already-sampled perturbations and their fitness scores – ‘theta’ = theta + (alpha / (N*sigma)) * sum_i(F_i * epsilon_i)`.
 
template<typename FitnessFn , typename RNG >
std::vector< double > pulsatrix::ESStep (const std::vector< double > &theta, FitnessFn fitness_fn, int population_size, double sigma, double alpha, RNG &rng)
 RNG-driven wrapper: samples population_size/2 standard-normal perturbation vectors, scores theta+sigma*epsilon and theta-sigma*epsilon for each (mirrored sampling), and returns the updated theta via ESUpdateGivenPerturbations.
 
template<typename FitnessFn , typename RNG >
ESResult pulsatrix::RunEvolutionStrategies (std::vector< double > theta, FitnessFn fitness_fn, int num_iterations, int population_size, double sigma, double alpha, RNG &rng)
 Runs num_iterations of Evolution Strategies starting from theta, tracking the best (theta, fitness) pair seen across every evaluated center point (global elitism, the same precedent this campaign's NEAT evolutionary loop already established) – vanilla ES itself has no such tracking, but a returnable "best point found" is genuinely necessary for this to be usable as an optimizer, so it is added here as a small, logged addition beyond the paper's own bare update rule.
 

Detailed Description

Evolution Strategies (Salimans et al. 2017, "Evolution Strategies as a Scalable Alternative to Reinforcement Learning"): a black-box, gradient-free optimizer over an arbitrary flat parameter vector, driven entirely by a scalar fitness function – zero inherent RL dependency, applicable to any fixed-topology network's flattened weight vector or any other real-valued parameterization.

Note
Uses mirrored (antithetic) sampling – every sampled perturbation epsilon is scored alongside its negation -epsilon – matching Salimans et al.'s own standard variance-reduction technique, not an invented addition. This is why population_size must be even.